Mission Specification: R1a — the guard half: freeze the SPEC_KITTY_HOME behaviour class
Mission Branch: feat/isolated-home-pin-guard Mission Slug: isolated-home-pin-guard-r1a-01KZNMA3 Created: 2026-08-10 Status: Draft — revision 2, post-gate. Three profile-loaded lenses failed the first revision with 13 blockers between them. All three independently reproduced its population exactly and found no fabricated number; every blocker was about an instrument, not a count. The operator's remediation decision — drop the @pytest.fixture limb — cascades through every figure below, so this revision re-derives all of them. Baseline: upstream/main @ 5d49d31ed6505627d98d8f95d8502c9bf6a2f5ac Landing base: upstream/main @ f6b90d34ef07f959fb6f6365dde884bf9bbf2474 (188 commits newer; the census was regenerated against it with the key set and both baseline hashes byte-identical) Relates to: #3121 Input: The operator's approved re-scope of isolated-home-pin-convergence-01KZCTWC, which halted on its own pre-convergence gate (|P| = 5). Record of the halt: spec.md amendment A21, ADR docs/adr/3.x/2026-08-07-1-a-mission-halting-instrument-is-worth-its-cost.md — which at specification time did NOT exist on feat/isolated-home-pin-guard; it existed only on spike/isolated-home-3121, and the WP that owns C-006's "halt-path ADR" must IMPORT it verbatim rather than author a new one, or R1a lands a second, divergent record of one halt — (Resolved at landing: imported verbatim, body-only sha256 identical to the spike's copy, with one added description: key for the docs frontmatter gate; registered in docs/adr/3.x/index.md and both generated docs inventories.) —, and https://github.com/Priivacy-ai/spec-kitty/issues/3121#issuecomment-5215920230. R1a is the guard half. R1b — adoption and adjudication — is deferred.
0. Why R1a exists, and the argument it is built on
0.1 THE OPENING ARGUMENT — dispersion, not site count, and the guard covers its class completely
Every figure in this section was re-derived from upstream/main @ 5d49d31ed by binding-resolving AST for this revision. Text search established nothing; it appears only as the pre-filter of C-003, whose premise is now stated and proved rather than assumed.
There are 191 SPEC_KITTY_HOME write sites under tests/ — 188 setenv calls in 83 files, 3 os.environ[...] = assignments, and 0 setdefault. They resolve into value buckets, and the buckets are the argument:
| Resolved value | Sites | Files | Where they live |
|---|---|---|---|
tmp_path/"global" | 42 | 3 | 42 test bodies, 0 fixtures |
tmp_path/"home" — the behaviour class | 40 | 36 | 30 fixtures, 10 test bodies, 0 helpers |
tmp_path/"kittify" | 17 | 4 | 16 test bodies, 1 fixture |
bare tmp_path | 10 | 10 | 7 fixtures, 3 test bodies |
tmp_path/"home"/".kittify" | 8 | 2 | 8 test bodies |
tmp_path/"consent-home" | 7 | 7 | 6 fixtures, 1 test body |
tmp_path/"global" is the LARGEST bucket — larger than the class this Mission freezes — and it is not the right target. It is 42 sites in 3 files. Three files is an afternoon's edit by one person who can hold all of it in their head; it needs no owner, no census and no guard. tmp_path/"home" is 40 sites across 36 files, in 6 directories, under 9 different fixture names, maintained by whoever last touched each file.
Dispersion, not site count, is what makes this the right convergence target. A count of sites measures how much text there is; the number of files measures how many independent decisions have to agree. That is the argument the first revision never made, and it is stronger than the one it did make.
The guard covers its class completely. Under the predicate of FR-001 — silhouette evaluated over the enclosing scope chain, no decorator limb — the classifier discovers 40 members in 36 files, which is every site and every file of the effect class. The three labelled reach figures, which FR-008 must publish separately and never merge into one:
| Figure | Value | What it means |
|---|---|---|
| members / effect-class sites | 40 / 40 = 100% | Every tmp_path/"home" write site is inside a member |
| members / effect-class files | 36 / 36 = 100% | Every file carrying one is covered |
members / all SPEC_KITTY_HOME pin sites | 40 / 191 = 20.9% | The class is a fifth of all home pinning; the rest is other buckets, out of scope by definition |
The figure the first revision led with — "49 of 188 — 26%" — is struck as wrong in both numerator and denominator, and it was stated as operator decision 1's entire justification. Of the 49 pin-bearing fixtures, 19 are not members (7 pin bare tmp_path, 6 pin tmp_path/"consent-home", and one each "kittify", "global_kittify", "_home", "trigger-home", "runtime-home", "spec-kitty-home"); and 188 counts setenv only, excluding the 3 environ-assignments. 26.06% was a ratio between two populations, neither of which is the guard's domain. Publishing one ratio that is none of the three above is how an inherited figure becomes a gate.
0.1a Why the @pytest.fixture limb was dropped — the limb that excluded everything the silhouette did not
The first revision's predicate required a @pytest.fixture decorator, and it defended the silhouette limb at length (A20(a), C-004) while never interrogating the decorator limb at all. Measured:
- The silhouette limb excludes 1 effect site under innermost attribution, and none at all under the scope-chain evaluation FR-001 actually binds (§0.1b).
- The decorator limb excluded 10 — every one of the ten non-fixture effect sites satisfies the superset silhouette at its keyed def. The decorator was doing all of the exclusion, and the silhouette none of it.
The decorator limb was a shape key, and re-keying the population from shape to effect is precisely what the parent Mission's name→effect correction existed to accomplish. Carrying a decorator limb re-introduced the same class of error one axis over.
Dropping it takes the class from 30 to 40 members, reach from 75.0% to 100% by site and from 83.3% to 100% by file, costs 10 census rows, and closes the "one decorator removal away" escape by construction. No new green path is needed: a test function requests the owner by parameter exactly as a fixture does.
Two claims the first revision made are struck as FALSE:
1. ~~"structurally invisible to any form of this guard"~~ — false, and Non-Goal 5 and SC-000 both rested on it. The true and stronger statement: all 10 of the 10 non-fixture effect sites carry the full (tmp_path, monkeypatch) silhouette at their keyed def, and were excluded by the decorator limb alone. They were not invisible; they were filtered out. (Measured at the innermost def the figure is 9 of 10 — the reading C-004 refuses, and the one an earlier revision published.) 2. ~~"three undecorated helpers … one decorator away from membership"~~ — false. Re-derived, the three resolve to tmp_path/"hangfix-home" (_stub_dossier_resolvers), tmp_path/<case>-home/".spec-kitty" (_make_checkout_with_grant) and tmp_path/"env-root" (_configure_root). Adding @pytest.fixture makes them fixtures, not members — they still fail the effect limb. The escape half of that sentence was true and is now closed by the limb drop.
0.1b The silhouette is evaluated over the SCOPE CHAIN, not the innermost def — the fourteenth finding
Three lenses found 13 blockers. This is a fourteenth, found while re-deriving the population for this revision, and it is the reason reach is 100% rather than 97.5%.
tests/sync/tracker/test_tracker_egress_refusal_3108.py:1165 sits inside _run_once, a nested closure whose own parameter list is (install_counter,). Attributed to the innermost def, it fails the silhouette and is the class's single blind-spot site — 39/40. But _run_once obtains tmp_path from its enclosing test function's parameter, by closure; that is precisely why its pin resolves to tmp_path/"home" at all. The site is not outside the silhouette; it is inside it, one scope up.
Therefore FR-001 evaluates the silhouette over the UNION of the parameter sets along the enclosing def chain, and keys the member by the OUTERMOST def in that chain which satisfies it. Re-derived: this yields 40 members rather than 39, changes nothing else (the symmetric difference against innermost attribution is exactly that one site), and takes site reach to 40/40 = 100%.
Left unstated, this was a live defeat waiting: an author who wants a pin the guard cannot see writes it in a two-line nested closure, and a spec that never named its attribution rule would have been read whichever way was cheapest at implementation time.
And :1165 is held in the class by a token nobody uses. Re-derived: its enclosing test declares monkeypatch in its signature and never references it — tmp_path is used, monkeypatch is not — and it is the only one of the 40 members held in by an unused silhouette parameter. Delete that token, which is routine cleanup, and the site leaves the class with its behaviour completely unchanged. Nothing will stop it: ruff's ARG rule is relaxed for tests/* (pyproject.toml:246), so no automation flags the unused argument — a human will remove it. This is recorded, not repaired: R1a cannot edit the module (C-001), and widening the predicate to catch it would mean inferring the silhouette from usage* rather than declaration, which is a larger change than R1a should make. It is the class's single most fragile membership, and the receiver form that would produce a miss (pytest.MonkeyPatch.context()) already landed in this same window, in #3108, and satisfied the predicate by coincidence.
0.2 The five operator decisions — settled, recorded, not open
| # | Decision | Justification of record |
|---|---|---|
| 1 | Guard-first stands. | Dispersion and complete coverage of the class — §0.1. Not any growth rate (§0.3), and not the struck 26% figure. |
| 2 | Census form, with the entitles-nothing check written in. | §0.4. |
| 3 | Both budgets re-derived from measurement; the byte pre-filter is MANDATORY. | §0.6, NFR-001, NFR-002. |
| 4 | A Mission, FIVE work packages, plus two carve-outs. | §0.8. Decision 4 was taken as "four work packages" and amended in the same table by WP-0, which precedes and gates the other four (§0.9). Stated as five here because a self-measuring row that says four while its own amendment says five is the shape §0.2's title carried too. The four are unchanged in scope and order. |
| 5 | Drop the @pytest.fixture limb. (post-gate remediation decision) | §0.1a. Cascades through every figure in this specification, all of which are re-derived here. |
0.3 PROVENANCE — there is no observed arrival rate, and stating one would be false
Two members arrived since the halted Mission's baseline. A previous framing described this as the class being re-created at a rate. That framing is false and must not be repeated.
| Member | Arrival commit | Authored | Committed | PR |
|---|---|---|---|---|
tests/cli/commands/test_sync_doctor_tracker_egress_3108.py::doctor_environment | 970852644 | 2026-08-04 | 2026-08-10 | #3108 |
tests/sync/tracker/test_tracker_egress_refusal_3108.py::_isolated_home_and_arming | 161a0c179 | 2026-08-04 | 2026-08-10 | #3108 |
Both were members at the commits that added them. 971fa0e3e (authored 2026-08-07) and b88d47728 (authored 2026-08-10) are last-touch, not arrival; all four share committer date 2026-08-10 because the branch was rebased before landing. Reading committer dates as arrival dates manufactures a four-day spread out of one event from one PR.
There is no observed arrival rate. Two points from one landing do not define one, and this specification refuses to compute one.
What is true and sufficient: the class grew 28 → 30 in a single landing the halted baseline could not see, and both arrivals are partition-B1 trap cases — verified against M4's labels in the plan phase — and the HOME orphaned-binding trap rose from 9 to 11. Growth landed in the expensive partition. (Corrected: this read "rose from 7 to 9", low by two at both endpoints. M4 measures B1 = 9 over its 28 (VERDICT.md:40, corroborated at RESIDUALS.md:58), and both arrivals are B1, so the trap went 9 -> 11. The figure had been cited in prose for five passes against a definition that lived only on spike/isolated-home-3121; importing the artefact is what falsified it. Over the current 40 the count is 11.)
Scoping note, binding on §0.9. The 28 → 30 figure was measured under the superseded decorator-limbed predicate. It is retained as the historical fact that motivated freezing and is labelled as such. The class's growth under the current predicate is unmeasured, deliberately — measuring it at the window start SHA is WP-0's job, and doing it here would fix §0.9's threshold after seeing its own measurement.
0.4 The census is not a manifest, and the difference is the whole design
R1a ships a frozen census: a sorted, checked-in set of 40 rows, one per measured member, under a header carrying the single frozen_at_sha and the single owed_to (FR-003 — they are scalars, not columns, because 40 rows held exactly two distinct values between them). Each row carries its MemberKey, a non-authoritative lineno, kind and home_partition. There is no reason column, and its absence is load-bearing.
The canonical owner is NOT a census row. It is a named singleton (FR-005), and the guard asserts discovered == census ∪ E (FR-004), where E is the enumerated exemption set: the owner, plus the single retained-pin probe of FR-011. E is fixed-arity by type (tuple[Exempt, Exempt]), hash-pinned outside the guard module, and enumerated — it is not an open allowance:
- The census can reach zero, and that is R1b's definition of done. With the owner outside it,
census == ∅means R1b is complete — one line, checkable by anyone, with every removal accounted for as a deletion, an adoption, or a manifest row carrying a distinct measured cause. That last clause is not decoration: without it the cheapest path to an empty census is bulk-migrating all 40 rows into the manifest under boilerplate reasons, which empties the census while burning down nothing. Put the owner in and an empty census is unreachable by construction, so the burn-down has no terminus at all. - The owner's row would be a manifest row wearing the census schema. The owner must exist forever; its presence is a decision, not debt.
- A carve-out costs the reviewer test its force. The moment C-007 reads "entitles nothing, except…", the next row that wants an exception has a precedent.
| Manifest row (R1b) | Census row (R1a) | |
|---|---|---|
| Means | may exist forever, because \<measured cause\> | existed at SHA X, owed an adjudication |
| Is | a decision — terminal | debt — transient |
| Columns | reason, mirrored at the definition site | MemberKey, lineno (non-authoritative), kind, home_partition — no reason, and frozen_at_sha / owed_to hoisted to header scalars |
| Direction | may grow when a cause is measured | monotonically non-increasing, mechanised (FR-004) |
> *Reviewer test, binding on every review of R1a and R1b — and stated over anything that makes a definition acceptable to the guard, not merely over a census row: > What does this entitle its definition to?* > The answer must be nothing. > Stated over census rows alone, the test does not reach the owner singleton, the tombstone list, or any future allowance — which is exactly how an exemption relocates out of the audited artefact and into the guard's own source.
E FAILS that test, and saying so is the point. An entry in E entitles its definition to exist forever, with no owed_to, no frozen_at_sha and no tombstone — strictly more than a census row, and terminal rather than transient. The remedy is not to pretend otherwise: the owner must exist forever, so no honest answer of "nothing" is available. The remedy is that E cannot grow — fixed arity by type, hash-pinned outside the module that declares it, two entries, both named here. The reviewer test earned its keep by detecting E; what closes E is mechanism, not prose. E is this Mission's one manifest-shaped artefact, and C-007's "no exception" is about the CENSUS, which is why E is audited by a stricter mechanism instead of a looser one.
The cost, stated honestly. The predicate is strong; the semantic content is empty. The guard proves the class has not changed since SHA X. It does not prove any member deserves to be there. That is entirely R1b's.
0.5 The owner is not optional, and one of its three safety limbs is genuinely untested
Without an owner, the author of a legitimate 41st member has no green path but to write the pin in a test body — and post-limb-drop that is still a member, so it reds too. The owner is what turns "you are blocked" into "request this fixture".
The halted Mission's C-010 discharged the owner's precedence safety three ways. The first revision called all three vacuous. That was over-stated:
| Limb | Status in R1a |
|---|---|
(i) ScopeMismatch makes scope inversion loud at setup | Exercisable — the probe modules R1a already builds (FR-011) can request the owner at a higher scope and assert the raise |
| (ii) A non-autouse owner is never instantiated in a module that does not request it | Exercisable — assert the owner's fixture never runs for a module that does not name it |
| (iii) Every retained and adopting definition stays function-scoped, asserted per manifest entry | Genuinely vacuous — stated over adopters, and R1a has zero |
PARENT-ID CITATION CONVENTION — binding on this whole specification. Requirement IDs belonging to the halted parent Mission are written with a non-breaking hyphen (U+2011): FR‑015, NFR‑007, and the parent's C‑009 / C‑010. They look identical and will not match \b(?:FR|NFR|C)-\d+\b.
Why this is mandatory and not cosmetic — upstream #3170. The finalize-tasks requirement-ID scraper "cannot distinguish a mission's own IDs from citations of another mission's." Measured on this specification: the scraper harvests 33 IDs where R1a defines 30. Two failure modes, and the second is the dangerous one:
- Mode 1 — visible hard block.
FR‑015,NFR‑007and (before §C's fix)C-014are scraped as phantom requirements this Mission never defines. - Mode 2 — SILENT, and this is the one that mattered. R1a defines its OWN
C-009andC-010and also cited the parent's. #3170 documents that such citations are "silently attributed to this mission's same-numbered requirement, inflating apparent coverage" — so a work package that never discharges R1a'sC-009(pre-existing reds) orC-010(PR #3285 coordination) could show as covering them. The earlierC-014remediation fixed the human-readable collision and left the machine-readable one exactly as it was.
The character is invisible to a grep, which is precisely why this note is mandatory: without it, a future editor retypes a plain hyphen and silently restores mode 2. #3170 is already filed; no issue is created here (C-013).
A SECOND, DISTINCT ID HAZARD — this Mission's own vocabulary is prefix-overlapping by construction. Where the above concerns another Mission's IDs, this concerns R1a's own: measured over the 46 IDs in this document, 21 ordered pairs are substring-contained — 14 of the form C-0NN ⊂ SC-0NN (so every constraint C-001…C-013 can be reported covered by its same-numbered success criterion), 6 of the form FR-00N ⊂ NFR-00N, and 1 SC-002 ⊂ SC-002b. Any tool or claim that matches IDs must use word boundaries; a substring match falsely greens 21 of the 46. This specification states coverage in prose and mechanises none of it, so nothing here is currently exposed — the hazard is recorded so that the first artefact to mechanise a coverage or traceability property does not have to rediscover it.
These three limbs are R1a's C-014, and every criterion refers to them by that ID. They originate in the halted parent's C‑010; R1a's own C-010 is "PR #3285 is a coordination dependency of R1b" and has no limbs, so "C-010 limbs (i) and (ii)" resolved two ways inside one document and an implementer reading SC-011 literally builds nothing. The same two-namespace collision affects the parent's C‑009 as cited in C-005, and FR‑015 / NFR‑007 are referenced but never defined here — all of them the parent's. Where a parent ID is meant, this specification now writes "the parent's <ID>".
The residual is limb (iii) — and one more, which had been carried without ever being named. SC-012's retained-pin probe is the single real-tree member in E, and it therefore consumes ONE OF ONLY TWO irrevocable, hash-pinned exemption slots — a slot that by construction "cannot grow" — for an assertion that, as first written, could not bite (see SC-012). That is this Mission's most irreversible decision and it was absent from every residual list. It is recorded here: the slot is spent, the repair makes the assertion falsifiable, and no third slot exists if the design is later found wanting. Both residuals are known costs of guard-first sequencing rather than discoveries left to R1b. SC-011 binds C-014 (i) and (ii).
0.6 The two budgets, re-derived from measurement
The recorded 4.47–4.54 s figure never measured the guard — it benchmarked bare ast.parse + ast.walk with no classification. Both budgets are re-derived from a full classifier over all 2737 files, warm, three runs, on two independent machines:
| Pass | Operator's hardware | Re-derived here (Python 3.12.13, pytest 9.0.3 venv) |
|---|---|---|
| Unfiltered classifier (parses all 2737) | 28.47 / 28.69 / 29.21 s | 17.64 / 17.47 / 17.34 s |
| Pre-filtered (parses the 100 byte-hit files) | 1.83 / 1.86 / 1.86 s | 0.986 / 0.991 / 1.003 s |
| Byte scan alone, no parsing | 0.22 s | 0.057 / 0.055 / 0.056 s |
| Member sets, pre-filtered vs unfiltered | symmetric difference EMPTY | symmetric difference EMPTY |
The old 2 s budget is violated 14× by the unfiltered classifier and was met on the operator's hardware only at 1.86 s — 7% headroom, which will not hold on a runner, and whose cheapest recovery is narrowing the walk. That is the A6/A8 defeat.
MEASURED FIGURES, post-tasks — these supersede the estimates above, and the denominators are published with them.
| Quantity | Measured | Against |
|---|---|---|
Guard scan, including the HOME limb and partitions | 0.88 s (3.14) · 1.24 s (3.11) | NFR-001's 6 s — about 7x / 4.8x headroom |
| SC-002 both-passes differential (0.684 s pre-filtered + 8.36 s unfiltered) | 9.0 s | NFR-002's 90 s — about 10x |
Full two-SHA gate driver (git archive x2 + discover() x2) | 1.86 s | not budgeted; the driver is one-shot |
git archive is the correct extraction: git worktree add is slower, mutates .git, and needs failure cleanup.
These are a FLOOR, not the CI figure. They were taken warm on a workstation under Python 3.14; arch-adversarial runs 3.12 on 4 vCPU with -n auto and coverage. They are the right figure for the timing-marked module, which runs -n0 without --cov (OD-003).
Re-measured at landing on the rebased tree (base f6b90d34e) under Python 3.11.15, five consecutive discover() runs: cold 1.211 s, warm worst 1.244 s, warm mean 1.226 s, 42 members — 4.82x headroom against NFR-001's 6 s. The interpreter, not the tree, accounts for the gap from the 3.14 figure; both are recorded above rather than one overwriting the other, because "which Python" is exactly the kind of denominator this table exists to publish. The margin narrows from 7x to 4.8x on the older interpreter and NFR-001 still holds with room.
Three inconsistent numbers existed for one quantity — §0.6/§0.7's 1.83-1.86 s / 0.986-1.003 s, a task's 0.056 s, and the measured 0.684 s / 8.36 s. The measured pair above is the one that stands; the argument for the pre-filter never depended on which, and does not change.
Budget (a) — the guard: 6 s warm, gating. Slowest measured pre-filtered figure (1.86 s) × 3, being the 1.86× observed two-machine spread doubled for a CI runner that has not been measured. The doubling is an assumption, flagged; OD-003 obliges the plan to measure a runner. The budget may be RAISED with runner evidence. The walk may never be narrowed.
Budget (b) — the coverage proof: 90 s warm, separable, non-gating. Slowest measured both-passes figure (30.5 s) × 3. (30.5 × 3 is 91.5, not 90 — the stated budget is 1.5 s tighter than its own derivation. Recorded in the plan phase rather than silently re-derived: the direction is conservative, so the figure stands as the binding one, but a published number that does not follow from its published rule is the shape this specification keeps catching.) Explicitly not bound by (a). Narrowing or skipping (b) to fit (a) relocates the defeat into the proof of the guard. OD-002 carries a sub-second alternative that is strictly weaker, and its weakness is stated there because 90 s is exactly the number that makes a weaker proof attractive.
No precedent exists in this repository for an executable pre-filter coverage proof. Two byte pre-filters ahead of ast.parse already ship — tests/architectural/_sole_door_scan.py:461-476 and tests/architectural/test_commit_target_kind_guard.py:186-188 — and both argue soundness in a comment and prove nothing executably. R1a's proof obligation is therefore stricter than anything the repo currently does, and that is deliberate: §0.7's pre-filter premise is the one this Mission was about to get wrong (§0.6a).
0.6a The pre-filter's premise was unstated, and it is now stated and proved
FR-002's soundness argument — a member pins SPEC_KITTY_HOME and must therefore contain that byte string — rests on an unstated premise: that the env key is a string literal at the call site.
Re-derived: of 1245 setenv/delenv calls under tests/, 73 use a non-literal key (70 Name, 3 Attribute) — SAAS_SYNC_ENV_VAR, OPT_OUT_ENV_VAR, PROVIDER_SELECT_ENV_VAR, _PACKS_ROOT_ENV, and tests/conftest.py:292 and :296 themselves. The indirection idiom is established and at scale; it simply has not been applied to this key yet — which is exactly the argument §0.1 makes to justify the whole Mission.
A member written monkeypatch.setenv(RUNTIME_HOME_ENV, str(tmp_path / "home")) contains no contiguous SPEC_KITTY_HOME bytes. The pre-filter drops the file, ast.parse never runs, the guard greens.
And the 90 s both-passes proof cannot detect it. Both passes run the same classifier; if it cannot see key-indirection members, neither pass finds them and the symmetric difference is empty by construction. OD-002 rejects its cheap candidate for "assuming the member set it is meant to help establish" — the expensive candidate assumes the classifier's member set one level up. With respect to this hole, 90 s buys nothing over 0.056 s.
Discharge — measured, and it is C-003 form (ii), not either option first considered. The population of the indirection is 0: re-derived across src/ and tests/ together, there are zero string constants whose value is "SPEC_KITTY_HOME". Nothing can be imported to indirect the key, because nothing holds it. So the premise is not merely stated — it is provable as an empty-set assertion, which C-003 form (ii) already admits as conclusive, and which ships as an executable test (FR-002, SC-002b). The hole is dormant, not live, and the assertion is what keeps it dormant: the day someone writes RUNTIME_HOME_ENV = "SPEC_KITTY_HOME", the empty-set test reds and the pre-filter must widen before the guard can be trusted again.
0.7 What R1a inherits, unchanged
| Inherited item | Status here |
|---|---|
A20(a) — the silhouette limb is a SUPERSET test over the parameter set, leading self stripped; never arity-exact, never order-sensitive | [re-derived] — extended to the enclosing scope chain (§0.1b) |
| A20(b) — SC-006 transition 6, and the C-010 manifest obligation: narrowing the silhouette is not an acceptable repair | Carried as C-004, SC-006 |
| A6/A8 — the mandatory pre-filter, its over-selection proof, two budgets, only (a) gating | [re-derived] — §0.6, and its premise now stated and proved (§0.6a) |
§0.7 UPHELD — the _isolated_worker_home precedence decision | [re-derived] — tests/conftest.py is byte-identical in full between spike/isolated-home-3121 and this baseline. Carried as C-005, NFR-005, FR-006 |
OD-002 M = 1; the baseid derived at runtime, never a literal | [INHERITED — not re-derived in this pass]. Carried as FR-005, NFR-006 |
Absent scope= is function scope | [re-derived] — 0 of 49 pin-bearing fixtures carries an explicit scope=. Carried as FR-007 |
Anchors, re-derived: _isolated_worker_home occupies tests/conftest.py 253–298 (decorator 253, def 254, body end 298); its precedence docstring 272–286; _enable_saas_sync_feature_flag 301–304.
0.8 The population, re-derived on 5d49d31ed
Membership predicate — binding, and the subject of FR-001. A def or async def at any nesting depth, with no decorator requirement, the union of whose enclosing def chain's parameter sets contains {tmp_path, monkeypatch} (leading self stripped), whose body writes SPEC_KITTY_HOME to a value resolving — through single-assignment local bindings, with str / Path / os.fspath unwrapping and f-string joining — to tmp_path / "home". The member is keyed by the outermost def in the chain satisfying the silhouette. autouse is not a limb. An absent scope= is function scope.
The write form is stated ONCE here, and FR-001 and §0.9 restate it verbatim. A write is any of: a call to an attribute or name setenv with a literal "SPEC_KITTY_HOME" first argument — receiver-agnostic; an assignment to os.environ["SPEC_KITTY_HOME"] or environ["SPEC_KITTY_HOME"]; or .setdefault("SPEC_KITTY_HOME", …). Receiver-agnostic is not an incidental word. tests/sync/tracker/test_tracker_egress_refusal_3108.py:1165 writes mp.setenv("SPEC_KITTY_HOME", str(tmp_path / "home")) where mp comes from pytest.MonkeyPatch.context(), and the whole file is absent at §0.9's window start, so every site in it is an arrival. A receiver-qualified reading drops it and r rises toward proceed; a receiver-agnostic reading admits it and r falls. One word flips the sign at a scale where one site moves the band, so :1165 is named here as the boundary case and WP-0 may not resolve it silently. Note also that mp is bound by a with … as item, not an ast.Assign. The rationale previously attached to that fact is FALSE and is struck. It claimed a resolver modelling only Assign "drops :1165". Measured by removing withitem binding entirely: all 40 members are still found, and :1165 is still admitted. The real discriminator is RECEIVER-AGNOSTICISM on the call, which the write test mandates independently — a receiver-qualified matcher is what drops :1165. withitem-bound-receiver resolution matters only when a value expression references a with-bound name, and that shape has real-tree population 0; it is therefore registered as asserted-inert (FR-007), not as the thing that saves :1165.
| Measurement | Value |
|---|---|
.py files under tests/ | 2737 — 0 parse failures |
Byte-hit files (b"SPEC_KITTY_HOME") | 100 |
All SPEC_KITTY_HOME write sites | 191 = 188 setenv (83 files) + 3 environ-assign + 0 setdefault |
The 3 environ-assign sites | tests/audit/test_no_legacy_path_literals.py:94,112; tests/concurrency/test_ensure_runtime_concurrent.py:43 |
Effect class (tmp_path/"home") | 40 sites in 36 files — 30 fixture / 10 test-body / 0 helper, kinds taken at the KEYED def (FR-001) |
| Members (FR-001 predicate) | 40, in 36 files, across 6+ directories |
| Members under the superseded decorator-limbed predicate | 30 |
| Members under innermost-def attribution | 39 (§0.1b) |
| Pin-bearing fixtures / of those, non-members | 49 / 19 |
Explicit scope= among the 49 | 0 |
| Members a byte pre-filter would miss | 0 |
String constants valued "SPEC_KITTY_HOME" in src/ ∪ tests/ | 0 (§0.6a) |
188 vs 191 is reconciled, not glossed: 188 is the setenv-only count and is the figure the first revision published as though it were the whole population. FR-001 admits three write forms; the population is 191. setdefault has population 0 and its limb is therefore asserted-inert — like scope= under FR-007, a limb matching nothing must be known to match nothing, or it will be read as enforcement.
Named escapes, each with a measured population. Recorded rather than discovered:
| Escape | Population today | Note |
|---|---|---|
request.getfixturevalue("tmp_path") | 0 | 6 call sites name tmp_path (not 7 — the first revision's count was a text artefact) |
getfixturevalue(<variable>) — dynamic | 0 | tests/saas/test_readiness_unit.py:71 passes a Name; statically unresolvable by any AST predicate |
| Env-key indirection via a constant | 0 | §0.6a; held dormant by an executable empty-set assertion |
Delegation — a silhouette-satisfying definition calling a sibling helper that does the setenv | 0 | 3 ready targets exist (the §0.1a helpers). Recorded as a named escape; not closed by inlining, because one level of inlining invites two |
Subprocess environment dicts — env["SPEC_KITTY_HOME"] = … into a dict passed to a child process | 1 site — tests/sync/_daemon_harness.py:263 | Correctly excluded from the 191: it mutates a local dict, not the test process's environment, so it cannot pin the runtime home for the test that runs it. Recorded because it is a live instance of a shape the three-form write does not claim to cover, and an enumeration of escapes that omits a live shape is the false-completeness defect Non-Goal 5 already carried once |
Unmodelled value forms — os.path.join, %-format, .format(), + concat | 3 sites, 0 members-in-waiting | Found by the post-plan gate; the enumeration above had been claimed complete. Closed by widening the resolver (FR-001), which admits 0 new members. Each new form is a limb and carries a positive control (FR-007) |
Work packages, for the tasks phase — named here, not specified here. WP-0 the SC-000 gate · WP-a the owner · WP-b classifier + guard + pre-filter · WP-c the 40-row census · WP-d the reduced record. WP-0 gates all four; WP-c and WP-d are parallelisable once WP-b lands.
WP-0 builds the shared module WP-b imports — not a second copy. WP-0 needs the site enumerator, the binding resolver and the scope-chain silhouette; WP-b needs all of those plus the census, the set equality and the pre-filter budget. WP-0 must ship them as one importable module that WP-b then extends, never as a parallel implementation cross-checked for agreement. This repository has that exact failure on record as a live incident: tests/architectural/_sole_door_scan.py:13-27 documents Gates 4 and 5 each rolling an independent, drifting copy of the same primitives, with Gate 4's copy having "already lost Gate 1's docstring rationale … a live drift, not a hypothetical one", and names the promoted shared module as the fix. R1a follows that precedent rather than re-earning it.
0.9 THE GATE — built so R1a could stop itself; now a published measurement under a pre-committed rule
The halted Mission halted because SC-012 was unsatisfiable without an executed run taken before any source was edited. R1a's first revision had no equivalent: every criterion was defined over an artefact R1a would itself build. The gate below is the question R1a can genuinely answer before building the census, whose answer could make the Mission not worth finishing:
> Does the guard catch a new member added the way members are actually added?
The measurement, its window, and its trigger — all fixed here, before the measurement is taken.
- Window:
709a59534a1b8aac7e55a1cf6f5d2106a32c31ea→5d49d31ed6505627d98d8f95d8502c9bf6a2f5ac. The start SHA is the merge-base of the halted Mission's branch with this baseline — the commit at which its 28-member baseline was correct. Non-tunable: moving it moves §0.3 with it. R= arrivals WITHIN THE EFFECT CLASS, not allSPEC_KITTY_HOMEarrivals: the sites resolving totmp_path/"home"present at the end SHA and absent at the start SHA. This is a correction to the first revision, and the first revision's definition was unusable. WithRdrawn over all 191 sites, 151 of them pin values that are nottmp_path/"home"— 42"global", 17"kittify", 10 baretmp_path, and so on — so a window in which every new member landed in a fixture, a 100% forward catch, could still have producedr < thresholdand HALTED; and symmetrically a flurry oftmp_path/"global"test-body arrivals could have produced a spurious proceed. The gate would have measured a population it was not about.- Established by AST at both ends — run the enumerator + resolver at both SHAs and difference the site sets, keyed by the member key (C-012, as interpreted in FR-003).
git log -Sandgit diff | grepare text search over a population and are barred by C-003. - The START SHA is anchored by the INDEPENDENT instrument too. C-011's artefact anchors
discover()at the end SHA only, whileRandR_fare measured at the start SHA by the same instrument, against a threshold that decides whether the Mission proceeds — the symmetric circularity is tune untilr >= 50%. WP-0 therefore runs the checked-inclf.pyagainst agit archiveextraction at the start SHA and publishes the symmetric difference againstdiscover()'s site set there as part of the verdict. One command, an artefact already in the tree, and it closes the only circularity C-011 left open. R_f= the members ofRthe FR-001 predicate catches.r = |R_f| / |R|.
The reference rate is the standing effect-class catch rate: 40/40 = 100% by site, 36/36 = 100% by file (§0.1). The threshold is half of it: 50%.
| Measurement | Consequence | Decided by |
|---|---|---|
| **`\ | R\ | < 10`** |
| The verdict changes under ±1 (either perturbation) | INADMISSIBLE — no verdict. Same path. | Implementer, mechanically |
r = 100% | Proceed. Forward reach equals standing reach. | Implementer, mechanically |
50% ≤ r < 100% | Proceed, with the degradation published. | Implementer, mechanically |
r < 50% | HALT. WP-a…WP-d do not begin. | The operator — explicit sign-off; the implementer may not proceed on their own authority |
The ±1 rule is evaluated over the CONSEQUENCE, not the label: {proceed, proceed-degraded} versus {halt}. The first three rows are two consequences, not three — SC-000 gates on "proceed or proceed-degraded" and FR-008 makes the publication obligation unconditional, so by this specification's own construction the top two rows differ in name only.
Evaluated over labels, the gate's success state is UNREACHABLE, and this is the fifteenth defect. At r = 100% we have |R_f| = |R|, and two perturbations are admissible and both leave the proceed label: |R_f| − 1 (the clamp only skips at |R_f| = 0) gives (|R|−1)/|R| ≤ 0.9, and |R| + 1 holding |R_f| gives |R_f|/(|R_f|+1) ≤ 0.909. Both land in proceed-degraded. Under "if any admissible perturbation lands in a different band, the sample is INADMISSIBLE" that is true for every |R| ≥ 2. Enumerated over |R| ∈ [10, 40]: 318 proceed-degraded, 364 halt, 124 inadmissible, and proceed in exactly 0 states. Widening cannot rescue it — a moved window measuring 100% is inadmissible again.
And r = 100% is the likely outcome, not a knife-edge. Because R is drawn from sites resolving to tmp_path/"home", and that resolution can only root at a tmp_path in the chain, tmp_path ∈ silhouette is entailed by membership in R. So r measures exactly one thing: the fraction of new sites whose chain also declares monkeypatch. Re-derived: of all 191 pins, every tmp_path-rooted site has a keyed def — the near-miss shape (tmp_path present, monkeypatch absent) has population 0 — and receivers are monkeypatch ×187, os.environ ×3, mp ×1. Left on labels, the gate's most likely trajectory was measure → 100% → inadmissible → widen ×5 → the operator decides: the instrument built so R1a could halt itself would most probably have returned no verdict, defaulting to the judgement call the halted parent died on.
Re-enumerated over consequences: 380 go, 364 halt, 62 inadmissible, and inadmissibility is confined to the 50% boundary — which is precisely what §0.9 argues the rule is for.
VOID is NOT a band. It is a precondition on |R|, evaluated before banding, and the ±1 perturbations are taken over the consequence classes only. A perturbation that would drop |R| below the floor is outside the admissible domain and is not evaluated, exactly like the |R_f| = 0 clamp. Stated because the alternative is silently off by one: were VOID a band, |R| − 1 at |R| = 10 would always land in it, every |R| = 10 sample would be inadmissible, and the floor would really be 11 — contradicting the argument for keeping it at ten.
Why 50%, and why it cannot be re-tuned. It is half of the standing 100% — the figure this specification already owns as operator decision 1's justification and already publishes in §0.1 and in its non-goals. Below half, the coverage argument has lost more than half its value on new debt. The threshold cannot move without visibly moving §0.1, decision 1 and the non-goals together — the property that made the halted Mission's N = 5 un-tunable.
Renames are excluded from both sides, and the signature carries kind. The difference key contains the enclosing qualified name, so a rename reads as one departure plus one arrival, inflating |R| and skewing r. A departure and an arrival in the same file matching on (resolved pin value, enclosing parameter set, KIND) is a rename, and is neither an arrival nor a departure.
On the detector's actual population, re-derived: in-file collisions are 3 files at n = 2 each — tests/upgrade/migrations/test_m_0_6_7_ensure_missions.py, tests/upgrade/test_compat.py, tests/upgrade/test_m_0_12_0_documentation_mission_unit.py — and all three are same-kind, so kind discriminates none of them. What closes those three ties is the unique-mutual-best-match rule below, not kind. kind is retained anyway because it is free, correct, and forward-looking: a fixture never renames into a test body, and the exact signature (tmp_path/"home", {monkeypatch, tmp_path}) already spans 38 sites at the keyed def across both kinds, so the collision surface exists even though today's in-file ties do not use it. 17 of the 100 byte-hit files changed inside the window (3 of the 36 member files), so unrelated delete-one-add-one pairs are a live hazard, not a hypothetical one. kind and the signature's parameter set are taken at the KEYED def (FR-001), not the innermost — measuring them at the innermost def is the reading C-004 refuses, and it is what produced the withdrawn 30/9/1 split and the 37-site span.
kindis carried, and the figure that first justified it is WITHDRAWN as measured over the wrong population. The earlier justification — "20 in-file collisions, worst n = 17, 17, 13" — reproduces exactly, but it was measured over all 191 pin sites, whileRis arrivals within the effect class. The three worst are twotmp_path/"global"files and onetmp_path/"kittify"file: the detector cannot see any of them. Quoting them reads as 17-way ambiguity where the detector faces 2-way.- The pair must be a UNIQUE MUTUAL BEST MATCH, or the pairing is refused and both sites are retained. Unresolvable values (
None) never match anything. - The first revision's claim that symmetric application implies no directional bias is a NON SEQUITUR, and is struck. Excluding a fixture rename gives
(f−1)/(n−1) < f/n— toward halt. Excluding a non-fixture rename givesf/(n−1) > f/n— toward proceed. Symmetric application, opposite-signed effects; the net depends on the rename mix, which is the one thing that must therefore be published. - Publish every pair, every refused-ambiguous candidate, and every unpaired departure and arrival.
±1 stability, in BOTH axes, over CONSEQUENCE classes, with the clamp stated. Recompute the consequence class — {proceed, proceed-degraded} or {halt} — under |R_f| ± 1 holding |R| fixed (one site reclassified) and under |R| ± 1 holding |R_f| fixed (one false rename pair, or one missed arrival). The second axis is what an earlier revision could not see: a false pair moves |R|, not |R_f|. Clamp: perturbations are taken only where they remain in the admissible domain — |R_f| − 1 is skipped at |R_f| = 0, |R_f| + 1 at |R_f| = |R|, |R| − 1 at |R| = |R_f|, and |R| − 1 is also skipped when it would fall below the floor, since VOID is not a band. If any admissible perturbation lands in a different consequence class, the sample is INADMISSIBLE.
|R| ≥ 10 is a VISIBILITY floor, not statistical adequacy, and this specification says so plainly. The first revision anchored it to the 10 non-fixture effect sites; that anchor belonged to a different population and is withdrawn. No floor can fix the boundary case: at a threshold of 1/2, every even n puts the threshold exactly on an attainable value of r and every odd n puts it strictly inside a grid cell, so a single reclassification is decisive at the boundary for every n. Only checking the actual measurement closes it, which is what the ±1 rule does. Ten is kept because a sample below it is not worth a reviewer's time.
THE FLOOR IS NOT MET AT THE STATED WINDOW, AND THE SPECIFICATION'S OWN PROSE SAID SO. Measured post-tasks: over 709a59534 → 5d49d31ed the class is 37 members at the start SHA and 40 at the end under the current predicate, with 0 departures, so |R| = 3 against a floor of 10. The window is VOID — no verdict. The three arrivals are tests/cli/commands/test_sync_doctor_tracker_egress_3108.py:124 and tests/sync/tracker/test_tracker_egress_refusal_3108.py:194,1165 — all of #3108, exactly as §0.3 describes.
This was derivable from §0.3 two sections earlier and six passes plus three adversarial gates did not derive it. §0.3 publishes the growth as 28 → 30 across precisely this window; 30 − 28 = 2 against a floor of 10 needed no measurement. §0.9 then stated outright that "|R| and r were deliberately NOT measured in this pass" while setting a floor five times the published growth. The transferable lesson, and it is the reason this is recorded as a named finding: a PRECONDITION must be checked against the published data that determines it. Same family as the struck 26.06% and FR-001's 9 against 1, but on a precondition rather than a figure — and the first in this Mission that would have cost a whole implementation session rather than a correction.
Lowering the floor to fit the window is REFUSED — by the operator, explicitly. It is exactly the tune-after-measurement §0.9 exists to prevent, and it would move a threshold after seeing the number it judges.
The widening schedule is PRE-COMMITTED, because "widen until it passes" is a garden of forking paths. A stopping rule defined by the property being measured is not a stopping rule. Therefore: walk the start SHA backwards along upstream/main's FIRST-PARENT history, one first-parent commit at a time, until BOTH |R| ≥ 10 and ±1 consequence-class stability hold. There is NO attempt cap. The step unit is first-parent COMMITS, and the reason is binding: upstream/main is rebase-merged, so there are no merge commits for "one merge at a time" to step over — an implementer would have stepped one commit at a time regardless, and the old 5-attempt cap therefore bounded the search to five commits inside a 132-commit, four-day window, which cannot reach the floor. Measured: |R| ≥ 10 first appears between 300 and 600 first-parent commits back.
THE WALK IS UNBOUNDED BUT NOT UNTERMINATED, AND THIS PASS NEARLY SHIPPED IT THAT WAY. Deleting the 5-attempt cap also deleted the escalation attached to it ("after which the OPERATOR decides"), leaving an implementer whose two conditions are never jointly satisfiable with no defined exit — a search that runs to the root of history. Restored, and it is not a cap on the walk:
- Reachability is evidenced, not assumed:
|R| >= 10is first met between ~300 and ~600 first-parent commits back, so the floor is reachable. What is not evidenced is that ±1 consequence-class stability holds at any such window. - Defined exit: if the walk reaches the root of
upstream/main's first-parent history without both conditions holding jointly, it stops and the OPERATOR decides — the same escalation the cap carried, now triggered by exhaustion rather than by an arbitrary count. - This is not a re-tuning surface: the exit is reached only by exhausting the history, so it cannot be used to stop early at a convenient band.
THE STOPPING RULE ITSELF DOES NOT CHANGE, AND THAT IS WHAT KEEPS AN UNBOUNDED WALK HONEST. It selects on |R| and ±1 consequence-class stability ONLY — never on the value of r. No candidate window is accepted or rejected because of the band it produces. That independence is the entire forking-path defence, and it survives the cap's removal untouched.
THE MEASUREMENT HAS LEAKED, and pretending otherwise would be dishonest. While verifying the VOID finding, an independent lens measured r = 100% at candidate windows roughly 300, 600 and 2000 first-parent commits back (with |R| = 9, 33, 34). The gate can no longer be run blind by anyone who has read this. What is PROTECTED is the stopping rule's independence from r, above — a leaked value cannot steer a rule that does not read it. What is LOST is the blindness of the analyst, which was a real part of the instrument's value. And the consequence must be stated rather than left for a reader to infer: the first window that meets the floor is around ~600 commits back, where the leak reports |R| = 33 and r = 100% — so on the published evidence THE GATE'S VERDICT IS ALREADY KNOWN TO BE proceed. The instrument built so that R1a could halt itself will, on everything now published, not halt it. That does not make running it pointless — the operands, the rename pairs and the stability check are still the record SC-000 requires, and the measurement can still surprise us — but R1a must stop describing the gate as its own stopping mechanism. It is now a published measurement with a pre-committed rule, and its power to halt this Mission was spent when the value leaked. Both facts are stated because only one of them is recoverable. Every attempted window is published with its (start SHA, |R|, |R_f|, band), including the discarded ones — beginning with the VOID result at 709a59534 (|R| = 3). The window WILL move, so §0.3's obligation is now concrete and falls on WP-0: the record must name, in so many words, whether §0.3's 28 → 30 figure is RE-DERIVED at the moved SHA or EXPLICITLY SUPERSEDED. Not "addressed" — one of those two words. Leaving it unstated would let a moved window silently contradict §0.9's own non-tunability claim, and §0.3 is the section whose published growth already determined this window's fate once.
SC-000 gates on the VERDICT, not merely on publication (see SC-000), and the verdict ships as a machine-readable verdict: field.
The band is MECHANISED, not eyeballed. band(|R_f|, |R|) and stability(|R_f|, |R|) ship as functions, and a test loads the published verdict artefact and asserts band(published) == published.verdict. The banding function, the clamp and the ±1 rule are fully specified above, and this specification already provides the 806-state oracle to test them against. Delegating the Mission's only stopping instrument to "checked by a reviewer" is the one thing this programme's history says will not hold: nothing else in the criterion set can catch verdict: proceed written above r = 0.6.
The verdict assertion gates EVERY package, not one. It ships as a collected test that reds until the artefact reads proceed or proceed-degraded — not as "WP-a's first task", which gates only whichever package happens to be sequenced first and which any concurrent package can skip.
~~|R| and r were deliberately NOT measured in this pass.~~ SUPERSEDED — dated record of the pre-measurement state, retained because it explains why the design was sound and the arithmetic was not. As authored (spec revisions 1-2, through 96b2b0910), this paragraph read: "|R| and r were deliberately NOT measured in this pass. The window was validated on member counts under the superseded predicate only (28 at the start SHA, 30 here); the enumerator was never run at the start SHA under the current predicate." Every clause of that is now false. The enumerator has been run at the start SHA under the current predicate: 37 members there, 40 here, |R| = 3, VOID — measured post-tasks and recorded at the head of this section. The principle stands unchanged and is why the paragraph is kept rather than deleted: fixing a trigger after seeing its measurement is the exact failure this instrument exists to prevent, which is why the floor was not lowered to fit the window (above). What failed was never the discipline of not measuring; it was that the floor was set without checking §0.3's published growth, a subtraction that needed no measurement at all. This is the THIRD stale claim to survive inside the section that documents stale claims — after FR-001's 9 against 1 and §0.3's 7 to 9 — and it survived by sitting 22 lines below its own correction, so a reader going top to bottom met the refutation first and the superseded claim last. The correction must be the last word on the page, not the first.
User Scenarios & Testing (mandatory)
User Story 1 — The class cannot change without someone deciding (Priority: P1)
A contributor adds a definition that pins SPEC_KITTY_HOME to tmp_path / "home" and can see (tmp_path, monkeypatch) — as a fixture, a test function, a helper, or a nested closure. Today nothing notices. After R1a the guard reds and the contributor must request the canonical owner or record a decision.
Why this priority: This is the Mission.
Independent Test: Add a 41st member under tests/ under a name used nowhere else; the guard reds. Remove a census row while its definition remains; the guard reds. Add a 41st member together with a census row for it; the guard still reds. All red-first.
Acceptance Scenarios:
1. Given the census-frozen tree, When the guard runs, Then it greens — attesting the class is frozen, asserting nothing about whether any member deserves to be there. 2. Given a 41st member anywhere under tests/, When the guard runs, Then it reds naming its composite_key. 3. Given a stale census row, When the guard runs, Then it reds on the row, not only on the missing member. 4. Given a 41st member plus a new census row absorbing it, When the guard runs, Then it still reds — additions are structurally impossible, not merely discouraged. 5. Given a 41st member that also requests the owner and restates the pin, When the guard runs, Then it reds. The remedy is a decision, never a narrowed silhouette.
User Story 2 — A legitimate 41st member has a green path (Priority: P1)
An author needs a private home for a new module. They request the canonical owner by parameter and write nothing.
Why this priority: Without it the census only blocks.
Independent Test: A probe module that requests the owner and asserts its SPEC_KITTY_HOME resolves to its own tmp_path/"home" with the directory present at test-body entry.
Acceptance Scenarios:
1. Given a module requesting the owner and writing no pin, When it runs, Then os.environ["SPEC_KITTY_HOME"] is its own tmp_path/"home" and the directory exists. (Bound by SC-012 — behavioural, not static.) 2. Given the owner added to tests/conftest.py, When the patch is inspected, Then it is added strictly after line 298 and the module's ordered list of definition names, WITH THE NEWLY-ADDED OWNER REMOVED, is unchanged — one known addition, then exact equality. Not "ignore differences": exactly one named definition is removed, the owner FR-005 mandates, and everything else must match including the ordering on both sides of the anchor. Asserted as the ordered list, never as a scalar definition index, which is satisfiable while the ordering it stands for has changed. (Bound by SC-005 and SC-010.) 3. Given a module that keeps its own setenv and also requests the owner, When it runs, Then the module's own value is observed in the test body — the owner never wins. (Bound by SC-012.)
User Story 3 — The guard is fast enough to keep, without seeing less (Priority: P1)
Why this priority: A guard over budget is deleted or narrowed, and narrowing is the failure this design exists to prevent.
Independent Test: Time the guard warm ×3 against budget (a). Prove coverage preservation executably, and prove the key-indirection population empty.
Acceptance Scenarios:
1. Given the guard, When timed warm ×3, Then every run is inside budget (a) while enumerating every .py under tests/. 2. Given both classifiers, When run, Then their full outputs — member sets and every member's home_partition — have empty symmetric difference (conditional on OD-002 — see SC-002). 3. Given the repository, When the key-indirection assertion runs, Then the set of assignment-bound string constants valued "SPEC_KITTY_HOME" across src/ ∪ tests/ is empty — 0 against 229 literal occurrences in 98 files, both published (SC-002b). (Corrected: this said "string constants" unqualified, which measures 229, not 0.)
User Story 4 — The record says what R1a proved and what it did not (Priority: P2)
Why this priority: The parent Mission was bitten five times by an inherited figure promoted to a gate; the first revision of this spec was bitten a sixth time by 26.06%.
Independent Test: A reader can recover the three labelled reach figures, the gate's outputs, and the census's entitles-nothing property from the record alone.
Acceptance Scenarios:
1. Given the record, When a reader looks for the guard's reach, Then they find three separately labelled figures, never one merged ratio. 2. Given the record, When a reader looks for a growth rate, Then they find §0.3's explicit refusal and the authored-vs-committed distinction.
Edge Cases
- A pin written in a nested closure. Caught — the silhouette is evaluated over the scope chain (§0.1b), and
:1165is the live instance. - A pin written via
pytest.MonkeyPatch.context(). Caught — the write test is receiver-agnostic and applies to the call, not the receiver (§0.8). - A fixture acquiring
tmp_pathviagetfixturevalue. Not caught. Population 0; recorded escape. The dynamic form is statically unresolvable by any predicate. - A definition delegating the
setenvto a sibling helper. Not caught. Population 0; recorded escape. - A member is deleted. The row goes stale and the guard reds; the repair is a tombstone (FR-004), never re-adding the fixture.
- A definition legitimately restates the pin. A member by effect; reds until adjudicated. The repair is never the silhouette (C-004).
- A member gains a parameter, is reordered, or loses
autouse. Still a member — superset, order-insensitive, decorator-agnostic. - A member is renamed or moved.
composite_keysurvives blank-line and comment drift; a genuine rename reds in both directions, which is correct. - A
SyntaxErrorin a test module. Must not be silently skipped —except SyntaxError: continuebuys budget headroom by narrowing the walk with nothing firing. Bound by SC-013. - A future member declaring
(monkeypatch, tmp_path)reversed. pytest resolves by name; the set-based silhouette is order-insensitive.
Requirements (mandatory)
Functional Requirements
| FR-001 | The classifier — effect-keyed, decorator-free, scope-chain silhouette | As a maintainer, I want the class discovered by an AST pass that resolves single-assignment local bindings (str/Path/os.fspath unwrapping, f-string joining, and — added post-gate — os.path.join, %-format, .format() and + concatenation), over def/async def at any nesting depth with no decorator requirement, whose enclosing scope chain's union of parameter sets contains {tmp_path, monkeypatch} (leading self stripped), keyed by the outermost chain member satisfying it, and whose body writes SPEC_KITTY_HOME — in the receiver-agnostic three-form sense stated once in §0.8 — to a value resolving to tmp_path/"home". The walk resolves TWO environment variables. SPEC_KITTY_HOME decides membership; HOME, enumerated with the same three-form receiver-agnostic write test and the same value resolution, decides home_partition (FR-003) over each member's own scope chain. The byte pre-filter stays sound and the argument is stated rather than assumed (C-003 form (i)): a member's scope chain lies entirely within one file, so every HOME write that can affect a member's partition is in the same file as that member's SPEC_KITTY_HOME write — which is by construction a byte-hit file. No widening of the pre-filter is required, and none is permitted as a substitute for this argument. And the argument is now MEASURED, with the amount the pre-filter hides published — because a soundness claim that never states its denominator cannot be audited. Of the 50 files containing a HOME write, 21 are invisible to the b"SPEC_KITTY_HOME" pre-filter (33 sites); zero of those 21 holds a member, and zero member files are non-hit. Re-derived end to end: the truly unfiltered HOME pass finds 85 sites in 50 files against the pre-filtered 52 in 29, and the two produce identical partitions for all 40 members — 0 disagreements. The locality argument holds by construction and by measurement. The HOME enumeration is a new limb and ships a positive control (FR-007); any sub-form matching nothing is labelled asserted-inert with its measured population. The decorator limb is dropped because it was a shape key excluding 10 effect sites against the silhouette's 0 (§0.1a), and re-keying shape→effect is what the parent Mission's correction was for. (Corrected in the plan phase: this row published 9 against 1, which are the innermost-attribution figures C-004 refuses — the fifth appearance of the withdrawn 30/9/1 split, in the row that binds the classifier. Under the scope-chain rule this row itself binds, §0.1a measures the decorator limb at 10 and §0.1b measures the silhouette limb at none at all. Independently confirmed by the C-011 evidence instrument: 40 members under the current predicate, 30 under the superseded decorator-limbed one, and 40/40 effect sites carrying the silhouette at their keyed def.) Re-derived: 40 members in 36 files. The resolver widening admits ZERO new members, verified: the population stays 40 and no published figure moves. All three UNRESOLVED sites are non-members for reasons independent of resolution — tests/audit/test_no_legacy_path_literals.py:112 — the second of _capture_nudge's two writes (:94,112) — restores a saved env value in a finally (os.environ[...] = old_env, not a tmp_path pin), and the enclosing def fails the silhouette; tests/concurrency/test_ensure_runtime_concurrent.py:43 is a top-level multiprocessing worker unpacking (home_path, asset_dir) = args and fails the silhouette (no tmp_path, no monkeypatch); tests/sync/test_routing.py:440 satisfies the silhouette but resolves statically to tmp_path / f"{case}-home" / ".spec-kitty" — a different value than tmp_path/"home". So the UNRESOLVED bucket is 3 sites but 0 members-in-waiting: the escape was PROSPECTIVE, not live. Each widened form is a new sub-form and therefore carries a positive control (FR-007); any that matches nothing is labelled asserted-inert with its measured population. "Sub-form" and "limb" are NOT interchangeable, and the inert registry is scoped to SUB-FORMS. A limb is a whole clause of the predicate (the silhouette limb, the write limb, the HOME enumeration); a sub-form is one shape within a limb (setdefault, bare-name setenv, %-format). A registry entry names a SUB-FORM. Any entry whose real-tree population is non-zero is a REGISTRY DEFECT, not a matcher defect — it means the registry described something that does fire. Stated because the HOME enumeration was registered as an inert limb, when measured it is the most heavily populated limb in the classifier: 85 HOME write sites in 50 files, and 13 of the 40 members re-pin HOME. It is what produces home_partition; it is the opposite of inert. | High | Open |
| FR-003 | The frozen census — 40 rows, no reason, constants hoisted, effect axis carried | As a reviewer, I want a sorted, checked-in census of 40 rows, one per measured member, the canonical owner not among them, and no reason column. frozen_at_sha and owed_to are FILE-HEADER SCALARS, not per-row columns. 40 rows times two columns held exactly two distinct values, which is §0.4's own argument against a reason column relocated one level over: a column whose every cell is identical is a header scalar wearing row clothing, and it invites per-row divergence that nothing needs. Row columns are therefore: the MemberKey 3-tuple, the non-authoritative lineno, kind, and home_partition. home_partition: Literal["A","B1","B2","other"] is the EFFECT axis, and its rule is imported, cited by path, and cross-checked: kitty-specs/isolated-home-pin-guard-r1a-01KZNMA3/research/m4_ablation_evidence/ (VERDICT.md, TABLES.md, RESIDUALS.md, verbatim and sha256-pinned). A = does not re-pin HOME; B1 = re-pins HOME to tmp_path/"home"; B2 = re-pins HOME to tmp_path/"user-home"; other = re-pins it to anything else — other has measured population 0 over the 40 and is therefore ASSERTED-INERT under FR-007's own rule, the same treatment as setdefault and scope=; a limb matching nothing must be known to match nothing or it reads as enforcement. (Recorded rather than discovered: this arm was added and left unlabelled in the very edit that cited the rule.) The partition keys on a SECOND environment variable, which is why no arrangement of the existing limbs could produce it — the plan phase first reported the rule as undefined when it was defined precisely, on a branch this one cannot see, over a variable this scanner did not enumerate (FR-001 now enumerates both). Re-measured over the current 40: A = 27, B1 = 11, B2 = 2 (0 other). Cross-checked against M4's independent per-member labels: intersection 28 — measured, not assumed — with 28 agreements and 0 disagreements, and the delta decomposes exactly (the 10 members from the limb drop are all test-body and all A; the 2 from #3108 are both fixtures and both B1). That makes this a second external anchor authored by a different actor, on the C-011 pattern. home_partition IS RECORDED AS REGENERABLE, and this is a condition of keeping it. It appears in no key, no hash and no equality: the baseline hashes MemberKey triples and the ratchet compares those, so the column cannot make the guard red or green. Combined with FR-004(2) — the census is generated by a single documented command, and the baseline hash is invariant under any change that leaves the key set alone — the column can be regenerated later without touching the ratchet. Stated so that a slip on the partition can never hold the ratchet hostage on the package that gates the rest. 17/9/2 of 28 remains labelled as the superseded decorator-limbed frame and is NOT R1a's figure. kind is retained but is a SHAPE key, kept only for the rename signature; recorded as such so that shipping a census keyed on shape does not contradict §0.1a's shape-to-effect correction, which is the whole argument for this Mission's predicate. owed_to matches ^#[0-9]+$ and nothing else — the first revision's second disjunct ("an issuer's own id pattern") permits ^SK-[0-9]+-[a-z-]+$, and SK-12-also-pins-home is a reason column in kebab case. R1a is barred from creating issues and #3121 exists, so one pattern suffices. Rationale that would otherwise be a per-row reason belongs in the file header, following tests/architectural/census/verdict_seam_IC01.yaml. | High | Open | | FR-004 | Set equality against census ∪ E, with direction MECHANISED | As a maintainer, I want the guard to assert the discovered class is set-equal to census ∪ E, where E is the enumerated exemption set — the owner (FR-005) and the single retained-pin probe (FR-011) — typed as tuple[Exempt, Exempt], a FIXED-ARITY tuple, so that a third exemption is a mypy error and not a one-line literal; and additionally pinned by a checked-in hash of its sorted entry set held OUTSIDE the guard module, following the same test_allowlist_shrink_only mechanism this row already cites. Per-entry typing was not enough: typing each entry bounds that entry's arity and says nothing about |E|, so a frozenset[tuple[str, str]] would accept a third entry with no type error while §0.4 asserted in prose that E was fixed-size by type. The spec asserted the stronger property and required the weaker one; both are now required. And to enforce direction mechanically rather than in prose: ~~(a) every row's frozen_at_sha equals the freeze SHA~~ — struck, because FR-003 hoists frozen_at_sha to a header scalar and a per-row equality over one value has no per-row divergence to catch. The load it carried is INHERITED BY LIMB (b), and naming that is the point of this sentence: (a)'s real function was to make an added row detectable by its SHA, and what detects an added row now is (b)'s shrink-only key-set hash, which reds because the key set changed — plus the exact accounting of discovered == census u E. (a) was a REDUNDANT defence, not a vacuous one. The plan phase first justified this strike with "nothing for it to catch", which is true only because (b) exists; left at that, a future pass could remove (b) on identical reasoning and leave nothing. Removing (b) is therefore barred while (a) is struck.; (b) a shrink-only baseline pinned outside the mutable artefact — a checked-in hash of the sorted freeze-time key set, plus an explicit tombstone list, so removals require a tombstone and additions are not expressible. The first revision left direction as prose after mechanising owed_to for exactly the reason prose was known insufficient, and the guard could not distinguish a legitimate removal from an illegitimate addition — both leave set equality green, and a 41st member was absorbable by a one-line append. The repository ships this shape at tests/architectural/test_charter_path_literal_authority.py::test_allowlist_shrink_only, which pins charter_path_literal_baseline: 49 and asserts len(keys) <= baseline. Its placement is NOT a precedent to copy, and R1a states its own — because "pinned outside the artefact" is a description, not a rule, and a rule is what stops the pin relocating the problem instead of closing it. Binding: (1) Paths, named — not described. The census is tests/architectural/census/spec_kitty_home_pin_R1a.yaml, following the tests/architectural/census/verdict_seam_IC0N.yaml fragments. The census baseline hash and E's hash are tests/architectural/spec_kitty_home_pin_baseline.yaml — a different file from the census, so the pin is not editable in the same hunk as its subject. (2) Generated, never hand-edited, with the generator named. Both files are emitted by tests/architectural/_home_pin_scan.py — the shared WP-0/WP-b module of §0.8, following _sole_door_scan.py — through a single documented regeneration command, and both carry a header saying so. A token-style field described as "never typed by hand" with no generator to produce it is a comment, not a mechanism. (3) The co-edit rule, and it differs between the two artefacts for a reason. For E: any delta to E or to its hash reds, unconditionally — E never legitimately changes, in R1a or in R1b, so there is no legitimate co-edit to permit. For the census: a blanket "touching both reds" is refused, because it would forbid the burn-down that is R1b's entire job. Instead the guard recomputes the baseline hash from the census and compares, and every delta must be accounted for by a tombstone — so a co-edit that removes a row and re-pins the hash still reds unless a tombstone explains it, while a legitimate adjudication passes. Git-state inspection is not used: it does not survive rebase or squash, and content comparison does. Recorded as an unsolved instance in the repo's own idiom (TG-4): load_baseline(ALLOWLIST_PATH) reads the scalar from the same file as the allowlist it bounds (charter_path_literal_allowlist.yaml:25), there is no generator in the tree, and no co-edit guard exists. What actually holds that gate together is not placement but exact accounting — len(load_allowlist(...)) == len(live_sites()) — which makes an added allowlist row require an added live site, which the gate itself reds. R1a inherits that insight: FR-004's discovered == census ∪ E is exact accounting, and it is the primary ratchet; the hash pin and the tombstone list exist to catch the one thing exact accounting cannot see, a removal with no adjudication behind it. | High | Open |
| FR-007 | Inert limbs are known to be inert | As a maintainer, I want the classifier to treat an absent scope= as function scope and reject only an explicit non-function value. The limb is STRUCTURALLY UNREACHABLE, which is a stronger argument than the empirical one it replaces: under a predicate with no decorator requirement scope= is undefined for the 10 test-body members, and for the 30 fixture members a non-function-scoped fixture cannot request function-scoped tmp_path, so no member can carry a rejectable value. (The previous justification — "0 of 49 pin-bearing fixtures carries one" — is an empirical count taken in the pre-limb-drop frame, the same frame that produced the withdrawn 26%; it is retained only as corroboration. The limb is decorator-derived and survived the decorator limb's removal.) And, on the same principle, setdefault's population is 0 and its limb is recorded as asserted-inert. A limb matching nothing must be known to match nothing or it reads as enforcement. THE INERT SUB-FORM REGISTRY, enumerated with measured populations — over the 191 SPEC_KITTY_HOME sites and the 85 HOME sites:
| ID | Title | User Story | Priority | Status |
|---|---|---|---|---|
| FR-002 | The byte pre-filter, and its now-stated premise | As a maintainer, I want ast.parse run only on files whose raw bytes contain b"SPEC_KITTY_HOME" while every .py under tests/ is still enumerated — a hard requirement, not an optimisation (§0.6). Its over-selection argument depends on a premise the first revision left unstated: that the env key is a literal at the call site. Measured, 73 of 1245 setenv/delenv calls already use a non-literal key, so the idiom exists at scale. The premise is therefore stated here and proved executably: the population of ASSIGNMENT-BOUND string constants valued "SPEC_KITTY_HOME" across src/ ∪ tests/ is 0 — against 229 literal occurrences in 98 files, and both figures are published — shipped as an empty-set assertion with a positive control under C-003 form (ii). (Corrected post-tasks: this row asserted the population of ALL such string constants was 0, which measures 229. SC-002b was restated and this row was left standing — the same defect, on the requirement the criterion implements.) The day that set is non-empty the assertion reds and the pre-filter must widen before the guard is trusted. | High | Open |
| FR-005 | The canonical owner — one, and structurally one | As a maintainer, I want exactly one owner fixture in tests/conftest.py, non-autouse, function-scoped, returning None (AST-asserted), providing the pin and the directory. The tuple[str, str] belongs to the owner's ENTRY IN E (Exempt), not to the fixture — the two were conflated here and in SC-005, and an implementer told to AST-assert both finds that whichever limb they pick the other is unsatisfiable, whose cheapest escape is to weaken one (C-002's named failure mode). return None is KEPT, and the reason is load-bearing for SC-012's non-circularity: a fixture that returns its own path invites assert os.environ["SPEC_KITTY_HOME"] == owner, which compares the environment against the fixture's own report and would pass for an owner that sets nothing. Returning None forces the probe to compute str(tmp_path / "home") itself. It is E whose fixed arity by type makes a second entry inexpressible. The first revision moved the owner out of the census to protect C-007 and thereby relocated the exemption from an audited, pattern-checked, reviewer-tested data file into an unaudited literal in the guard's own source, where C-007's "carries no exception" does not reach — and no criterion asserted the owner set had exactly one entry, so a second name would have greened the guard with every written criterion satisfied. The addition must be strictly after line 298. | High | Open |
| FR-006 | The owner establishes by env var only, and never wins | As a maintainer, I want the owner to establish the home only through monkeypatch.setenv — never monkeypatch.setattr(Path, "home", …) or any process-global patch — and never to override a definition keeping its own pin, upholding the decision at tests/conftest.py:272–286. Demonstrated behaviourally (SC-012), not merely asserted from source. | High | Open |
Every row carries a canonical id, and the id column is what makes this table an EXTERNAL operand. A row whose sub-forms are packed into one cell, or whose meaning is anaphoric on the row above, cannot be parsed into ids without the parser inventing them — and a parser written to emit the author's own set is the circularity the fourth operand exists to close, relocated one level out. Hence one sub-form per row, no four-way cells, no back-references:
| id | Inert sub-form | Population | Limb |
|---|---|---|---|
SKH-SETDEFAULT | setdefault("SPEC_KITTY_HOME", …) | 0 | write |
SKH-BARE-SETENV | bare-name setenv (unqualified call) with a literal "SPEC_KITTY_HOME" | 0 | write |
SKH-VAL-OSPATHJOIN | os.path.join in a SPEC_KITTY_HOME value | 0 | value resolution |
SKH-VAL-PERCENT | %-format in a SPEC_KITTY_HOME value | 0 | value resolution |
SKH-VAL-FORMAT | .format() in a SPEC_KITTY_HOME value | 0 | value resolution |
SKH-VAL-CONCAT | + concatenation in a SPEC_KITTY_HOME value | 0 | value resolution |
HOME-SETDEFAULT | setdefault("HOME", …) | 0 | write |
HOME-VAL-OSPATHJOIN | os.path.join in a HOME value | 0 | value resolution |
HOME-VAL-PERCENT | %-format in a HOME value | 0 | value resolution |
HOME-VAL-FORMAT | .format() in a HOME value | 0 | value resolution |
HOME-VAL-CONCAT | + concatenation in a HOME value | 0 | value resolution |
SCOPE-EXPLICIT | explicit scope= among the 49 pin-bearing fixtures | 0 | fixture shape (structurally unreachable) |
PARTITION-OTHER | home_partition == "other" | 0 | partition |
WITHITEM-VALUE-REF | a withitem-bound receiver referenced by a value expression | 0 | value resolution |
Fourteen rows, fourteen ids. The only id in the positive-control set that is not in this table is SC-002b's assignment-bound-constant assertion, which is a criterion in its own right rather than a sub-form of the classifier — so the delta between this table and the control set is exactly one, and it is named.
Two of these were missing from the first enumeration — bare-name setenv and withitem-bound-receiver resolution — which is why a completeness mechanism shipped incomplete. An enumeration that claims to list every inert sub-form and does not is the same false-completeness defect as Non-Goal 5's "each with population 0".
AND EVERY POPULATION-0 ASSERTION SHIPS A POSITIVE CONTROL. C-003 form (ii) otherwise makes the assertion its own proof: an over-narrow matcher returns set(), greens forever, and still greens the day someone writes RUNTIME_HOME_ENV = "SPEC_KITTY_HOME". That is DIR-041's canonical pass-for-the-wrong-reason, and this row diagnoses the disease and then shipped it. A positive control runs the same matcher over a synthetic module containing the shape and asserts it returns exactly that hit. Required for SC-002b, for this row's setdefault limb, and for every resolver limb added by the FR-001 widening. | High | Open |
| FR-008 | The reduced record, with every figure labelled and the gate's outputs unconditional | As the issue's reader, I want #3121 updated with: the three separately labelled reach figures (§0.1) and an explicit retraction of the struck 26.06%; the census-is-not-a-manifest distinction and its reviewer test; §0.3's provenance correction; the statement that R1a adjudicates nothing; and — unconditionally, not only in the degraded band — r, ` | R | , | R_f | , both window SHAs, every attempted window, and the machine-readable band verdict. The first revision made the degraded band's publication obligation invisible by enumerating four items that did not include r`, while SC-009 asserted "FR-008's four items". | Medium | Open |
|---|---|---|---|---|---|---|---|---|
| FR-009 | The classifier is root-parameterised, and synthetic trees are MATERIALISED, never checked in | As an implementer, I want the classifier's walk root to be a parameter, not a module constant, so that every guard behaviour can be exercised against a synthetic tree without editing a real test module. The synthetic trees are materialised into tmp_path at test time from source strings; NO synthetic tree is checked in under tests/. SC-004's word is "permanent test", not "permanent tree". A checked-in synthetic directory would sit inside the guard's own walk: enumerate_py_files is bound to every .py under the root and never narrowed (C-003), the real-tree pass roots at tests/, and these trees exist precisely to contain 41st members — so they land in discovered and discovered == census ∪ E can never green. Every cheap repair is barred: directory exclusion by C-003/NFR-001, a census row by C-007, a third E entry by fixed arity. (Note the shape: FR-011 found ONE instance of "the guard's own fixtures are members" and closed it with a fixed-arity exemption; a checked-in tree would reintroduce a whole DIRECTORY of the same instance — the signature defect operating on this Mission's own newest rule. "Not collected" is a pytest property and says nothing about the classifier's walk.) A counterpart guard asserts discover(Path("tests")) finds no member under any test-owned fixture root, so this cannot be reintroduced. This also resolves SC-013 independently: a checked-in deliberate SyntaxError reds ruff check with invalid-syntax, which is not a rule code, so pyproject.toml's per-file-ignores cannot reach it and no exclude covers _-prefixed directories — and the repository has zero precedent for .py.txt or broken* fixtures. Without this, SC-004 has no C-001-compatible demonstration path — making a census row stale requires deleting or mutating a real member, which C-001 and C-006 forbid — and the cheapest repair available to an implementer who discovers that is the subset-only ratchet §0.4 spends a page refusing. | High | Open | ||||
| FR-010 | The owner parameter counts as tmp_path | As a maintainer, I want a resolver limb treating a parameter naming the canonical owner as equivalent to tmp_path for both the silhouette and value resolution, because R1a WIDENS an existing escape — (corrected twice, and the direction is the point. First post-gate: this row said the escape was "manufactured" by R1a; it is not. Second post-WP01, by construction: the shape is live as a SHAPE but population 0 as a MEMBER. The single instance is _capture_nudge — the parameter runtime_home at tests/audit/test_no_legacy_path_literals.py:82, whose enclosing def writes at :94 and :112 (:82 is the parameter declaration, not a write; this row previously cited :82,94, conflating the two). It is refused on both limbs: the chain union {argv, module_name, runtime_home} has no monkeypatch, so the silhouette fails; and the value resolves to the bare parameter at :94 and is unresolvable at :112, never tmp_path/"home". Re-derived: it is the only def in tests/ declaring runtime_home, and canonical_home has zero defs anywhere in the repository — that name came from this row's own prose and matches nothing.) — a member can be written (monkeypatch, <a fixture yielding the home>) with no tmp_path declared, failing the predicate on two limbs at once. Because FR-005 binds the owner to return None, the canonical owner is NOT itself such a fixture, and this row's limb is therefore stated over the SHAPE, not over the owner: the resolver treats any parameter naming a fixture whose declared purpose is to yield a private home as equivalent to tmp_path, and the canonical owner is named in that set for the silhouette limb even though it yields nothing. THIS LIMB IS MEASURED INERT TODAY — member-level population 0 — and FR-007's rule applies to it as the SIXTEENTH inert sub-form. "A limb matching nothing must be known to match nothing, or it will be read as enforcement": FR-007 applies that to fifteen sub-forms and this specification had not applied it to FR-010's. The limb exists prospectively, to catch the shape once WP03 adds the owner — and the obligation that makes it fire is WP03/T012, which binds the owner contract's declared fixture name into OWNER_PARAM_NAMES. Until then it matches nothing, and that is recorded rather than discovered. And the rationale is now stated as a HYPOTHESIS, not a finding: it has been weakened twice by measurement, both times in the same direction — manufactures → widens an existing escape at population 1 → population 0 at member level, with the one backing site failing the predicate on two limbs. A rationale that keeps shrinking under measurement is one to hold as a hypothesis until the owner exists to test it. SC-006 transition 6's witness is rewritten accordingly: it declares (monkeypatch, <a synthetic-tree fixture that YIELDS a home root>) — modelled on the live population-1 runtime_home shape — not (monkeypatch, canonical_home), because a witness requesting a None-returning fixture could not restate the pin from its value. This contradiction survived planning because the one seam two parallel packages must agree on had no contract; contracts/canonical-home-owner.md is now WP-a's first deliverable. Without this limb the guard goes blind in proportion to R1b's adoption — the more the owner is adopted, the more members become invisible — and SC-006's transition 6 becomes satisfiable only with a hand-picked witness that gratuitously re-declares tmp_path. | High | Open | ||||
| FR-011 | Probe modules, and their exemption | As an implementer, I want R1a to ship probe modules exercising the owner behaviourally (SC-011, SC-012), with exactly one of them exempted. Most probes need no exemption: a probe that requests the owner and writes no pin of its own is not a member, and SC-006's transition witnesses live in the synthetic tree of FR-009, outside the guard's real walk. Exactly one real-tree probe is a member — SC-012's second limb, which must keep its own setenv to prove the owner never wins — and it is therefore declared in E beside the owner, typed as a fixed singleton, and counted by SC-005. Recorded rather than discovered: the C-010 limb-(i)/(ii) probes are constructible only because FR-010's escape exists, so closing that escape turns the retained-pin probe into a class member needing a carve-out C-007 forbids — hence one enumerated exemption rather than an open allowance. | High | Open |
Non-Functional Requirements
| ID | Title | Requirement | Category | Priority | Status |
|---|---|---|---|---|---|
| NFR-001 | Guard runtime — budget (a), gating | Under 6 seconds warm, three consecutive runs, enumerating every .py under tests/. Derivation in §0.6: slowest measured pre-filtered figure 1.86 s × 3 (1.86× observed spread, doubled for an unmeasured runner). The budget may be RAISED with runner evidence (OD-003); the walk may never be narrowed. | Performance | High | Open |
| NFR-002 | The two proofs — budget (b), separable, non-gating | The pre-filter's coverage preservation ships as an executable test, never a recorded claim, at 90 s warm under OD-002's both-passes form, marked separable and not bound by NFR-001; under the file-set-containment form the budget is 2 s and the weaker proof and cheaper budget move together. Separately and unconditionally, the FR-002 key-indirection empty-set assertion ships at negligible cost and is not substitutable by either. No precedent for an executable pre-filter proof exists in this repository (§0.6); R1a sets it. Budget (b) may be RAISED with runner evidence on the same terms as NFR-001, and the walk may never be narrowed and the form may never be weakened to (b) to fit it. (Added in the plan phase: NFR-001 carried a raise-with-evidence clause and this row did not, while both budgets face the same unmeasured runner — an asymmetry whose only cheap resolution under load would have been the weakening this row exists to forbid.) | Performance | High | Open |
| NFR-003 | Parallel-run correctness | Identical verdicts under -n0 and -n auto --dist loadfile, both demonstrated. | Reliability | High | Open |
| NFR-004 | Static gates clean | ruff check and mypy --strict clean on every touched file, no new # noqa, # type: ignore or per-file ignore. ruff format is not run. | Maintainability | High | Open |
| NFR-005 | No interference with the existing home owner | The owner must not alter, shadow, or reorder relative to _isolated_worker_home (tests/conftest.py:253–298). Demonstrated by an executable check — an AST assertion that the module's ordered list of definition names, WITH THE NEWLY-ADDED OWNER REMOVED, is unchanged — one known addition, then exact equality. Not "ignore differences": exactly one named definition is removed, the owner FR-005 mandates, and everything else must match including the ordering on both sides of the anchor. (not a scalar definition index: a scalar can hold while the order it summarises has changed, and it is the ordering conftest resolution depends on). The carve-out is what makes the check satisfiable at all, since FR-005 requires this very WP to add the owner. Plus the strictly-after-298 placement of FR-005. "Demonstrated, not assumed" needs a concrete check, and the first revision supplied none. | Correctness | High | Open |
| NFR-006 | Pinned-version re-measurement | Measured on pytest>=9.0.3,<9.1 (pyproject.toml:102), including the [INHERITED] runtime baseid derivation of FR-005. The external pytest-9.0.3 venv is reused; no bare uv run/uv sync, and no venv inside this tree. | Correctness | High | Open |
Constraints
| C-012 | Use the repository's canonical key and idiom, not an improvised one | Keys use composite_key from specify_cli.contracts.anchoring:220, re-exported at tests/architectural/_ratchet_keys.py. ROW IDENTITY IS THE PATH-QUALIFIED 3-TUPLE MemberKey = (rel_path, enclosing_qualname, normalized_token_line). This APPLIES this constraint; it does not amend it — the parenthetical (enclosing_qualname, normalized_token_line) is a gloss on the primitive's return type, and all three authorities this row names already use the 3-tuple for row identity: tests/architectural/_sole_door_scan.py:87 ("rel_path/qualname/token form the authoritative composite key", with lineno explicitly non-authoritative), _ratchet_keys.py's "Key shape — reuse, not fork" docstring (CompositeKey = tuple[str, str, str]), and tests/architectural/surface_resolution_audit/audit.py:90-91 ("`(rel_path, enclosing_qualname, token) -- the frozen-comparand row identity", CompositeKey = tuple[str, str, str]). Measured: the bare 2-tuple yields 19 distinct values over the 40 members, so a census keyed on it cannot hold 40 rows and the ratchet goes blind to the 29 members that sit in a collision class. The justification this row previously carried is measurably FALSE and is struck: ~~"an improvised (file, qualified_name) cannot express two pin sites inside one definition"~~ — the 3-tuple carries the token line, so it expresses them fine; the claim confused (file, qualname) with the 3-tuple. composite_key is content-addressed and survives blank-line/comment drift. THE KEY IS FORMED AT THE WRITE SITE, never at the definition line. composite_key_from_file is a pure function of (file, lineno), and the two readings are not merely different: measured, they produce 40-row censuses with ZERO overlap (19 write-site keys against 21 def-line keys), and under the def-line reading two pin sites in one definition collapse to one key by construction — the very thing this row exists to prevent. C-004 governs membership ATTRIBUTION, not site IDENTITY, and the two must not be conflated — they are separately measurable and they disagree. Binding: (0) KEYS ARE FORMED FROM THE ALREADY-PARSED TREE discover() HOLDS, never by re-reading the file. anchoring.py:192-195 — reached through composite_key_from_file — does try: ast.parse(source) except SyntaxError: return "<module>", so a broken file silently yields (rel, "<module>", "") instead of raising, and SC-013's guarantee would hold only on the parse_module path. Zero SyntaxErrors across all 2737 files today, so the hole is narrow — but SC-013 exists precisely so that a future one reds. composite_key_from_file is used only for the independent live recomputation, where the file is known-parseable because discover()` has just parsed it. Incidental win: not re-reading saves a measured 0.19 s per pass.
| ID | Title | Constraint | Category | Priority | Status |
|---|---|---|---|---|---|
| C-001 | No test module is edited | R1a edits no existing test module. The only edits to existing files are the owner's addition to tests/conftest.py (strictly after 298). No adoption, no deletion, no annotation repair, no parity run, no discriminating red. New files — the guard, its instrument, the census, the probes, the synthetic trees — are not edits to test modules. | Technical | High | Open |
| C-002 | No counted definition-of-done | Not gated on a collected-test count or a count of definitions. 40 and 36 are content, published as key sets (C-011), never thresholds. The only number that is a threshold is SC-000's, and it is fixed before its measurement. | Technical | High | Open |
| C-003 | Populations by binding-resolving AST, never text search — two carve-out forms | Every population, count and set comparison is produced by an AST pass resolving local bindings. Text is admissible only as a matcher provably over-selecting w.r.t. the AST population it precedes: (i) a pre-filter ahead of an AST pass; (ii) an over-selecting matcher returning the empty set. Narrowing by directory or filename qualifies under neither. Every carve-out matcher ships its own executable proof — form (i) via NFR-002, form (ii) via FR-002's empty-set assertion. | Technical | High | Open |
| C-004 | Narrowing the silhouette is not an acceptable repair | When a definition that legitimately restates the pin reds the guard, the remedy is an adjudication. Narrowing to arity-exact, re-adding a decorator limb, keying on a name, or attributing to the innermost def instead of the scope chain are all refused. | Technical | High | Open |
| C-005 | The _isolated_worker_home decision binds unchanged | tests/conftest.py:272–286 records env-var-only establishment with no Path.home patch, after the setattr form silently broke ~16 tests/sync cases and was reverted. Re-derived: the file is byte-identical to the halted branch. The parent's C‑009, C‑010, FR‑015 and NFR‑007 bind unchanged — written with U+2011 per the convention in §0.5, so the scraper cannot mistake them for R1a's own C-009 and C-010. | Technical | High | Open |
| C-006 | Blast radius | Limited to: tests/conftest.py (the owner); WP-0's enumerator/resolver module, which exists before the guard module; the guard module and its tests; the census artefact and its baseline; the probe modules and synthetic trees; tests/_arch_shard_map.py (OPTIONAL balance-control pin only — see below); the C-011 evidence artefact at kitty-specs/isolated-home-pin-guard-r1a-01KZNMA3/research/spec_kitty_home_pin_evidence/; the imported M4 ablation evidence at .../research/m4_ablation_evidence/ (extracted with git show from spike/isolated-home-3121 — no merge, rebase or cherry-pick; it supplies home_partition's rule and its independent 28-member labels); the halt-path ADR under docs/adr/3.x/; and this Mission's record. No file under src/ changes. No existing test module changes. The plan phase first justified this entry as "not optional" and that premise was FALSE — withdrawn and re-justified honestly. tests/_arch_shard_map.py:419 sets default_fallback=True, and tests/_shard_registry.py:181 hash-buckets any unregistered under-root file into a deterministic shard, so a brand-new file is auto-covered by construction. The module's own docstring says so twice: "no manual table edit is required just to keep main green" and the explicit tables "remain the authoritative balance control, not a keep-green obligation". The historical incident the plan quoted as current behaviour ends "...until the default_fallback hash-bucket auto-cover picked them up (FR-011/#2671)". So editing this file is OPTIONAL load-balance pinning, not a keep-green obligation — a legitimate reason to touch it, just not the one first given. It stays in the blast radius on that basis. It is not a test module (it collects nothing), so C-001 is untouched either way. | Technical | High | Open |
| C-007 | The census is monotonically non-increasing, and entitles nothing | A row may be removed when R1b adjudicates its definition, accompanied by a tombstone. A row may not be added. A row entitles its definition to nothing. This constraint carries no exception and none may be added — the exemption set E — the owner (FR-005) and the single retained-pin probe (FR-011) — is kept out of the census and typed so neither entry can grow, precisely so no carve-out is needed. The reviewer test is stated over anything that makes a definition acceptable to the guard (§0.4), not over census rows alone. | Technical | High | Open |
| C-008 | `\ | P\ | = 5` is NOT an input, and R1b cannot inherit it | Measured over 28 members at a merge-base predating #3108, with two current members never ablated — and the class is now 40 under a different predicate. R1b must re-run WP01 over the current class. | Technical |
| C-009 | Pre-existing reds are not this Mission's to fix | Baseline-red classification applies. Per DIR-013, pre-existing failures encountered are reported before being treated as accepted baseline. | Process | Medium | Open |
| C-010 | PR #3285 is a coordination dependency of R1b | 35 files, all under kitty-specs/; intersection with the member files empty. Not an R1a blocker. | Process | Medium | Open |
| C-011 | The measured key sets are published as evidence, and the count is never the anchor | The member composite_key set, the site split, and the value-bucket table ship as a checked-in evidence artefact produced by the spec-phase classifier, which is preserved and identified by path. The path, satisfied in the plan phase: kitty-specs/isolated-home-pin-guard-r1a-01KZNMA3/research/spec_kitty_home_pin_evidence/ — instrument clf.py, producer step3.py, evidence members.json (40 entries), all checked in verbatim and sha256-pinned in that directory's README.md. Authored by the post-spec gate's independent third lens from the specification's predicate text alone, never compared against the spec author's classifier before publication, and verified during planning to reproduce members.json byte-for-byte on a tests/ tree that is byte-identical to the baseline. members.json carries (path, qual, line, sites); the key encoding is derived at test time by the repository's own composite_key_from_file, so the anchor — which sites are members — never passes through the instrument under test. SC-001 and SC-003 assert set equality against those published keys, not against a number. Without this the only anchor is the count, and it is circular: a reviewer re-running the shipped instrument compares its output against the census the same instrument produced, so the cheapest WP-b path is tune the predicate until it yields 40. | Technical | High | Open |
(0a) The type composition is stated because mypy --strict will not accept the bare comparison. composite_key_from_file returns tuple[str, str] while MemberKey is tuple[str, str, str]; wherever this specification asserts equality between them the composition is *MemberKey = (relpath_posix, composite_key_from_file(path, lineno)). Comparing the two directly is a type error**, not merely a wrong assertion.
(1) The MemberKey qualname component is the WRITE-SITE enclosing qualname — the repository primitive's value, which is the innermost dotted qualname (anchoring.py:173,211 — the def and its docstring at :173, the min(candidates) narrowest-span selection at :211). It is correct for identity precisely because the dotted form contains the whole chain: at tests/sync/tracker/test_tracker_egress_refusal_3108.py:1165 it is test_bind_counter_wrapper_changes_no_outcome_committed_red._run_once. (2) The KEYED def determines membership, kind, and the rename signature's parameter set, and is NOT in the key. (3) Measured, and this is why (1) is normative rather than advisory: keying at the innermost gives symmetric difference 0 against the C-011 anchor; keying at the outermost gives 2, both at :1165. An implementer told to key at the keyed def ships a discover() that fails SC-001's set equality two packages downstream. (4) C-004 is therefore MECHANISED BY THE kind DISTRIBUTION, not by the key set — measured, kind at the keyed def is 30 fixture / 10 test-body / 0 helper and at the innermost is 30 / 9 / 1, which is exactly the withdrawn split. That makes C-004 falsifiable where it is actually observable, instead of resting on a key that must be innermost for an unrelated reason. (5) Both discriminators currently rest on a SINGLE real-tree site — :1165 — which is held in the class by an unused monkeypatch parameter that ruff's relaxed ARG for tests/ will not defend (§0.1b). The permanent guard therefore ships a SYNTHETIC outermost-versus-innermost witness, so C-004's mechanisation does not depend on one real row surviving a routine cleanup. The key is formed at the boundary, from the Member record — never inside the write-site finder or the silhouette keyer. The census artefact follows tests/architectural/census/, and the guard follows the repo's sole door idiom (tests/architectural/_sole_door_scan.py, five existing gates). The 3-tuple is not injective either — though at MEMBER level the collision population is 0, and the earlier "LIVE" framing overstated it. Measured over all 191 write sites the guard walks: 190 distinct 3-tuples, one collision class of two — tests/paths/test_runtime_root_spec_kitty_home.py:91,93, a definition carrying the full (tmp_path, monkeypatch) silhouette with two setenv sites that collapse to one key because normalized_token_line strips "one" and "two". Neither site is a member — their values are tmp_path/"one" and tmp_path/"two" — so the member-level collision population is 0*. The pair is the right illustration of the hazard and the wrong word for its status: it is one string literal away from being two members with one key, in a file named test_runtime_root_spec_kitty_home.py. Hazard real, population 0. Therefore discover() ships an import-time exactly-one assertion over its own output, stated at MEMBER level: if two MEMBERS ever produce the same MemberKey, the guard reds at import rather than silently deduplicating. The literal mechanism this clause first prescribed — _ratchet_keys.py's assert_descriptor_unique_within_qualname applied per member — is FALSE and is struck. Measured against the real tree with the D-1 rule (occurrence=None), it raises on 11 of the 40 members. Cause: code_tokens_by_line strips string literals, so at tests/cli/commands/test_sync_commands.py the fixture _isolated_home's three consecutive monkeypatch.setenv calls — SPEC_KITTY_HOME, HOME, LOCALAPPDATA — collapse to one normalized token line, and only one of the three is a member. Source-scoped descriptor uniqueness is therefore not the property this clause wants: it fires on sites the guard does not own. ContentDescriptor is retained as the diagnostic vehicle for reporting a collision, not as the uniqueness predicate. Record the shape: this clause was added to close BLOCKER-1's key collision and named a mechanism scoped to the wrong population — the same family as §0.9's floor set against the wrong window, and as the withdrawn "20 in-file collisions, worst n = 17" measured over all 191 sites rather than over the detector's own. A guard prescribed for the wrong population is this Mission's most-repeated defect. Member and every census row also carry lineno, explicitly NON-AUTHORITATIVE, per the sole-door precedent at _sole_door_scan.py:87-89 — normalized_token_line takes only 4 distinct values across all 40 members, so without it a maintainer working 40 debt rows has no way to reach the site. CLAUDE.md's canonical-sources rule is explicit that improvising here propagates drift. Tension recorded, because "follow the idiom" would otherwise import a defect: the sole door does except SyntaxError: at _sole_door_scan.py:477, and so do test_commit_target_kind_guard.py:169 and _ratchet_keys.py:169, which R1a imports. SC-013 forbids that construct. Follow the idiom's key shape and artefact layout*, never its syntax-error handling. | Technical | High | Open |
| C-014 | The owner's three precedence-safety limbs | Inherited from the halted parent's C‑010 and given an R1a ID because SC-011 binds them normatively and the bare reference resolved two ways. (i) ScopeMismatch makes scope inversion loud at setup — exercisable. (ii) A non-autouse owner is never instantiated in a module that does not request it — exercisable. (iii) Every retained and adopting definition stays function-scoped — genuinely vacuous in R1a, which has zero adopters, and recorded as the Mission's residual (§0.5). SC-011 binds (i) and (ii) behaviourally and states that both are properties of pytest's fixture machinery rather than of the owner's body. | Technical | High | Open |
|---|---|---|---|---|---|
| C-013 | Nothing is merged, and no issue is created | No gh pr merge, no git merge, no un-drafting, no gh issue create. Explicit-path git add; never git add -A, git reset --hard, git checkout -- ., git clean or git stash. Long commands bounded with timeout; a timeout is a datum, never silently retried. | Process | High | Open |
Key Entities
- Behaviour class — 40 members in 36 files on
5d49d31ed, defined by effect per FR-001: no decorator limb, superset silhouette over the enclosing scope chain, receiver-agnostic three-form write, value resolving totmp_path/"home". Comprises every site (40/40) and every file (36/36) of the effect class. - Effect class — the 40
SPEC_KITTY_HOMEsites resolving totmp_path/"home", across 36 files: 30 fixture / 10 test-body / 0 helper. There is no helper member. The kind of:1165istest-body, because FR-001 keys it to the enclosing test; calling it a helper reported the innermost def, which C-004 names as a refused reading. The populationRis drawn from this, not from all 191 pins. - Census — the frozen, sorted, checked-in set of 40
MemberKeyrows. Header scalarsfrozen_at_shaandowed_to(^#[0-9]+$); row columnsMemberKey, non-authoritativelineno,kind(shape, rename signature only),home_partition(effect — rule imported and cited atresearch/m4_ablation_evidence/, re-measured A=27 / B1=11 / B2=2 over the 40, cross-checked 28/28 against M4). Noreason. Direction mechanised by a baseline hash plus tombstones (FR-004).census == ∅is R1b complete, with every removal accounted for. - Canonical owner — one fixture, typed
tuple[str, str], non-autouse, function-scoped, returningNone, added strictly aftertests/conftest.py:298. Itself a class member under FR-001, which is whydiscovered == census ∪ Eis the only correct form — excluding it by predicate would require narrowing the silhouette, which C-004 forbids. R1b must not "simplify" this. - Named escapes —
getfixturevalue(static and dynamic), env-key indirection, delegation, and unmodelled value forms (os.path.join,%-format,.format(),+concat). The first four measure population 0; the fifth measured 3 UNRESOLVED sites but 0 members-in-waiting and is closed by FR-001's widening. The enumeration is not claimed complete — it was, and a fifth escape was found by the post-plan gate (Non-Goal 5). Each entry is recorded with its measurement (§0.8). - The guard's non-domain — the other value buckets:
tmp_path/"global"(42 sites, 3 files),"kittify"(17/4), baretmp_path(10/10), and the rest. Out of scope by definition, not by blindness.
Success Criteria (mandatory)
Each criterion states what it cannot see.
Cannot see: how contributors will behave; arrivals that never became sites; a rename that also changed the pin value, which the detector will correctly read as a departure plus an arrival.
- SC-000 — THE GATE. WP-a…WP-d may not begin until §0.9's measurement is published and its verdict is
proceedorproceed-degraded. Publication alone does not gate — a published HALT is still a halt. The verdict comprises: (i) the site sets at both SHAs — the operands, so the difference is recomputable by a reviewer without re-running anything; (ii) the excluded rename pairs with their matching keys, every refused-ambiguous candidate, and every unpaired departure and arrival; (iii)R,|R|,|R_f|,r; (iv) both SHAs, every attempted window, and the exact invocation; (v) a machine-readableverdict:field; and (vi) the published label —proceed,proceed-degradedorhalt— which must equal the band recomputed from the publishedr,|R|and|R_f|, checked by a reviewer from operands FR-008 already publishes unconditionally. Without (vi) the two go-labels are distinguishable only in prose: consequence-class stability (§0.9) deliberately makes them one class for admissibility, so nothing else would catch a verdict recorded asproceedatr = 0.6. WP-0's enumerator ships checked in and regenerable by a single command. Bands, thresholds, ±1 stability in both axes over CONSEQUENCE classes ({proceed, proceed-degraded}vs{halt}) with the stated clamp, and the pre-committed widening schedule are fixed in §0.9 before the measurement. Stability tested over labels rather than consequences makesproceedunreachable in every state (§0.9), which would turn the halting instrument into a guaranteed no-verdict.
Cannot see: whether any member deserves to be a member. R1b's, entirely.
- SC-001: The classifier's discovered member set is set-equal to the published
MemberKeyset derived from C-011's evidence artefact — the path-qualified 3-tuple of C-012, formed at each write site inmembers.json'ssites. Asserted as a set, never as the number 40. (Post-gate correction, and it was the worst consequence of the key defect: under the bare 2-tuple the published set has 19 elements, so a classifier finding exactly ONE member per collision class discovers 19 members and produces the SAME 19-element set — SC-001 green. C-011 could not distinguish a 40-member classifier from a 19-member one, which is strictly worse than the tune-until-40 failure C-011 was written to prevent. The 3-tuple restores 40 distinct keys and with them the criterion's power.)
Cannot see: key-indirection members. Neither form can (§0.6a); SC-002b covers that.
- SC-002: The pre-filtered and unfiltered classifiers agree, in the form OD-002 selects. Under form (a) their full
discover()outputs — member sets and every member'shome_partition— have empty symmetric difference, executed under NFR-002's 90 s budget, AND the same test asserts the two passes PARSED DIFFERENT FILE SETS — prefiltered< 2737, unfiltered== 2737. The comparison covers BOTH variables because the pre-filter now serves both (FR-001). Extending the pre-filter's scope toHOMEwithout extending its proof would leave the second variable's soundness resting on prose — which is exactly what §0.6 and TG-3 condemn the repository's two existing pre-filters for. (The plan phase first claimed this measured — "pre-filtered and unfiltered agree on all 40 partitions, 0 disagreements" — and that measurement was VACUOUS: both arms ran through the C-011 instrument, whoseclassify()hardcodesif b"SPEC_KITTY_HOME" not in b: continue, so the "unfiltered" arm was silently pre-filtered and the symmetric difference was empty by construction. That is precisely the defect this criterion's differential exists to catch —prefilteras a parameter the body ignores — committed in the same pass that wrote the criterion. Re-measured honestly against a patched copy: the truly unfilteredHOMEpass finds 85 sites in 50 files against the pre-filtered 52 in 29, and the two still give identical partitions for all 40 members, 0 real disagreements. The conclusion survived; the proof of it did not, and only the second measurement is evidence.) Without that differential the criterion is satisfiable withprefilteras a parameter the body ignores: the symmetric difference is then empty by construction at about 1 s, and because 90 s is a ceiling and not a floor, a suspiciously fast pass reads as good news. Under form (b){files containing a member} ⊆ {files hitting the needle}, and the criterion must then carry the clause "this proves strictly less" naming what it does not prove. Written conditionally because a criterion hard-coded to (a) while OD-002 stays open is unsatisfiable if the architect picks (b) and would be reworded at implementation time — which C-002 names as the failure mode.
Cannot see: a constant assembled at runtime rather than written as a literal.
- SC-002b: The set of ASSIGNMENT-BOUND string constants valued
"SPEC_KITTY_HOME"acrosssrc/∪tests/is empty — targets ofAssign/AnnAssign/ class-attribute binding — asserted executably, and both figures are published: 0 assignment-bound against 229 literalast.Constantoccurrences in 98 files. The 229 is what makes the 0 meaningful. (Corrected post-gate: the criterion previously asserted that ALL string constants valued"SPEC_KITTY_HOME"are empty. Measured, there are 229 — every literalsetenvfirst argument, including all 40 members — so the predicate was falsified by its own subject matter and the only available green was a narrowing at implementation time, which C-002 names as the failure mode. The population that is genuinely 0 is the assignment-bound one, which is what §0.6a actually argues.) This is the pre-filter's stated premise and the only check that can falsify it. It ships with a POSITIVE CONTROL (see FR-007).
Cannot see: whether owed_to names the right ticket — only that it is a well-formed reference and not prose.
- SC-003: The census is set-equal to the published
MemberKeyset minusE, sorted, no row carrying areason, and the owner not among the rows. The header carries exactly onefrozen_at_shaequal to the freeze SHA and oneowed_tomatching^#[0-9]+$; no row carries either. Every row carrieskind,home_partitionand a non-authoritativelineno.
Cannot see: whether the row was stale for a good reason.
- SC-004: The guard reds on a stale census row, demonstrated as a permanent test over a synthetic tree (FR-009), not as a red-first demonstration that leaves no artefact. A subset-only ratchet passes this tree and fails this criterion; that is its purpose.
Cannot see: whether the owner works. SC-012 covers that (SC-011 does not — it is pure shape, see its own cannot-see line) — and their absence in the first revision meant an owner that silently did nothing satisfied every criterion, in a Mission whose thesis is that shape is not effect.
- SC-005: The owner exists in
tests/conftest.py, is non-autouse, function-scoped, returnsNone, contains no process-global patch (all AST-verified), and the patch adds it strictly after line 298. Thetuple[str, str]is asserted ofE's entry, not of the fixture (FR-005): asserting both of the fixture is unsatisfiable, since it returnsNone. A test asserts that adding a third entry toEis a TYPE ERROR — the disjunction "or reds the guard" is deliberately dropped, because a test that reds is precisely the artefact a contributor adding a third exemption edits in the same change, with the first two entries as the template. AndE's membership is asserted against its checked-in hash held outside the guard module (FR-004), not against the guard's own literal: assertingEis "exactly as declared" while both the declaration and the assertion live in the guard module is circular in exactly the way C-011 describes for the census.
Cannot see: whether a member's presence is justified.
- SC-006 — the EIGHT transitions, red-first: (1) reds on an empty census; (2) greens on the census-frozen tree, attesting freezing, not convergence; (3) reds on a 41st member anywhere under
tests/; (4) reds on a 41st under a name used nowhere else; (5) reds when a row is removed while its definition remains; (6) reds on a 41st that also requests the owner and restates the pin — with the witness declared(monkeypatch, <a synthetic-tree fixture yielding a home root>), notmp_path— modelled on the live population-1runtime_homeshape, not theNone-returning canonical owner (FR-010) — named in the criterion, because a witness that gratuitously re-declarestmp_pathtests nothing FR-010 is for. (7) reds on a 41st member accompanied by a new census row — the absorption path, which none of the first six touched. (8) reds when one of a COLLIDING PAIR is removed — a synthetic tree carrying two members in different files whose barecomposite_keyis byte-identical (same fixture name, same normalized token line), one row removed, guard reds. Without this transition the key-type interpretation of C-012 is unpinned: it can be reverted, or a future key change can re-introduce the collapse, with all seven other transitions still green. Nobody hand-authoring a synthetic tree produces this shape by accident, which is why it is named rather than left to the author. This is the single cheapest test that would have caught the 19-key collapse, and the only one that keeps it caught.
Cannot see: CI-runner behaviour (OD-003).
- SC-007: The guard completes inside NFR-001's budget on three warm runs, and the set of files it enumerated is equal to the set of every
.pyundertests/— asserted as set equality, with the count reported, not asserted. *The expected set MUST be computed by aPath(root).rglob(".py") written inline in the test, never obtained from the module under test: the seam exposes exactly one enumerator, so the natural assertion is a self-comparison, and a narrowed walk — NFR-001's exact defeat — passes a self-comparison. The anti-drift rule forbids the guard modules from callingast.parseor subclassingNodeVisitor; it does not** forbidrglob, and this criterion requires it.2737is stale the moment anyone adds a file, and a criterion that asserts it trains reviewers to re-baseline on noise. Identical verdicts under-n0and-n auto --dist loadfile.
- SC-008:
ruff checkandmypy --strictclean on every touched file, no suppression added.
- SC-009: #3121 carries FR-008's items, including all three labelled reach figures, the retraction of 26.06%, and the gate's outputs unconditionally.
Cannot see: ordering effects arising outside this file.
- SC-010: in
tests/conftest.py, the module's ordered list of definition names, WITH THE NEWLY-ADDED OWNER REMOVED, is unchanged — one known addition, then exact equality. Not "ignore differences": exactly one named definition is removed, the owner FR-005 mandates, and everything else must match including the ordering on both sides of the anchor. AST-asserted (NFR-005), and never as a scalar definition index — a scalar is satisfiable while the ordering it summarises has changed (two definitions swapping either side of the anchor leaves the index intact), and conftest fixture resolution depends on the ordering, not on the number. (Corrected: this criterion briefly required the list to be unchanged outright, which no implementation can satisfy —tests/conftest.pyis in WP03'sowned_filesand FR-005 requires WP03 to ADD the owner to it, so the list necessarily gains an entry. The fix for a criterion that could not FAIL produced one that could not PASS — same axis, opposite end — and the cheapest repair available to an implementer facing an unsatisfiable criterion is to weaken it back toward the scalar form, which is what the change was for. Recorded because the failure mode is symmetric to the one being fixed and the repair path leads straight back to it.) Inserting the owner above line 253 would shift it down and change conftest fixture ordering while the old line range shows no modification — so the diff-shaped criterion alone is satisfiable with its risk untouched.
Cannot see: limb (iii), which needs adopters — and whether the owner does anything at all. Limbs (i) and (ii) are properties of pytest's fixture machinery, not of the owner's body: an owner whose body is return None with no setenv and no mkdir satisfies both. In a Mission whose thesis is that shape is not effect, SC-011 is pure shape and SC-012 carries the entire behavioural load alone. A reviewer seeing SC-011 green must not read it as evidence the owner works.
- SC-011: C-014 limbs (i) and (ii) are exercised: a probe requesting the owner at a higher scope raises
ScopeMismatch; a module not naming the owner never instantiates it. C-014 limb (iii) remains vacuous and is recorded as R1a's residual (§0.5).
Cannot see: behaviour under real adoption, which R1a has none of.
- SC-012: Behavioural, binding US2 AS-1 and AS-3. A probe requesting the owner observes
os.environ["SPEC_KITTY_HOME"] == str(its own tmp_path/"home")with the directory present at test-body entry; and a probe that keeps its ownsetenvand requests the owner observes its own value, failing if the owner's is seen. *LIMB 2 AS FIRST WRITTEN IS VACUOUS, verified by construction: FR-005 forces the owner to pinstr(tmp_path/"home"), any class member pinsstr(tmp_path/"home"), andtmp_pathis function-scoped and shared — so within one test the two strings are IDENTICAL and "fails if the owner's value is seen" cannot fail. The repair is a second, NON-MEMBER probe pinningtmp_path/"probe-home"— a value the owner can never produce — which is the only assertion in the pair that can be falsified. The probe's own pin MUST live in a module-local FIXTURE, not in the test body, and the criterion states the required declaration order. Fixture setup completes before the test body, so a body-levelsetenvoverrides the owner unconditionally and proves nothing about the fixture-versus-fixture* precedence decision attests/conftest.py:272-286that this limb exists to demonstrate. R1a spends one of only two irrevocable, hash-pinnedEslots on this probe; an assertion that need not bite is not worth that price.
Cannot see: a file unreadable for reasons other than syntax.
- SC-013: A test module containing a deliberate
SyntaxErrorreds the guard, red-first.except SyntaxError: continuenarrows the walk with nothing firing and buys budget headroom, which is NFR-001's defeat wearing an exception handler.
Explicit Non-Goals
1. No existing test module is edited. No adoption, no deletion, no annotation repair. 2. No adjudication. The census records debt; it decides nothing. 3. No parity run, no discriminating red, no ablation. 4. R2's three deletion candidates are NOT in R1a — tests/delivery/test_purge_all_body_uploads_3030.py, tests/delivery/test_purge_all_events_3030.py, tests/specify_cli/identity/test_identity_value_faults_3030.py::TestThePolicyGateAnswersInsteadOfCrashing._isolated_home. They belong to R1b's WP08 and appear here only as ordinary census rows. 5. The guard covers its class completely and other buckets not at all. 40/40 sites and 36/36 files of the effect class; 0 of the 42 tmp_path/"global" sites, the 17 "kittify", the 10 bare tmp_path, and the rest — out of scope by definition, not by blindness. The first revision's "structurally invisible to any form of this guard" is struck as false (§0.1a). What remains genuinely uncoverable is the named escapes. (Corrected post-gate: this previously read "the four named escapes, each with population 0", which is a false completeness claim — it asserted both that the enumeration was complete and that every entry measured 0. A fifth escape existed unnamed: value forms the resolver did not model (os.path.join, %-format, .format(), + concat), whose UNRESOLVED bucket is 3 sites. FR-001's widening closes it, and the widening admits 0 new members, so the population figures stand — but "each with population 0" was carrying an unearned guarantee about a set whose membership had not been established.) 6. No growth-rate claim (§0.3). 7. |P| = 5 is not used, inherited, or cited (C-008). 8. Nothing is merged and no issue is created (C-013). 9. SC-000 measures arrivals, not people — where new effect-class sites landed in one bounded, published window. Not a forecast. 10. The two corrections to the halted Mission's artefacts are NOT in R1a. scripts/mutants/ablate_home_pin_3121.py and evidence/ablation/VERDICT.md exist only on spike/isolated-home-3121 — verified absent from this baseline. C-013 bars merging, so R1a has no in-scope path to either file, including any escape hatch, because the artefact to mark is also absent. Both move to R1b, which will have the branch in scope. See "Corrections Deferred".
Open Decisions for the Plan Phase
- OD-001 — WHICH tracker reference
owed_tonames. The shape is closed (^#[0-9]+$, FR-003). R1a cannot create issues, and #3121 exists. The plan chooses between #3121 and a ticket the operator opens before WP-c lands. A row whoseowed_toresolves to nothing is a row with no creditor, which is indistinguishable from a permanent one.
(a) Both-passes set identity — strong over the classifier's own reach; ≈ 30.5 s recorded / ≈ 18.4 s here; NFR-002's 90 s budget. (b) File-set containment — {files containing a member} ⊆ {files hitting the needle}; 0.056 s here / 0.22 s recorded; never runs the classifier twice. (b) is strictly weaker: it proves no currently-known member's file is dropped, and assumes the member set it is meant to help establish. Neither form closes §0.6a's key-indirection hole, which is why SC-002b is separate and unconditional and why (a) buys nothing over (b) with respect to it. A hybrid — (b) in the always-on pole, (a) marked — is available. SC-002 is written conditionally on this decision.
- OD-002 — how C-003 form (i) is discharged, and at what strength.
- OD-003 — the CI-runner factor in NFR-001. The 3× is a measured 1.86× spread doubled for an unmeasured runner. Measure the guard on the actual runner. The budget may be raised; the walk may not be narrowed.
- ~~OD-004 — where the guard lives.~~ SETTLED BY MEASUREMENT, and it is not what the first revision assumed.
tests/architectural/runs inarch-adversarial(.github/workflows/ci-quality.yml:2026) — a standalone, always-on, 3-shard job gated byif: always(), with no path filter and noneeds:edge to the fast lane, so it runs on 100% of pushes and PRs. It was deliberately extracted fromintegration-tests-core-misc, which now carries an explicit "do NOT re-add it here" note (:1872-1875). The guard therefore lands intests/architectural/with a shard assignment intests/_arch_shard_map.py, and NFR-001's budget protects something every contributor feels on every PR. CLAUDE.md:256 is stale on this point — filed as TG-2.
The "100% of pushes and PRs" conclusion rests on THREE limbs, and the plan phase found only one of them stated here. Landing in tests/architectural/ is necessary and not sufficient:
1. A selected marker must be declared on the module. The pole selects -m '<arch_shard_N> and not windows_ci and (git_repo or integration or architectural) and not timing', so the module must declare at least one of architectural, integration or git_repo — architectural is the semantically correct one. No conftest hook auto-applies it; declaration is by convention, and a module declaring none of the three is collected and then deselected on every PR, leaving NFR-001's budget protecting something no contributor ever feels. (Two corrections from the post-plan gate. First, the count "161 of 164" is withdrawn: it is a ratio between two populations that neither yields — measured by AST, test_.py recursive gives 162 of 165 and top-level 161 of 164, and the figure is sensitive to whether a per-function @pytest.mark.architectural counts alongside a module-level pytestmark, which is exactly how test_resume_non_reemission_guard.py differs. It was also originally a text-search figure, which C-003 bars. The qualitative claim carries the argument and the count is dropped. Second, the requirement is the disjunction, not the architectural marker alone — test_resume_non_reemission_guard.py is selected via git_repo while carrying architectural on one function only.) 2. A shard marker must be present — which default_fallback=True supplies automatically. tests/_arch_shard_map.py:419 opts the arch group into the fallback and tests/_shard_registry.py:181 hash-buckets any unregistered under-root file, so test_arch_shard_marker_completeness.py's total-partition invariant holds without a table edit. An explicit row is optional load-balance pinning (C-006). (This limb previously read "a shard row must exist … or the invariant reds", which was the false premise C-006 already withdrew — corrected here too, on the checking surface.) 3. E's seed must not trip the positional-anchor ratchet — the FOURTH hard ratchet, and it had zero mentions anywhere in this Mission. tests/architectural/test_ratchet_positional_anchor_ban.py::test_no_int_line_sink_in_architectural_python_seeds walks every tests/architectural//.py and flags an int literal reaching the second positional argument of composite_key_from_file(path, N) / code_tokens_by_line(...), including a module-level allow-list seed constant that embeds a positional path:NNN anchor (its #2564 clause). The cheapest form of E — a module-level tuple of (rel, lineno, …) rows — trips it. Compliant form, binding: E's entries are content-addressed MemberKey 3-tuples, and any lineno used in a recomputation is obtained from discover() at RUNTIME, never written as a literal in the seed. 4. The file must not become a zero-gate orphan. tests/architectural/test_gate_coverage.py::test_no_new_orphan_surfaces is a hard ratchet whose committed baseline is empty by design since the #2296 drain, so every new test file must be positively selected by at least one CI job. tests/_arch_shard_map.py records a prior mission landing a file unregistered and reddening both that ratchet and the shard-completeness guard.
5. The job's if: excludes two PR labels — pr:deferred and pr:skip-ci — on both arch-adversarial and timing-nfr-serial. 6. A docs-only narrowing branch exists (.github/workflows/ci-quality.yml:2110-2136) that swaps selection to -m '<arch_shard_N> and docs_scoped and not windows_ci' and tolerates exit code 5. Effect on R1a is benign — a docs-only PR cannot mutate tests/ — but "100% of pushes and PRs" is stated as measured fact, and limbs 5 and 6 were missed by the same pass that found 1 through 3 (and limb 3 by the pass after that), in the section whose own subject is limbs a derivation never states.
Consequence for NFR-001's own gate: because the pole excludes timing, SC-007's wall-clock assertion belongs in timing-nfr-serial (.github/workflows/ci-quality.yml) — -m timing -n0 over tests/ on the same blacksmith-4vcpu-ubuntu-2404, if: always(), no filter gate, no needs: edge, and wired into quality-gate.needs so a red timing gate blocks merge. That is gating, serial, uncontended and always-on at once. It measures the guard uncontended while the pole runs it under -n auto, so its figure is a floor, not the worst case — raising the budget against it requires the contention headroom to be stated, not assumed.
Corrections Deferred to R1b
Both target files that do not exist on this baseline. Verified: neither is present on feat/isolated-home-pin-guard, and both live only on spike/isolated-home-3121. They are struck from C-006's blast radius and carried here so R1b inherits them.
| # | Correction | Where it lives |
|---|---|---|
| D-1 | evidence/ablation/VERDICT.md §8 reads GREEN 147/147 for the identity member; that fixture is self-bound and governs exactly 6 tests — live collection on this baseline: module 147 nodes, class 6. It should read GREEN 6/6 governed (147/147 module nodes green). A presentation defect, not an evidence gap. | spike/isolated-home-3121 only |
| D-2 | scripts/mutants/ablate_home_pin_3121.py:482–499 — home_ok's within_tmp tests the current test's tmp_path (item.funcargs.get("tmp_path")) and misses tmp_path_factory siblings, producing 15 false-positive home violations, all tests/cli/commands/test_sync_routes.py. Repair before reuse. (The defect was confirmed by reading the code on the spike branch; the count of 15 was not re-measured.) | spike/isolated-home-3121 only |
Tooling Gaps Filed by This Mission
Filed in the record for the operator to route. No issue was created — C-013. (Housekeeping from the post-plan gate: the rows below are ordered TG-1, TG-2, TG-4, TG-3 — read by ID, not by position. TG-1 is referenced from meta.json's hand-edit rationale and TG-3 from §0.6's no-precedent argument; both were otherwise defined and never cited.)
| # | Gap | Evidence | Worked around how |
|---|---|---|---|
| TG-1 | No CLI surface updates a Mission's purpose_tldr / purpose_context after creation. spec-kitty agent mission offers ten subcommands, none of which edits the stakeholder blurb, so it goes stale the moment the spec is revised. | spec-kitty agent mission --help (v3.2.5) | meta.json's purpose_context hand-edited and re-validated as JSON — an edit to a tool-managed artefact, which the canonical-sources rule discourages. |
| TG-2 | CLAUDE.md:256 is stale on where repo-wide gates run. It says "Some repo-wide gates run only in CI's integration-tests-core-misc job, NOT in the fast-tests- suites"*. That routing was superseded by the ci-topology-shrink extraction; tests/architectural/ now runs in the always-on arch-adversarial pole on 100% of PRs. The practical advice in that line remains sound. | .github/workflows/ci-quality.yml:2007-2026, :1872-1875 | OD-004 resolved against the workflow file rather than the doc. |
| TG-4 | The repo's own shrink-only baseline is co-located with the artefact it bounds, has no generator, and has no co-edit guard. load_baseline(ALLOWLIST_PATH) reads charter_path_literal_baseline: 49 from tests/architectural/charter_path_literal_allowlist.yaml — the same file as the allowlist. The token field is documented as "a FROZEN tool-derived code_tokens_by_line string — never typed by hand", but no generator exists in the tree. | Read directly at test_charter_path_literal_authority.py:116, 934-935 and charter_path_literal_allowlist.yaml:25 | Not worked around — the gate is sound in practice because exact accounting (len(allowlist) == len(live_sites())) makes the co-location survivable. R1a states its own placement rule (FR-004) rather than copying, and records the insight. |
| TG-3 | No executable pre-filter coverage proof exists in this repository to follow. Two byte pre-filters ahead of ast.parse ship today — tests/architectural/_sole_door_scan.py:461-476 and tests/architectural/test_commit_target_kind_guard.py:186-188 — and both argue soundness in a comment only. | Read directly; no test_prefilter exists | R1a sets the precedent (NFR-002, SC-002b) rather than inheriting one. |
Note Carried to R1b
Three defects in this specification's history had one shape: a rule changed in one section, its consequences left standing in others. An exemption moved out of an audited artefact into the guard's source; an attribution rule bound at the keyed def while three published figures still reported the innermost; a hash pinned "outside the artefact" with no rule governing where outside. Each was caught by review, none by the author.
The check is cheap and belongs in R1b's working habit: after changing any derivation rule, grep the derived figures. The 30/9/1 split survived four separate publications of a document that had already changed the rule producing it. That one line is more useful to R1b than any figure in this specification.
Provenance
| Claim | How it was established |
|---|---|
| Baseline | git fetch upstream; upstream/main = 5d49d31ed…, 132 commits ahead of the halted branch |
| Every population figure in §0.1, §0.6a, §0.8 | A binding-resolving AST classifier written for this specification and run on this baseline. Text search established nothing. |
| Reach figures, value buckets, dispersion | Same classifier; buckets computed over all 191 write sites with per-bucket file counts and kind splits |
| Scope-chain attribution (§0.1b) | Measured both ways: innermost → 39 members, scope-chain → 40; symmetric difference is exactly …::_run_once |
| Env-key indirection population | AST over src/ ∪ tests/: 0 constants valued "SPEC_KITTY_HOME"; 73 of 1245 setenv/delenv calls use a non-literal key |
| Rename-signature collisions | On the detector's population (the effect class): 3 files, n = 2 each, all same-kind. Exact signature (tmp_path/home, {monkeypatch, tmp_path}) spans 38 sites at the keyed def (37 at innermost). The all-191 figure previously quoted is withdrawn — wrong population (§0.9) |
| Window churn | 17 of the 100 byte-hit files, and 3 of the 36 member files, changed between the window SHAs |
| Timings | Three warm runs each on the external pytest-9.0.3 venv (Python 3.12.13). Operator-recorded figures quoted separately and attributed. |
tests/conftest.py byte-identity | git diff --stat spike/isolated-home-3121 HEAD -- tests/conftest.py — empty |
| Arrival commits | git log --diff-filter=A per file; authored and committed dates read separately |
| 6 vs 147 collected nodes | Live pytest --collect-only, class and module separately |
| Window validation | git merge-base → 709a595…; classifier over a git archive extraction yields 28 members there under the superseded predicate, 30 here |
composite_key, _ratchet_keys, census/, sole door, shrink-only baseline | Read directly at the cited paths and line numbers |
| CI routing (OD-004) | .github/workflows/ci-quality.yml read directly |
| D-1 / D-2 file absence | git cat-file -e against this baseline — both absent; present only on spike/isolated-home-3121 |
| ~~Deliberately NOT measured~~ — SUPERSEDED | As authored: `\ |
| **`\ | R\ |
| Not re-derived, flagged | The runtime baseid derivation ([INHERITED], NFR-006); the 15 false-positive count in D-2 |