Rank 5 on the docket, and the first item this arc whose expected outcome is
LOSING capability.
conservative_floor is a one-sided Wilson lower bound, and Wilson assumes
independent trials. CORE's pipeline is deterministic, so replaying an identical
case is not a second trial — it is the same trial with a guaranteed outcome.
The deduction sealer folded the raw practice corpus and counted every replay,
so bands recorded 720/720 committed whether that was 720 decisions or 28
decisions seen 26 times each. conservative_floor(720,720)=0.990868 cleared
theta_SERVE=0.99; conservative_floor(28,28)=0.808413 is not close.
RESULT: 25 licensed bands -> 4. Twenty-one demoted.
Survivors: en_conditional_chain, en_disjunctive, en_verb_fact,
en_verb_universal — the only bands whose corpus holds enough independent cases.
What the demotion did and did not do, stated exactly. NO ANSWER CHANGED. No
answer became wrong. `wrong` stayed 0 across all 25 bands. The engine is
exactly as correct as it was. What changed is the CLAIM attached to the answer:
21 bands now serve the same sound conclusion prefixed with "(reasoned, but I
haven't yet earned a verified track record on arguments of this shape)". The
reasoning did not get worse; the boast did.
The producer is hardened, which is the half that stops this recurring.
seal_ledger now refuses outright — assert_sealed_evidence_distinct compares the
ledger about to be written against the corpus's distinct-case count and raises
BEFORE anything is written. The curriculum sealer has carried this guarantee
since ADR-0264 R9; the deduction sealer had no equivalent, which is exactly how
the exposure arose. Catching it at seal time rather than in an audit matters:
an audit finds a padded ledger after it is committed, trusted, and gating a
live flag, and unwinding that then needs a ruling — it needed one.
A HOLLOW GUARD OF MY OWN, caught by sabotage rather than trusted. The first
version compared all_gold_problems() with distinct_gold_problems() — two
functions that agree with each other by construction, and neither of which is
what build_ledger folds. Reverting build_ledger to the raw corpus walked
straight past it AND re-sealed the inflated artifact. Same defect class the
ledger already suffered: a guard checking something adjacent to the thing it
protects. It now takes the BUILT LEDGER as input, because the invariant is
about the artifact. Re-running the identical attack: refused, wrote nothing.
Pinned as test_sealer_refuses_a_ledger_that_counts_replays.
A SECOND MISS, caught by the gate rather than by me. My blast-radius scan
grepped for deduction_serve_license / deduction_serve_ledger / 720 and found
nine files. It missed tests/test_construction_inventory.py, which asserts
deduction SURFACE STRINGS and never mentions the ledger — so a pattern search
over the wrong noun could not see it. The full gate found it. Recorded because
"I scanned for the blast radius" is worth exactly as much as the scan's key,
and mine was the wrong key.
That test is the #138 fabrication defect pin, and it is updated WITHOUT being
weakened. Both surfaces now carry the disclosure prefix, and the fabrication
survives untouched — which is a sharper finding than before: a disclosure
prefix looks like it might cover the hazard, and it does not. The user is still
told they said "furthermore". Exact-match assertions kept on the full strings
so neither the fabrication nor the licence state can drift unnoticed, plus an
isolating assertion that the only difference between clean and leaked recitals
is the fabricated premise.
Four records corrected in the same PR, each of which would otherwise have gone
on asserting capability that no longer exists:
- CAPABILITY_LEDGERS' manifest note ("25 bands at 720/720 wrong=0")
- the published CLAIMS.md claim — caught immediately by G-22's own pin, which
is that fix earning its keep three days after landing
- test_volume_honesty.py's module docstring, which said the exposure was
"NOT silently repaired ... changing it is Shay's ratification, not a
test's". R-13 IS that ratification, so the line went stale by being acted on
- the exposure inventory's framing: it now pins the repair, not the gap
test_volume_honesty's sealed-vs-producer assertion is strictly stronger than it
was. It compared sealed committed counts to the producer's committed counts and
both were the inflated 720 — the assertion held perfectly while the number it
agreed on was the wrong number. Two records agreeing is not evidence that
either is right. It now pins committed == distinct, plus a non-vacuity check.
Nine tests encoded the unearned expectation and were updated to the honest one.
Closes G-19 and H-1.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
Phase 5 item 1 asked for the overlap to be sized "by measuring the reader's
construction set against the writer's rather than by growing corpora blindly".
New lane `evals/construction_inventory` does that: it sweeps the writer's whole
parameter space (every RhetoricalMove x IntentTag x predicate x quantifier x
tense x aspect, both public entry points, 32292 cells), quotients it by the
reader's OWN function-word skeleton, and comprehends each construction under
three vocabularies.
writer constructions 1739
reader constructions 19 (mint-site AST-guarded)
overlap, faithful 6
accepts but MIS-READS 22
refuses 1711
vocabulary-dependent 0
Faithfulness, not acceptance, is the criterion — and that reverses two claims
in the plan's §6 RESULT, using a metric that was already on the page:
* g_read_rate is 1/293 but **g_args_rate is 0.0**. The one surface that "reads",
`all molecules are defined as compounds`, is comprehended as
`subset(molecule, defined_as_compound)` — the reader chunked the writer's verb
phrase into a class name. It accepted; it did not comprehend.
* "one construction wide" was wrong both ways: zero on that corpus, six over the
writer's actual output space. A corpus cannot report an inventory's size.
Three findings re-order Phase 5 item 1:
1. The reader FABRICATES on 22 constructions — neither reads nor refuses.
`every dog is a mammal` -> member(every_dog, mammal);
`furthermore, all dogs are mammals` -> asserted(furthermore).
Both are ordinary user English, not writer artefacts. Root cause: _RESERVED
lacks the function words the writer emits, and _parse_propositional accepts
any single token as a fact.
2. It reaches SERVED output. deduction_surface recites
`Given: furthermore; p implies q; p.` — a premise the user never stated — and
chat/runtime.py realizes declarative turns into the held self, so a
fabricated atom is vault-writable. Widening the inventory first would widen
the fabrication surface with it.
3. ADR-0265's defect class survives inside ADR-0265's designated owner of clause
grammar: the four aspect arms of _inflect_predicate bind `negated` to a
wildcard and never read it, so `dog has been defined as mammal` serves both
the assertion and the denial (10530/16146 points). Unreachable today — no
producer sets aspect — so a loaded gun, not a casualty. It survived because
ADR-0265's invariant is structural (is `negated` threaded?) and cannot see an
arm that receives the flag and ignores it. A behavioural sweep can.
The two fixes are written already, as the mutations that turn the defect pins
red. They are NOT applied here: they change what CORE comprehends from user
input, which is a serving change on the truth path — authorization-gated, ADR
first, on the ADR-0261 §5.1 refuse-don't-drop precedent.
Guards, so the tables cannot rot: reader constructions are pinned to an AST
count of the reader's 10 mint sites; the declared tense/aspect axes are pinned
to _inflect_predicate's match arms; the committed corpus is pinned to the
fillers in use. 14 pins, 13 mutations observed RED.
[Verification]: in-worktree, CPython 3.12.13, uv sync --locked.
deductive 517 (was 503, +14 — count moved, registration confirmed)
smoke 641 (unchanged; the file is registered in `deductive`)
lane SHAs 11/11 match, no pin edited
pyright 0 errors on all new files
mutations 13/13 red, including both fabrication fixes