core/CLAIMS.md
Claude 9cf6a164d5
feat(evals,chat,tests): PR-12 — the Wilson re-count. 25 licensed bands become 4 (R-13)
Rank 5 on the docket, and the first item this arc whose expected outcome is
LOSING capability.

conservative_floor is a one-sided Wilson lower bound, and Wilson assumes
independent trials. CORE's pipeline is deterministic, so replaying an identical
case is not a second trial — it is the same trial with a guaranteed outcome.
The deduction sealer folded the raw practice corpus and counted every replay,
so bands recorded 720/720 committed whether that was 720 decisions or 28
decisions seen 26 times each. conservative_floor(720,720)=0.990868 cleared
theta_SERVE=0.99; conservative_floor(28,28)=0.808413 is not close.

RESULT: 25 licensed bands -> 4. Twenty-one demoted.

Survivors: en_conditional_chain, en_disjunctive, en_verb_fact,
en_verb_universal — the only bands whose corpus holds enough independent cases.

What the demotion did and did not do, stated exactly. NO ANSWER CHANGED. No
answer became wrong. `wrong` stayed 0 across all 25 bands. The engine is
exactly as correct as it was. What changed is the CLAIM attached to the answer:
21 bands now serve the same sound conclusion prefixed with "(reasoned, but I
haven't yet earned a verified track record on arguments of this shape)". The
reasoning did not get worse; the boast did.

The producer is hardened, which is the half that stops this recurring.
seal_ledger now refuses outright — assert_sealed_evidence_distinct compares the
ledger about to be written against the corpus's distinct-case count and raises
BEFORE anything is written. The curriculum sealer has carried this guarantee
since ADR-0264 R9; the deduction sealer had no equivalent, which is exactly how
the exposure arose. Catching it at seal time rather than in an audit matters:
an audit finds a padded ledger after it is committed, trusted, and gating a
live flag, and unwinding that then needs a ruling — it needed one.

A HOLLOW GUARD OF MY OWN, caught by sabotage rather than trusted. The first
version compared all_gold_problems() with distinct_gold_problems() — two
functions that agree with each other by construction, and neither of which is
what build_ledger folds. Reverting build_ledger to the raw corpus walked
straight past it AND re-sealed the inflated artifact. Same defect class the
ledger already suffered: a guard checking something adjacent to the thing it
protects. It now takes the BUILT LEDGER as input, because the invariant is
about the artifact. Re-running the identical attack: refused, wrote nothing.
Pinned as test_sealer_refuses_a_ledger_that_counts_replays.

A SECOND MISS, caught by the gate rather than by me. My blast-radius scan
grepped for deduction_serve_license / deduction_serve_ledger / 720 and found
nine files. It missed tests/test_construction_inventory.py, which asserts
deduction SURFACE STRINGS and never mentions the ledger — so a pattern search
over the wrong noun could not see it. The full gate found it. Recorded because
"I scanned for the blast radius" is worth exactly as much as the scan's key,
and mine was the wrong key.

That test is the #138 fabrication defect pin, and it is updated WITHOUT being
weakened. Both surfaces now carry the disclosure prefix, and the fabrication
survives untouched — which is a sharper finding than before: a disclosure
prefix looks like it might cover the hazard, and it does not. The user is still
told they said "furthermore". Exact-match assertions kept on the full strings
so neither the fabrication nor the licence state can drift unnoticed, plus an
isolating assertion that the only difference between clean and leaked recitals
is the fabricated premise.

Four records corrected in the same PR, each of which would otherwise have gone
on asserting capability that no longer exists:
  - CAPABILITY_LEDGERS' manifest note ("25 bands at 720/720 wrong=0")
  - the published CLAIMS.md claim — caught immediately by G-22's own pin, which
    is that fix earning its keep three days after landing
  - test_volume_honesty.py's module docstring, which said the exposure was
    "NOT silently repaired ... changing it is Shay's ratification, not a
    test's". R-13 IS that ratification, so the line went stale by being acted on
  - the exposure inventory's framing: it now pins the repair, not the gap

test_volume_honesty's sealed-vs-producer assertion is strictly stronger than it
was. It compared sealed committed counts to the producer's committed counts and
both were the inflated 720 — the assertion held perfectly while the number it
agreed on was the wrong number. Two records agreeing is not evidence that
either is right. It now pins committed == distinct, plus a non-vacuity check.

Nine tests encoded the unearned expectation and were updated to the honest one.

Closes G-19 and H-1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
2026-07-28 07:42:58 +00:00

4.3 KiB

CLAIMS

Every row below is mechanically derived from in-tree state and verified by CI. Tier 1 rows come from core.capability.ledger_report; Tier 2 rows come from scripts/verify_lane_shas.py's pinned SHAs.

Tier 1 — Ratified domains

Each row asserts: the domain's packs pass all nine ADR-0091 contract predicates, declared operator chains meet the ≥8 / ≥3-intent floor, the reviewer registry resolves the primary reviewer, and the ledger status predicate evaluates to reasoning-capable with no open gaps.

Domain Status ADR Packs Open gaps
philosophy_theology reasoning-capable ADR-0085 2 0
mathematics_logic audit-passed ADR-0097 1 0
physics audit-passed ADR-0100 1 0
systems_software audit-passed ADR-0101 1 0
hebrew_greek_textual_reasoning reasoning-capable ADR-0102 4 0

Tier 2 — Pinned lane reports

Each row asserts: running the lane's runner produces a JSON report whose SHA-256 matches the pinned value below. Mismatch is a CI failure (.github/workflows/lane-shas.yml).

ADR Lane Purpose Report path Pinned SHA-256
ADR-0092 reviewer_registry Reviewer registry schema validates + bootstrap entry self-seals evals/reviewer_registry/results/v1_dev.json 681a2aab5aa4ffd58cd837ce5673c8b2a9545b570117aec3c02726a12f6876e6
ADR-0093 domain_contract_validation All ratified packs satisfy the 9 ADR-0091 contract predicates evals/domain_contract_validation/results/v1_dev.json 98ace04e3f02bbc5a8ad655bb6593c3f1ee64cb67014f1122fe6c3c85f48d22f
ADR-0095 miner_loop_closure Miner-sourced proposals route through single reviewed teaching path evals/miner_loop_closure/results/v1_dev.json 537094fe21d7e6cfbaf42bfc32b82d669fa9bb05a132d2bc93c72b3ceb7762a6
ADR-0096 fabrication_control_summary Phantom endpoints / cross-pack non-bridges / sibling collapses refuse evals/fabrication_control/results/v1_summary.json 01e1b6b711141f2b4a14551d7df3ea482d8d6dd7b364a25c509f4f8d08cda8a8
ADR-0098 demo_composition Demos compose from shipped modules; no parallel mechanism evals/demo_composition/results/v1_dev.json f0611a2ce41721dd40767fc6a83a08470d3c7fd7fc8f1ae8ba003abf8a25ec97
ADR-0099 public_demo Public showcase runs deterministically under 30s; all claims supported evals/public_demo/results/v1_dev.json da7fad654e77aac4573a6fcf6e9eaaf84540be8e135d2e033d9cfd15119df3fc
ADR-0104 curriculum_loop_closure Curriculum-sourced proposals route through single reviewed teaching path evals/curriculum_loop_closure/results/v1_dev.json cb94ca0042d78ec2624129ff6493d52e767b69feea32d2997b85d88f1c0883af
ADR-0131 math_teaching_corpus_v1 Math teaching corpus replays deterministically; all chains pass exit criterion (correct_rate=1.0, wrong=0) evals/math_teaching_corpus/v1/report.json eaf160d145da29f9050ede8d58bf111b0f651dd40aeae9201857d0b97e014dd4
ADR-0206 deductive_logic_v1 Propositional entailment scored against an independent truth-table oracle; dev+holdout+external 716/716 correct, wrong=0, refused=0 evals/deductive_logic/report.json 97a230949016e38d5e3f37a69e4245b320575ee70e5af92ff7607f7b05f74b5f
ADR-0256 deduction_serve_v1 Flag-gated deduction serving decides real English/member/fused/verb/existential arguments end-to-end, wrong=0 across all splits; 4 of 25 shape-bands hold an earned SERVE licence on DISTINCT evidence (R-13 re-count, 2026-07-28) and the other 21 decide with disclosure evals/deduction_serve/report.json c855d55cf316471fdfe092aa0d5c954e5ceb6f30c9a8db283e0b9a5d5e8b419a
ADR-0262 curriculum_serve_v1 Flag-gated curriculum serving answers exam questions from a subject's RATIFIED chain corpus only; untaught facts return UNKNOWN (anti-recall probes enforced), wrong=0 evals/curriculum_serve/report.json d9e7ba500f040b865870413a940ee9a49910ac22e1a89c9feec1a60bdd2513f1

Verification

python3 scripts/verify_lane_shas.py    # verify all Tier 2 SHAs
core test --suite full -q              # exercises Tier 1 invariants