Commit graph

5 commits

Author SHA1 Message Date
Claude
71d6366fa3
feat(ledger,tests,adr): PR-14 — the outcome-mix rule, and the four expected licences dissolve under it (R-8)
Rank 6, the last docket item before the posture statements, and the throughput
frontier N-5 called "one ruling wide". The ruling landed. It closed the frontier
rather than opening it, and that is the correct outcome.

R-8 ruled C: committing to an entailment and correctly declining to commit are
DIFFERENT capabilities, licensed on DIFFERENT evidence, and may not be pooled.

THE FINDING: split, ZERO bands license — not four.

  band                              unknown   floor      entailed  floor
  philosophy_theology_contrast          652   0.989925          8   0.000000
  philosophy_theology_modal             652   0.989925          8   0.000000
  physics_causal                        653   0.989940          7   0.000000
  systems_software_causal               653   0.989940          7   0.000000
  (pooled, the old basis)               660   0.990046  -- clears by 0.000046

N-5 recorded that "four bands would earn SERVE the moment a ledger is sealed."
That is true on the POOLED basis and false on the ruled one. Those four cleared
theta_SERVE=0.99 only because 7-8 entailments were counted alongside 652-653
correct refusals to reach 660. The licence was manufactured by the pooling, not
earned by the evidence. conservative_floor(9,9) is 0.000000 — the Wilson lower
bound at nine trials is not merely below theta, it is zero.

This is exactly why the plan blocked PR-14 on R-8 instead of sealing first and
ruling after. A ledger sealed on the pooled basis would have granted four
licences that the mix rule then had to revoke — and revoking a granted licence
is the expensive direction, as PR-12 had just demonstrated across 21 deduction
bands in this same session.

DELEGATED RULING — R-8 C's entailed floor N: no new constant.
The packet recommended "C, with a floor from A applied to the entailed
capability only," and N was never named. It does not need naming. theta_SERVE
=0.99 through conservative_floor — the bar every other capability already meets,
657 distinct correct decisions — applied to each separated capability on its own
evidence licenses nothing, by a factor of 73. Inventing a second, weaker
constant for the capability that most needs the strong one would institutionalise
two standards, which is precisely the objection that sank option B. Recorded
under the standing delegation with its reasoning, so it can be overturned.

NO LEDGER IS SEALED, AND THAT IS THE DELIVERABLE. Under the rule there is
nothing to license, and the registered missing_ok=True absence already says so
once. An artifact whose only content is its own emptiness would be a second
statement of the same fact — the defect class registered as G-23 this same week.

THE USEFUL RESULT IS THE SHAPE OF THE GAP, WHICH POOLING HAD HIDDEN.
Non-commitment serving is FOUR TO FIVE distinct query atoms short per band.
Entailed serving is ~648 short. Those are content tasks of completely different
size, and a single pooled figure could not distinguish "almost there" from "two
orders of magnitude away". Four cases is a morning's work; 648 is a program.
Nobody could see that before the split.

Delivered:
  - curriculum_serve_entailed registered in CAPABILITY_LEDGERS, missing_ok=True
    (absent = nothing licensed, the honest state), with the rule declared in the
    manifest table rather than at a call site (ADR-0263 rule 5)
  - an audit source for it — demanded immediately by the manifest's own pin
    (test_every_licensed_capability_has_an_audit_source went red the moment the
    capability was declared, which is PR-5's declared-table discipline working)
  - ADR-0264 §5 amendment recording the ruling and correcting §4.1's "four bands
    would earn SERVE" expectation. Changes no decision in that ADR;
    curriculum_serving_enabled stays False
  - tests/test_curriculum_outcome_mix.py on the gate: both capabilities pinned,
    the shortfall pinned exactly, AND the counterfactual pinned — pooling
    licenses 4 — so the argument for the rule cannot drift away from the number
    it rests on

The serving path needs no change, and that is stated rather than left implicit:
under the rule nothing is licensed, so there is no licensed-vs-disclosed branch
to route. Building that machinery now would be building for a case that cannot
occur yet.

Closes G-10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
2026-07-28 07:59:01 +00:00
Shay
b970054226 docs(adr): ratify ADR-0264 — Accepted, and §4.1 self-superseded by its own R5
Ratified by Joshua Shay 2026-07-26. All nine rules are implemented and merged to
main @ 803703fc: R5-R7 (#120), R9 (#117), R1-R4/R8 (#125). The status line names
the enforcement pins so the record points at what holds it up.

Also discharges a staleness the ratification would otherwise have blessed. §4.1's
heading — "No curriculum band can earn a SERVE license under the current
architecture" — is NO LONGER TRUE of the running system, and it was falsified by
this ADR's own R5: the finding rested on "the premise cap holds a family to <=16
chains", and query-scoped compilation removed that cap. Post-R5, four bands would
earn SERVE the moment a ledger is sealed. Leaving that heading unqualified in an
Accepted ADR is precisely the ADR-0256 shape — a governance record asserting
something false about live code.

The §4.1 CONCLUSION survives, for a different and stronger reason, and the note
says so: reliability is commitment precision and a correct UNKNOWN is a
commitment, so a band clears theta_SERVE on non-commitments alone
(conservative_floor(660,660)=0.990046). Licensable evidence is 99.0-99.98%
non-entailed; max entailed volume in any band is 9. So §4.1's "≥97.6%
non-entailed" arithmetic was directionally right and its CAUSE was wrong — it is
not the premise cap that makes a curriculum license dishonest, it is the absence
of a mix rule. The binding constraint is now the outcome-mix ruling (§5, still
open), and nothing is licensed while it is unmade: the ledger is absent and
`core proposal-queue reseal` refuses to grant a license without
--allow-new-licenses.

[Verification]: in-worktree on canonical CPython 3.12.13 with `uv sync --locked`:
smoke 621 (unchanged), and 335 passed across test_adr_status_governance.py /
test_adr_index.py / test_volume_honesty.py. Status parses as "Accepted …" under
the governance regex and the banned `ratify-on-merge` predicate is absent. No
flag moved; no code changed.
2026-07-26 14:39:56 -07:00
Shay
923e60c136 docs(arc): reconcile the arc's records with what the units measured
Every correction here is to something I wrote earlier in this arc. Original text
is preserved in place; nothing was rewritten to look right in hindsight.

ADR-0264 §4.2 — CORRECTED block. I sized the eleven band ceilings from per-TERM
exclusivity; `resolve_domain`'s predicate is per-PAIR, which is looser, so every
ceiling was understated and one verdict was wrong: systems_software_causal is
720, not 630, i.e. reachable rather than impossible. Seven of eleven bands
cannot reach 657, not eight. Downstream conclusions survive — physics·modal (480)
is still impossible and philosophy_theology·modal is still the target.

ADR-0264 R5 + DIVISION-OF-WORK ×2 + compile_premises docstring — the
"verdict-identical" claim is exact only where the full family was READABLE. Over
MAX_PREMISE_SENTENCES, full-family compilation refuses outright, so narrowing is
verdict-IMPROVING there. Read literally the unqualified phrasing implies a
17-chain family keeps declining, which is backwards — the opposite of the
unblock.

E2 rename staleness, 4 sites: ADR-0256 (Accepted) named
`gold.py::assert_corpus_sound`, which #119 deleted; the plan's E2 item and
DIVISION-OF-WORK's §2 row, §4 heading and "renaming after that is strictly more
work" reasoning were all still written as pending.

DIVISION-OF-WORK §0b — new status table: which PR each unit is, what stacks on
what, and the four things the units found that the plan did not predict.

Adds the arc-close brief from ARC-CLOSE-TEMPLATE.md. Its §3 (falsified
assumptions) is the section that cannot be recovered from the diff: nine
entries, five of them corrections to my own claims. Records that the
articulation trigger is not merely unmet but STRUCTURALLY unmeetable — the
previous arc left it unrecorded, which is what kept it re-triggerable — and
names four pieces of stale reasoning still in the tree, with files, since that
is the class of staleness nobody greps for.

[Verification]: in-worktree on canonical CPython 3.12.13 with `uv sync
--locked`: smoke 569, deductive 291 — both baseline, since the only non-doc
change is one docstring.
2026-07-26 12:33:30 -07:00
Shay
13a4b05d62 test(reliability): volume-honesty invariant + distinct-evidence audit
Phase B of the curriculum-license-loop arc. Stacked on the ADR-0264 branch.

The invariant: conservative_floor is a one-sided Wilson lower bound and Wilson
assumes INDEPENDENT trials. CORE's pipeline is deterministic, so replaying an
identical case is not a second trial -- it is the same trial with a guaranteed
outcome. A ledger's `committed` count is an upper bound on its evidence; the
defensible figure is its distinct-decision count.

Phase B was scoped as a precaution for a curriculum producer that does not
exist yet. Running the instrument against the producers that DO exist found it
already violated:

  21 of the 25 ratified deduction_serve bands do not clear theta_SERVE=0.99 on
  distinct evidence. Three inflate 28 distinct cases into 720 committed
  (honest floor 0.8084 vs claimed 0.9909). deduction_serving_enabled was
  ratified ON 2026-07-24 and is live.

The estimation producer (ADR-0175) is clean -- 660 committed, 660 distinct,
zero repeats, evidently chosen just above the 657 a perfect record needs. So
this is a regression from a standard already established in the repo, not an
architectural gap. `claimed` is identical for all 25 deduction bands because
every band commits 720 with a perfect record: the gate sees 25
indistinguishable passes while honest floors span 0.8084..0.9909.

NOT REPAIRED. The ledger is SHA-sealed, ratified, and gating a live flag, so
re-sealing it is Shay's ratification, not a test's. wrong=0 still holds -- no
band answered anything incorrectly. What is not established is reliability at
the volume claimed. Measured, pinned in both directions, written down.

Added:
- core/reliability_gate/evidence.py -- pure measurement, imported by no serving
  path. volume_for_theta() DERIVES 657 from conservative_floor rather than
  restating it, so a WILSON_Z change cannot leave a stale literal behind.
- tests/test_volume_honesty.py -- 13 tests. AUDIT_SOURCES is checked against
  CAPABILITY_LEDGERS, so a new licensed capability cannot ship unaudited; the
  curriculum entry is a declared None that fails the moment its ledger exists.
  That is Phase C's forcing function against repeating the pattern.
- docs/research/distinct-evidence-audit-2026-07-25.md -- full audit + the open
  outcome-mix ruling for Shay.
- ADR-0264 R9 amended with the measured exposure.

Decision key is case TEXT deliberately: a tighter key (normalized atom) would
collapse spelling variants and report MORE inflation, so text-identity
under-reports and every number is a floor on the real gap.

Outcome mix is deliberately NOT pinned. ClassTally has no verdict axis, the
deduction producer already balances 240/240/240 by construction, and imposing
a per-verdict-class floor would retroactively fail all 25 bands on a criterion
no ADR has ratified (smallest per-band class count is 120). Recorded for a
ruling with numbers instead.

[Verification]: uv sync --locked on canonical CPython 3.12.13; in-worktree
smoke 569 passed (556 baseline + 13), deductive 285 passed. Registered in the
`smoke` curated suite -- not left to `full`, which gates nothing. Mutation-
checked: doubling CASES_PER_BAND -> 2 failed; reporting committed as distinct
-> 4 failed; gating on `claimed` instead of `honest` -> 3 failed; tree restores
to 13 passed.
2026-07-25 16:03:10 -07:00
Shay
a35f6f7bd7 docs(adr): ADR-0264 — negative curriculum, premise scope, and band ceilings
Phase A of the curriculum-license-loop arc. Design only; no production code,
no flag, no default changed.

Rules R1-R9: polarity is a row-level field (absent => affirmative), negatives
compile to the sentential-not form and MUST reuse the affirmative connective
(atom identity), contradiction is rejected at ratification, premise
compilation becomes query-scoped, an empty scope is UNKNOWN, premise_count
keeps reporting family size, the oracle learns polarity independently, and a
band's honest ceiling must be reported before anyone authors for it.

Three findings, each measured in-tree and each reversing something:

1. `polarity` is read by nothing. A row authored to refute a question today
   serves a confident wrong "Yes" -- and the independent oracle ignores the
   field for the same reason, so gold agrees and wrong=0 stays green. This is
   the silent-failure class Phase A existed to prevent, now demonstrated.

2. The 16-premise honesty cap holds a family to <=16 chains; at 17 the band
   declines everything, including the taught positive that works today. So
   ADR-0262 §5.1's own remedy (author ~219 relations) destroys the capability
   at row 17, tripping the fix §6 deferred as a future contingency. Jointly:
   a band has <=16 entailed cases against a 657-committed threshold, so no
   curriculum band can earn SERVE until scoping lands. The blocker is
   engineering, not content -- inverting the conclusion two arcs closed on.

3. Honest volume is bounded by taught vocabulary, not corpus size. 8 of 11
   bands cannot reach 657 at all, including the plan's chosen first target
   physics·modal (ceiling 480, counting every possible question). Retargets
   Phase F to philosophy_theology·modal (ceiling 44,104, same 8-chain start).

R5 is the only rule that could change an existing answer, so it is the only
one carrying evidence: 8,520 routable questions, 0 verdict mismatches for both
candidate scopes. Verdict-identity holds for any scope that is a superset of
the query-atom rows, since each premise mints one independent atom. That is
stronger than ADR-0262 §6's soundness argument, and it is why this narrowing
is not the ADR-0261 §5.1 premise-dropping failure.

ADR-0262 §5.1/§5.2/§6 carry forward-pointers; original text preserved.

[Verification]: uv sync --locked on canonical CPython 3.12.13; in-worktree
`core test --suite smoke -q` 556 passed (555 baseline + 1: the derived
ratify-on-merge pin picked up ADR-0264 and discharged it), `--suite
deductive -q` 285 passed. Four probes reproducible from
docs/research/curriculum-premise-scope-2026-07-25.md; the corpus-mutating one
restores the file and asserts byte-identity.
2026-07-25 15:44:25 -07:00