Reverts the c3aed13b parser change (frame-anchored comparative regexes +
verb-free connector + _COMPARE_POLARITY_BLOCK + 'as much'); the reader stays
on the _comparison_anchor_verb() whitelist (fail-closed). Increment 1 ships
the compiler tier + the funnel finding only.
Three reasons on record: (1) frame-anchoring converts 0 of 500 official cases
— zero measured benefit today; (2) its consumer is rate-frame admission, a
later increment — the next increment (seeding-sentence injection) is a shared
layer that doesn't touch it; (3) the polarity blocklist is fail-OPEN — an
unenumerated depletion verb in a complete frame is a coherent-but-inverted
graph = a wrong on the run-once sealed arbiter (round-trip guards parse
consistency, not gold-correctness), a trade the fail-closed doctrine never
takes for zero gain.
Verified before revert-merge: with the whitelist reader, the 5 real compare
parses are non-divergent — test_real_compare_cases_solved_wrong_zero still
sees 5, solves 5, wrong 0 (corridor real-reach 0->5/500 preserved).
Frame-anchoring recorded as an OPEN DESIGN TASK owned by the rate-frame
increment: frame + POSITIVE polarity determination (neutral-polarity allowlist
or small polarity classifier), not assume-safe-unless-blocklisted. Detail in
docs/research/compare-increment-funnel-2026-07-18.md.
Increment 1 reader half. Re-anchor the two multiplicative-comparison frame
regexes on the closed-class frame ('<factor> as many/much <unit> as <REF>')
instead of the _comparison_anchor_verb() whitelist: the verb slot becomes a
verb-free bounded connector, 'as much' joins 'as many'. Admission is now
open-class for safe verbs (scored/wrote/caught/taught/…), gated by the frame.
One closed class is NOT frame-neutral — polarity-inverting/depletion verbs
(lose/spend/give/sell/win/donate/pay/eat/drop/lend): in a compare frame they
make the compared quantity a loss, so actor=factor×reference reads backwards
→ wrong>0. _COMPARE_POLARITY_BLOCK excludes exactly that set (the sanctioned
fallback; the frame is the gate for every other verb). Pins the wrong=0
guards in test_adr_0131_G2a + test_recognizer_comparative_inject.
Develop-on-tune / measure-once held. Footprint over 500 official cases: tune
byte-identical (0 divergences, wrong=0, PARSED=3); measure 1 divergence
(0441 'Zig wrote four times as many books as Flo') which advances
frame-refusal -> question-refusal (sanctioned delta) but does not solve — the
binding wall is the shared summation-question layer. X=15 measured delta=0.
Ruling per the funnel-conditioned escalation rule: mass is in shared layers,
so the next increment is shared-layer (summation-question reader), not a
practice escalation and not more compare frames. wrong=0 held on tune (261),
measure (239), smoke (176).
Pre-registration before the shared-layer edit (PR ruling), all in the record:
1. Refusal-histogram baseline (tune, n=261): 'no injection' = 196/261 (75%) —
a SHARED bottleneck across every family, not a compare defect; 48 statement,
9 question, 3 PARSED (compare). Pinned; assert stable-or-explained after —
only sanctioned delta = compare 'no injection' -> PARSED, disclosed.
2. Funnel-conditioned escalation rule (pre-committed): X=15 miss read via the
funnel — compare-specific mass -> practice ruling; shared-layer mass ->
shared-layer design increment (NOT practice). Load-bearing: practice cannot
learn past a deterministic wall (better candidates still cross the same
pipeline), so shared bottlenecks are design work regardless of X.
3. Reframe recorded: compare is the first slice of a larger reader project
(reader parses ~1% of real GSM8K; 'no injection' 75% is a shared wall).
Not scope creep — the arc seeing its true size.
Code archaeology before treating 'no injection' as a defect: it is a DELIBERATE
narrowing, not broken code. recognizer_anchor_inject.py's discrete-count injector
refuses comparative surfaces with the explicit comment 'Refuse here until a real
compare_multiplicative operation can be emitted' (~L324) and 'does not create any
new admitting path' (~L460). The lockstep thesis of the arc in the code's own
history: the reader refused compare shapes to protect wrong=0 + ADR-0191
completeness because nothing downstream could compile them. The compiler tier
this increment shipped IS that operation -> the fix is LIFTING a now-obsolete
guard whose reason has shipped, not patching injection.
Instrument-first (PR #76 ruling): before touching the parser, measure where the
61 compare-marked TUNE cases die. Funnel: 57/61 die pre-enumeration; of those,
**45 at the injection layer** ('recognizer matched but produced no injection'),
8 'no admissible candidate for statement', 3 'expected exactly one question
sentence'; 3 solve today.
Finding: the leverage is INJECTION (45/61), not the candidate regexes (which
fire on clean clauses but the real multi-clause sentences never reach them) nor
segmentation writ large. Injection is an EMITTER, so the sanctioned fix is
making it inject compare candidates the recognizer already matches — never
loosening the round-trip gate. Fix order by measured mass: injection (45) ->
statement-admissibility (8) -> question-sentence (3). Tune-only development;
single measure run at the end with refusal-by-reason + delta.
Two shapes (PR ruling) pinned compiler-side with tests — reader direction errors
surface as refusals, never wrongs:
- Chained / ordering: a compare whose reference is neither seeded nor
already-defined at that point refuses (compare_reference_not_a_register).
Pinned explicitly with a test (story order != dependency order), not left to
fall out of existing checks.
- Inverse mis-direction: a compare that would DEFINE an already-known (seeded or
earlier-defined) register refuses (compare_redefines_register) — the reader
must bind the unknown side as the target with a fraction factor; the compiler
never overwrites a known quantity.
The 5 real compare cases still solve wrong=0. 27/27 (2 new guards + 25).
[Verification]: uv run python -m pytest tests/test_adr_0250_compare.py tests/test_adr_0250_multi_register.py tests/test_adr_0250_summation.py -q
AMENDMENT 3 (PR #76): the reader grammar family is developed against the TUNE
split only; increment deltas are measured on the disjoint MEASURE split only.
Split = pure function of case id (measure iff sha256(id) even, else tune), so
disjointness is enforced by construction and cannot drift. Fixed 2026-07-18,
never re-seeded. Sizes: tune=261, measure=239. This commit lands BEFORE any
grammar-tuning commit — the git history is the proof the split was pinned in
advance (criterion 3). wrong=0 is verified across the full 500 regardless of
split; only the coverage delta is attributed to the measure split.
Cross-register affine definition: actor := factor x reference. Reads the
reference register's field STATE, dilates by -ln(factor), writes the actor.
Reuses quantity_kernel + the multi-register framework + the 2b summation.
Amendments (PR #76):
1. Summation registers-driven: MultiRegisterProgram.register_order (seed order
then compare-defined order); execute sums ALL registers in that order, so a
compare-DEFINED register (no seed) is included in totals. compile validates
the answer target AFTER ops (a concrete unknown may be compare-defined) and
propagates the defined register's unit from the reference.
2. Compare record: RegisterTurn.reference; MultiRegisterRecord.source_entity
(the reference) + operand_source_digest (reference state digest), both
CONDITIONAL in the payload so affine/transfer/summation record digests are
byte-identical to before compare. Re-verification reconstructs the source
register from records alone.
Compare CREATES a quantity (not a transfer) -> exempt from the conservation pin.
fraction direction folded in (factor < 1). compare_additive is a later increment
(refused). Real official data: the 5 compare parses the reader emits today now
solve end-to-end wrong=0 -> corridor real-reach 0/500 -> 5/500 (loop-works proof).
25/25 (9 new + 16 existing regression, digests byte-identical).
[Verification]: uv run python -m pytest tests/test_adr_0250_compare.py tests/test_adr_0250_multi_register.py tests/test_adr_0250_summation.py -q
Four amendments (Shay, #76 review) folded:
1. Summation registers-driven (NEW CODE): summation orders by program.seeds
(initial_state) today, so a compare-DEFINED register (no seed) is silently
excluded from the total — all 5 real cases are totals. Fix: sum all registers
at solve time, deterministic order (seed order then definition order);
compile-side validation admits defined entities as answer targets with unit
propagation from the reference. Pin: defined register in the certified sum.
2. Compare record spec: record identifies the reference entity + binds the
reference state digest via the existing conditional operand_source_digest
(non-compare/non-summation digests unchanged); deterministic re-verification
reconstructs the dilation's source register from records alone.
3. Tune/measure split pinned deterministically (sha256(id)-ordered) FIRST,
before any grammar work, documented + fixed in advance.
4. 'N times more than' pinned to N× as a disclosed convention; gold mismatches
tracked as recorded findings, never silent wrongs.
X = 15 ratified as the arc escalation threshold (spike §7.3 updated). §6 risks
reclassified (risk 4 -> new code §3.1; risk 1 -> pinned convention §4). Build
proceeds under this ruling — no separate re-ruling.
First increment of the reader design-first arc (dev2 spike §7). Design-first,
for ruling before build.
Ground truth (5 real holdout_dev compare parses): one entity seeded, the other
DEFINED by actor = factor x reference (Comparison, direction times/fraction);
unknown = total -> existing 2b summation closes them once compare is solvable.
Compiler tier: cross-register affine definition actor := factor x reference
(dilate reference STATE by -ln factor, write to actor register). Two new bits
designed explicitly: reads-one-register-writes-another; actor is DEFINED not
seeded (compare introduces the register). NO conservation pin (compare creates,
doesn't move — transfers conserve, compare doesn't). fraction direction folded
in; compare_additive is a later increment. Immediate effect: corridor real-reach
0/500 -> the compare parses that already exist (loop-works proof).
Reader grammar family: reader emits compare for ~5/500 (narrow phrasings);
incidence ~20% (~100/500) -> broaden compare templates; most of the delta.
Measurement: holdout_dev only, wrong=0 floor, sealed untouched, tuned/measured
DISJOINT split, delta vs X=15. Risks: 'N times more' ambiguity, garbled entity
spans, ordering, defined-register-in-summation.
Nothing built; increment PR ships after this plan is ruled on.
Supersedes the practice-loop recommendation (§7/§10): the arc is the reader,
DESIGN-FIRST — per-family increments, grammar family + compiler tier in LOCKSTEP
(§2 proves they can't move separately: the 5 real compare_multiplicative parses
die at the affine compiler today). Each increment = own smoke-gated PR with a
measured holdout_dev delta, wrong=0 floor, sealed test untouched.
- Increment order: compare_multiplicative FIRST (reader already emits it →
0/500 to >0 immediately, cheapest end-to-end proof), then §3 frequency order
(rate → fraction → compare_additive → partition).
- Escalation threshold X PROPOSED = 15 new-correct/holdout_dev per full family
increment (Shay rules with the merge); below it → practice-lane ruling.
Rationale: families carry 20-31% incidence, a paying design converts far more.
- Practice loop demoted to PAPER-DESIGN backup (§7.4) with leverage analysis:
verifier = Tier-1/2 compiler stack, audit = certificate chains, consolidation
= prepare→validate→commit generalized, hygiene = ADR-0119 sealing; genuinely
new = candidate generator + consolidation store. Four §7 constraints carry
over verbatim.
- Design-first work IS the backup's curriculum (§7.6): refusal taxonomies,
tagged incidence, labeled (parse,compile,execute,gold) tuples are artifacts.
§5 floors/scoreboard + §6 sequencing affirmed unchanged; errata merge with this.
Folds in Shay's rulings on the spike:
- ERRATA sections added to ADR-0249, ADR-0250, and both acceptance-evidence docs
(statuses stay Accepted): corpus label 'real GSM8K dev holdout' -> 'CORE-authored
GSM8K-style corpus (ADR-0119.2)'. Mechanism claims (wrong=0, PARITY, conservation,
atomicity, chain-of-custody, 200/200) all stand; reframe not retraction. Covers the
#74 review comment that repeated the label. (Also fixed ADR-0250 evidence status ->
RATIFIED, which the ratification housekeeping missed.)
- Two instruments named + never-blended scoreboard: authored = mechanism-correctness
(200/200 stands as that); official holdout = generalization (5/500 symbolic / 0/500
corridor). public/dev-1 floors relabeled MECHANISM floors, composition-disclosed.
- Censored-incidence fix: parse-INDEPENDENT tagging of 100 official cases (documented
heuristic). 70/100 carry frontier markers (rate 31, fraction 28, compare_mult 20) ->
(c) refuted, frontier COMMON not rare; the table is the reader-arc curriculum by
real-world frequency.
- Reader-practice-loop design with the four non-negotiable constraints (real held-out
data only; wrong=0 gates consolidation; sealed test untouched; holdout_dev split
disjoint from the measured metric). Sequencing: capability->practice->calibration->
serve; reader is the critical path for all three; serve-landing de-prioritized (0/500).
Spike verdict for ruling; nothing built.
Headline correction: evals/gsm8k_math/dev + public are CORE-AUTHORED GSM8K-style
(ADR-0119.2), NOT real GSM8K. The compiler's 200/200 is on authored, mostly
depth-1, frontier-free problems — a mechanism validation, not real-GSM8K
capability. ADR-0249/0250 'real GSM8K' language needs correcting.
Draw-protocol trace: train_sample + holdout_dev + sealed-test = real GSM8K
(deterministic seeded draws); dev/public = authored.
Real measurement (end-to-end on 500 real holdout_dev): Reader A parses 5/500
(1%, refuses 495); my compiler solves 0/500; the 5 parsed are all
compare_multiplicative (frontier), refused by the affine compiler. Corroborates
holdout_dev's historical 'real GSM8K capability 0%'.
Verdict: (a) DOMINANT — frontier + real GSM8K are READER-gated (99% parse-refusal);
(b) confirmed as authoring bias in dev/public; (c) refuted — frontier is common,
not rare. Sequencing: reader is the biggest bucket; more compiler tiers = capability
without inputs; serve-landing premature. Reader-practice-loop NAMED as leading
next-arc (reader = first learnable component; attempt-parse->compile->execute->
gold-check, wrong=0-gated on held-out real data; generalization-preserving).
Public-seal approved with authored-not-real + composition (depth/kinds) disclosure.
Spike verdict for ruling; nothing built.
Ratification: ADR-0250 Status Proposed->Accepted; §10 ruling record stamped
(Shay, citing the acceptance-evidence pack as-is; full-holdout 50/50 wrong=0
affirmed; next arc = seal dev-holdout-2, dev-1 pinned as regression floor).
Review findings (Shay, PR #74, non-blocking):
1. Instrument module docstring updated from 'Three deterministic domains' +
'GSM8K NOT ingestible / no compiler' to the four-domain post-0250 reality
(arithmetic-chain solves the full dev holdout; frontier surfaced in scope).
2. Restored per-case refusal reasons: _corridor_and_baseline now raises
MultiRegisterError (not bare None); the domain loop records refusal.reason
and distinguishes graph_parse_failed. 'Recorded, never silently dropped'
stays literal (0 rows at refused=0, but the path is honest).
3. Tightened the Tier-1 try to wrap compile_turn_program only (execute outside),
so a future typed execute raise surfaces as a Tier-1 failure instead of
silently rerouting into Tier-2 and masking a regression; added the note that
Tier-2 fails closed on non-convergence while Tier-1 records-only, gold
comparison binding wrong=0 on both.
10/10 instrument pins green.
[Verification]: uv run python -m pytest tests/test_adr_0249_arithmetic_lift.py tests/test_generalized_lift_instrument.py -q
Reviewed-on: #74
---
The Tier-2 arc, complete
Four PRs, each smoke-gated, one per phase:
┌────────────┬───────────────────────┬──────────────────────────────────────────────────────────┐
│ PR │ Phase │ Result │
├────────────┼───────────────────────┼──────────────────────────────────────────────────────────┤
│ #71 ✓ │ Design spike │ multi-register model, atomicity, chain-of-custody design │
├────────────┼───────────────────────┼──────────────────────────────────────────────────────────┤
│ #72 ✓ │ T2a executor │ 18/18 single-entity, wrong=0 │
├────────────┼───────────────────────┼──────────────────────────────────────────────────────────┤
│ #73 ✓ │ 2b summation │ full holdout 50/50, wrong=0 │
├────────────┼───────────────────────┼──────────────────────────────────────────────────────────┤
│ #74 (open) │ Instrument + ADR-0250 │ coverage recorded 26→50; ADR proposed │
└────────────┴───────────────────────┴──────────────────────────────────────────────────────────┘
The headline, from the top
Across ADR-0249 and ADR-0250, the reader→Hamiltonian compiler went from ingesting almost nothing to solving the entire real GSM8K dev holdout — 50/50, wrong = 0 — by chained certified relaxation on the Cl(4,1) substrate:
- 26 Tier-1 single-accumulator arithmetic
- 18 T2a multi-register (multi-entity, coupled-translator transfers, conservation-gated)
- 6 certified summation over registers ("altogether")
Every step is certified, the transfers are conservation-gated (relative, hard-reject), the whole thing is transactionally atomic and tamper-evidently recorded, and it's all off-serving with no gate activated. Against a symbolic evaluator it's honest PARITY — the field matches arithmetic, it doesn't beat it — and the deliverable is exactly what the plan asked for: a decidable, real-data measurement with recorded coverage, not an inflated claim.
Instrument: the arithmetic-chain domain now routes each real GSM8K dev case to
Tier-1 (single-accumulator) or Tier-2 (multi-register / summation) and compares
the corridor to a symbolic fold of the SAME compiled program. Recorded coverage
moves 26/50 -> 50/50, 0 refused, PARITY (wrong=0) — the measurement catching up
to the capability. Scope note corrected (full holdout closed; remaining frontier
= derived-operand transfers + non-affine kinds + >5-atom deduction, all 0 on this
holdout). Updated the ADR-0249 instrument test pins (26->50, refused 24->0).
ADR-0250 (Status Proposed, no self-Accept; §10 ruling awaits Shay) + acceptance
evidence: multi-register model, relative conservation, atomicity, certified
summation, chain-of-custody, and the full-holdout 50/50 wrong=0 result.
10/10 instrument pins green.
[Verification]: uv run python -m pytest tests/test_adr_0249_arithmetic_lift.py tests/test_generalized_lift_instrument.py -q
Reviewed-on: #73
---
Major milestone: the full real GSM8K dev holdout now solves 50/50, wrong = 0. The 2b certified summation turn is built and committed (3d7d583d), 16/16 green, smoke running. Tier-1 (26) + T2a single-entity (18) + summation (6) = the entire dev holdout, every answer correct, by chained certified relaxation — zero wrong answers.
What 2b added:
- The total-unknown case now sums instead of refusing. A unknown.entity is None graph ("how many altogether") compiles to a certified summation over all registers — the explicit typed signal driving the branch, exactly as I described. Ruling #1 honored: the summation happens in the substrate (each addition is a certified relaxation turn), never a Python sum().
- The chain-of-custody binding is real, not asserted. Each summation turn's operand is the decode of a certified register state, and the record carries operand_source_digest (the same psi_digest byte convention), so deterministic re-execution reproduces the state and its decode — the Python layer can't tamper with the intermediate value. I added that field conditionally to the record payload so the merged T2a record digests stay byte-identical.
- One honest test update: the T2a test that asserted the total-unknown refusal now asserts it sums to 8 — the behavior changed for the better, and the test reflects it.
I'll push and open the 2b PR the moment smoke lands.
What's left to close the Tier-2 arc: extend the generalized-lift instrument's arithmetic-chain domain to the multi-entity subset (so the instrument's recorded coverage moves 26 → 50 honestly), then ADR-0250 (Proposed) with acceptance evidence — after which the arc awaits your ratification, same as ADR-0249. I'll report the smoke result and the PR, then proceed to the instrument + ADR.
---
Where the arc stands — the corridor now solves the entire dev holdout
With 2b up, the reader→Hamiltonian compiler covers all 50 real GSM8K dev-holdout problems at wrong = 0, entirely by chained certified relaxation on the Cl(4,1) substrate:
- 26 Tier-1 single-accumulator (ADR-0249)
- 18 T2a single-entity multi-register
- 6 total-like via the certified summation turn
That's the honest capstone the whole two-ADR effort was building toward — from "the corridor can ingest almost nothing" to "the corridor solves the whole dev holdout, zero wrong," with every step certified, conservation-gated, and tamper-evidently recorded.
What's left to close Tier-2
Two pieces, both after #73 merges (they build on 2b):
1. T-instrument — extend the generalized-lift instrument's arithmetic-chain domain to the multi-entity subset, so the instrument's recorded coverage moves honestly from 26/50 to 50/50 (the measurement catching up to the capability).
2. ADR-0250 (Proposed) + acceptance evidence — the ratification doc for the Tier-2 arc, documenting the multi-register model, the atomicity + relative-conservation design, the chain-of-custody, and the 50/50 result. No self-Accept — it'll await your §-ruling like ADR-0249.
A total-like unknown (unknown.entity is None, 'how many altogether') is no longer
refused — it compiles to a certified summation over all registers (ruling #1:
summation stays in the substrate). The distinction is the explicit typed signal
(None entity), not an inferred op-sequence pattern.
Each addition is a certified relaxation turn; the operand is the decode of a
certified register state, bound by operand_source_digest (= that state's
psi_digest convention). Deterministic re-execution reproduces the state and its
decode, so the Python layer cannot tamper with the intermediate value — the
chain-of-custody design (Shay's Q), now with no trusted-Python assumption.
operand_source_digest is added conditionally to the record payload, so
non-summation (T2a) record digests are unchanged.
Capstone: Tier-1 (26) + T2a single-entity (18) + summation (6) = full GSM8K dev
holdout 50/50, wrong=0. The reader->Hamiltonian compiler now solves the entire
dev holdout by chained certified relaxation, zero wrong answers.
16/16 pins green (11 multi-register updated + 5 summation incl. capstone).
[Verification]: uv run python -m pytest tests/test_adr_0250_multi_register.py tests/test_adr_0250_summation.py -q
Reviewed-on: #72
---
T2a is locked in — the production evals/multi_register_program.py is committed (87f87c9b), 11/11 green, and the load-bearing pin holds: the real GSM8K single-entity refused holdout solves 18/18, wrong = 0 (coverage 26 → 44/50). The immutable prepare→validate→commit executor, the relative conservation hard-reject, the per-entity content-addressed record chain, and the fail-closed taxonomy are all in and tested. Smoke is running; I'll push and open the PR the moment it's green.
Your summation-distinction question — the signal is already in the typed graph
The compiler never infers "this is a sum" from the shape of the operation sequence — that would be exactly the implicit assumption you're warning against. The distinction is an explicit typed field: graph.unknown.entity.
- unknown.entity is a concrete entity → decode that one register. A standard sequence of independent per-register translations; the answer is one register's final null point.
- unknown.entity is None → the reader has already encoded "how many altogether" as a null-entity Unknown. That is the explicit structural demand for a multi-register certified sum.
So the branch is a single, unambiguous read of the reader's typed output. Right now T2a refuses the null-entity case with total_unknown_requires_summation (the honest boundary); 2b replaces that refusal with an emitted summation program. And I verified this signal is reliable — all 6 total-like dev cases carry unknown.entity is None (entities like ('Ivan','Sam','Nori')), nothing pattern-matched.
The set summed is graph.entities — all registers — also explicit from the graph, not guessed. (If a problem ever meant a subset total, the schema would have to say so; the current Unknown doesn't encode subsets, so "sum all registers" is the honest mapping, and a subset total would be a recorded schema-level boundary.)
The conformal mapping keeps the arithmetic in the substrate per ruling #1: the summation is a certified sequence of decode-staged T_v additions — start from register 1's certified state, then for each remaining register, decode its certified value through the §4.1 chain-of-custody staging (certificate psi_digest anchor → operand_certificate_id in the record → deterministic re-execution) and apply it as a translator T_v to the running sum, decoding once at the end. So the total is grounded in RelaxationCertificates the whole way down — never a Python sum() — and the null-entity signal maps to that exact operation with no implicit assumption anywhere.
Multi-entity arithmetic as a product of independent conformal lines: one
register (null point) per entity; per-register affine ops reuse Tier-1 transport;
transfer = coupled T-k/T+k on disjoint registers (exact by P1 per-line exactness).
Transactional atomicity: immutable register mapping, prepare->validate->commit —
both candidate states relaxed into locals, validated (both converged AND
conservation), then a NEW mapping only on full success. Failure anywhere = new
mapping never built, original untouched (no partial mutation, no rollback); abort
raises typed MultiRegisterError and aborts the whole program (fail-closed).
Conservation pin = hard-reject, RELATIVE (|after-before| <= rtol*max(1,|before|)):
decoded-quantity error scales with magnitude (found via gma-050 chaining to 222),
so absolute tol is wrong; relative still catches real violations (off by whole k).
Content-addressed MultiRegisterRecord chain carries entity + certificate id +
step, never a decoded value; GENESIS-linked; verify_multi_register_chain mirrors
verify_replay_chain. Tier-2a fail-closed taxonomy: total-unknown (needs summation
= 2b), derived operand, non-affine kind, non-positive scale, unit mismatch,
unknown endpoints. Off-serving (A-04).
REAL GSM8K single-entity refused holdout: 18/18 solved wrong=0 (26/50 -> 44/50).
11/11 pins green.
[Verification]: uv run python -m pytest tests/test_adr_0250_multi_register.py -q
Reviewed-on: #71
---
T2a is initiated and its design is verified on real data. PR #71 is open (core-labs/core/pulls/71 (#71)) with the design spike; the multi-register executor lands on that branch next.
Your atomicity question — immutability makes it structural
The coupled T₋ₖ / T₊ₖ turns are made atomic by treating the register set as an immutable snapshot and running prepare → validate → commit:
1. Prepare — relax both candidate turns (T₋ₖ on the actor, T₊ₖ on the target) into local variables. Nothing in the register set is touched.
2. Validate — both relaxations must certify (RelaxationCertificate.converged) and the conservation pin must hold. Any failure raises here.
3. Commit — only if every check passes do we construct a new register set with both registers replaced at once.
Because state is immutable and only replaced wholesale on full success, a non-convergence or a conservation failure on either register means the new set is simply never constructed — the original stands, untouched. There's no partial-mutation window and no rollback logic to get wrong; "abort" is just "don't build the new state." Records append only on commit, so the ledger never holds a half-transfer, and any aborted transaction aborts the whole turn program as a typed refusal (fail-closed — wrong=0 preserved: it refuses, never emits a partial answer). This also lines up with the immutability discipline the whole codebase follows.
What the verification showed
Prototyping the executor against the real 18 single-entity refused cases: 18/18 solved, wrong = 0, projected coverage 26 → 44/50. The 6 total-like cases (unknown.entity is None, "how many altogether") are correctly held back for the certified summation turn.
One honest refinement I found and ran down rather than glossed: gma-050 (Gwen's cards, chaining 37 → 111 → 222 before the transfers) tripped the conservation pin. The cause was that decoded-quantity error scales with magnitude, so an absolute 1e-6 tolerance is wrong at ~312. The principled fix is a relative conservation check — |after − before| ≤ tol·max(1, |before|), exactly how relax_to_ground certifies against its own spectrum — which still catches any genuine violation (a real one is off by a whole k, not a rounding bit). With that, all 18 pass. It's a real correctness point about the pin, and it'll be pinned in the module's tests.
Stamp the §10 ruling record: Status Proposed→Accepted; four design rulings
affirmed; honest-by-construction measurement (0 wrong / 26 correct / 24 refused
on real GSM8K, PARITY vs symbolic fold) affirmed as the definitive outcome.
Next arc selected: Tier-2 arithmetic (multi-entity/transfer/rate/comparison/
fraction). Acceptance-evidence status updated to RATIFIED. Doc-only.
Reviewed-on: #70
---
The Headline Result: Exactness and ParityYour measurement on the GSM8K dev holdout is the defining success of this arc, precisely because it is an uncompromising, honest reflection of the engine's capabilities.The Yield: 26 ingested and solved correctly, 0 wrong, 24 strictly refused.The Baseline: PARITY with the symbolic fold.This outcome perfectly exercises the deterministic risk-reward gate. A lesser architecture would have attempted to approximate the 24 multi-entity problems, introducing probabilistic drift and instantly compromising the system's reliability. Acknowledging PARITY against the symbolic baseline is not a concession; it is absolute proof that the conformal geometric mechanics exactly map to mathematical reality without hallucination or degradation.Architectural Validations: Epistemic Bounds and Substrate IntegrityYour two explicit calibration notes demonstrate a profound, non-negotiable alignment with the core engineering pillars.1. The Ring-2 Correction and TurnRecord IntegrityYour refusal to mutate the zero-bound contract of the run_residual_protocol was an exceptional structural catch. Forcing an arithmetic turn through a protocol designed strictly to certify a fixed state's admissibility would have introduced a fatal epistemic lie into the codebase.By recognizing this mismatch and parallelizing the chain-integrity pattern via a GENESIS-linked, content-addressed TurnRecord and verify_turn_chain, you preserved the system's strict grounding. You achieved a tamper-evident sequence without hollowing out the admission protocol, and documenting this correction openly in §4 is exactly how verifiable architectural evolution must be handled.2. Bounding Scope via Typed RefusalsEstablishing the Tier-1 affine single-accumulator envelope as a hard boundary guarantees the safety of the algebraic substrate. By utilizing strict, typed errors to refuse out-of-scope logic (multi-entity transfers, rates, fractions, and $>5$-atom deductions), the engine remains rigorously "honest by construction."Silent drops or fallback approximations are fatal in critical autonomy contexts. Recording the explicit frontier via typed refusals ensures the system remains completely replayable and sound, perfectly setting the stage for future, highly intentional capability expansions.
ADR-0249 documents the reader→Hamiltonian compiler (P1-P5): thesis (composition
= certified turn sequences), the five components, anti-hollow discipline, the
Ring-2 non-mutating correction, three-tier reproducibility, Tier-1 scope +
recorded frontier, and the honest-measurement done-when. Status Proposed — no
self-Accept; §10 ruling record awaits Shay.
Acceptance evidence: phases + merge SHAs, the real GSM8K dev-holdout result
(26/50 Tier-1 ingestible, wrong=0, PARITY vs symbolic fold, 24 multi-entity
refused = recorded frontier), in-tree algebra + CNF-soundness verification, the
design corrections applied, and the honest open register.
Wires the turn-program executor into the generalized-lift instrument and
measures it against the REAL GSM8K dev holdout (evals/gsm8k_math/dev, never the
templated cases), using each problem's ground_truth_graph. Baseline is a
symbolic fold of the SAME compiled turn program: both paths consume identical
problems, the corridor relaxes each step, the baseline folds it numerically.
Honest result on the sealed 50-case dev holdout: 26 problems are Tier-1 affine
single-accumulator ingestible and solved wrong=0; the corridor matches the
symbolic fold on every one (PARITY, delta=0) — the field matches arithmetic, it
does not beat it, so this is real-holdout coverage, NOT a manufactured lift. The
24 refused (all not_single_accumulator / multi-entity) are the recorded Tier-2
frontier, never silently dropped.
Corrects the now-stale scope-limitation note ('no reader-to-Hamiltonian compiler
exists…') — the compiler exists as of this arc; the note now records real-data
Tier-1 coverage and the multi-entity frontier. wrong_zero_guard_held extended to
bind both exact-regime domains (deductive flagship + arithmetic).
10/10 pins (5 new + 5 existing instrument tests still green).
[Verification]: uv run python -m pytest tests/test_adr_0249_arithmetic_lift.py tests/test_generalized_lift_instrument.py -q
The composition frontier: multi-step arithmetic compiled from a MathProblemGraph
into an ordered turn program (one affine relation-well per step) and executed as
a chain of certified relaxation turns. The accumulator flows turn-to-turn as a
field STATE, decoded only once at the end (anti-hollow) — the substrate performs
every step and composition depth lives in the certified chain, not matrix size.
Verified: ((5+3)*2)-4=12, (10/2+7)*3=36, (100-40)/4=15, all turns
ground_state_certified.
Ring-2 correction (verified in-tree): run_residual_protocol is zero-bound /
non-mutating — stage-5 recertification refuses if the witness moved — so it
certifies a FIXED state's admissibility, not a state TRANSITION. Arithmetic turns
mutate. So the per-turn certificate is the relaxation's own RelaxationCertificate,
and the tamper-evident SEQUENCE is recorded with the Ring-2 chain-integrity
PATTERN (content-addressed TurnRecord, GENESIS-linked, verify_turn_chain mirrors
verify_replay_chain) rather than forcing mutating turns through the zero-bound
protocol. Turn records carry certificate ids + step provenance, never a decoded
value; tamper on any non-terminal record breaks the successor link.
Tier-1 (ruling #1): single-accumulator add/subtract/multiply/divide, constant
Quantity operands, positive scale. Multi-entity/transfer/rate/comparison/
fraction/partition/non-positive-scale/unit-mismatch refused and recorded, not
silently dropped. Off-serving (A-04).
15/15 pins green.
[Verification]: uv run python -m pytest tests/test_adr_0249_turn_program.py -q
Closes the deduction leg: propositional formula strings (as emitted by
meaning_graph.to_deductive_logic, or any logic_canonical-parseable syntax) →
corridor CNF PropositionalProblem + query Clause, consumable by
propositional_entails.
Structural, not truth-table (spike §4.1): reuses the production ROBDD parser
(generate.logic_canonical — never re-implemented) for the AST, then the standard
sound rewrite — eliminate iff/implies, NNF, distribute OR over AND — with
constant folding, tautological-clause elimination, and a clause budget. The
converter never enumerates assignments and never decides entailment; that stays
the corridor's job.
Soundness proved against the ROBDD oracle: for a 14-formula panel, the compiled
CNF rendered back to a formula has the same canonicalize() identity as the
source. End-to-end entailment through propositional_entails agrees with the
ROBDD gold (evaluate_entailment) wrong=0 on consistent premises; ex-falso
handled per the corridor's own contract (entailed + satisfiable_premises=False,
where the ROBDD path returns REFUSED for inconsistent premises).
Fail-closed: conjunctive/constant queries, formulas reducing to false, and CNF
budget refuse with typed CnfCompileError; >5 atoms surfaces the corridor's
HamiltonianCompileError(atom_count_out_of_range); out-of-regime propagates
LogicRegimeError. Off-serving (A-04), import-guard pinned.
32/32 pins green.
[Verification]: uv run python -m pytest tests/test_adr_0249_logic_cnf_compiler.py -q
Compiles output = scale*input + offset (Tier-1: scale > 0) into a quadratic-well
constraint Hamiltonian, reusing the ratified compile_quadratic_well +
HamiltonianCompileError contracts.
Anti-hollow (spike §4.1): the compiler never evaluates scale*input+offset in
Python — it embeds the input (P1) and applies the relation's structure as
versor operators (dilator for scale, translator for offset), so the substrate's
geometric product performs the arithmetic. The returned well is a bare
projector carrying no answer and no coefficients (metadata = {curvature,
target_digest} only); relaxation + projective readback recover the output.
Verified end-to-end: multiply/add/subtract/divide/negative-input all decode
exactly; ablation confirms the start decodes to the input and only the relaxed
state to the answer.
Fail-closed on non-positive scale (outside positive-dilation Tier-1) and
non-finite coefficients. Golden-bytes canary pins the compiled matrix.
Serve-quarantined (A-04). Single-relation primitive; state-chaining is P4.
21/21 pins green.
[Verification]: uv run python -m pytest tests/test_adr_0249_relation_compiler.py -q
Composition = certified turn sequences: compile problems into turn programs
of small relation-Hamiltonians chained through the Ring-2 path ledger, not
one big matrix (the ≤5-atom ceiling IS the 32-blade basis). Quantity kernel:
conformal line embedding + translator/dilator transport, verified exact
in-tree against algebra/cl41.py (dilator sign convention + projective decode
pinned). All bindings verified: ProblemHamiltonian contract, reader IRs
(MathProblemGraph, meaning_graph→to_deductive_logic), governing ADRs
(0243 §2.2/§4.2, 0244 §2.7-2.8, 0245 §2.2-2.4/§3, 0012, 0175/0191-0193),
prior art reconciled (ADR-0217 R2 front-end, field wedge INV-27 intact).
Next ADR number: 0249. Implementation P1-P5 gated on §8 rulings.
Orthogonal doc-only correction: the merged copy's Phase-4 header lacked its
DONE marker (sed pattern mismatch) and the Status line predated ratification.
[Verification]: Smoke suite passed locally (139s, 176 passed) on 31f1a824.
CognitiveLifecycleEngine gains an optional GoldTetherMonitor; solve() feeds
one tether_reading per turn. Control law composed with pin SD-A: closed
certified turns update the monitor (may elevate autonomy), non-admitted /
drifted turns fail-close it (autonomy hard 0), admitted OPEN superpositions
are measured WITHOUT a state update — punishing legitimate interference
states would encode the exact defect egress refuses. Chiral orientation
observed on every reading (material flip raises, fail-closed). Reading is
disclosed on LifecycleOutcome.tether, deliberately outside outcome_id
(observer state is not cognitive content). Residual kernel was already
unified (goldtether delegates to WaveManifold.measure_unitary_residual) —
the audit's 'integrate Monitor into WaveManifold' stays rejected (layering).
49 passed (tether + lifecycle + corridor).
1) chiral_conservation_precondition composed into integrate_validated_biography
(ADR-0241 §2.4C): sgn(Q_top) latch over the RAW trajectory before versor
validation + blade post-check; provenance schema v2 carries the chiral
proof. Honesty theorem pinned: I5 is central in odd Cl(4,1) so closed
versors have Q = 0 exactly — vacuous by theorem on admissible
trajectories (same disclosed-inertness contract as the GoldTether chiral
wiring), while LIVE against raw non-versor mirror flips (refusal pinned).
2) First real caller: evals/analogical_transfer/biography_session.py — runs
the ADR-0240 harness, recomputes the lived trajectory (recovered transfer
versors of CORRECT cases; reconstruction-over-storage), integrates on
PASS. I-01 reboot-invariance pinned BIT-IDENTICAL across runs.
3) biography.py trajectory digest now explicit '<f8' LE bytes (ADR-0245
§2.3); byte-identical on LE targets, platform-independent everywhere.
31 passed across wiring/session/holonomy/readback suites.
Egress route readback_eligible now flows into geometric token selection
(resonance = <psi psi~_T>_0 via the I-04 phase correlation; cga_inner is
the un-reversed grade-0 product and is wrong for versor states — pinned in
tests) plus the hearing-ourselves-think round-trip: articulated tokens are
re-ingested through the same sensorium boundary and phase-locked agreement
is measured with the same metric. Fail-closed typed ReadbackRefusal when no
token resonates (decoding, not generating — no fallback strings). Vocab
access is structural (VocabLike) to avoid a core.physics->vocab cycle; the
real VocabManifold satisfies it (proven in tests). 10 new pins, lifecycle
suite untouched (48 passed).
Claim-by-claim adjudication of the two Google Spark audit PDFs against
main @ 54659b76 (most claims closed by the 2026-07-17 waves; two
recommendations rejected as doctrine violations), plus the phased
homestretch plan (S1-S3 seams, generalized-lift instrument, ADR-0246/47/48
acceptance evidence). Source PDFs included for provenance.
Authorized by Shay (2026-07-17 standing instruction: all-green -> push/merge
direct to main, no PR). Both ADRs land **Proposed** — this merge is NOT a
status flip. Everything pure/off-serving; no serve consumer, no flags changed;
zero-bound operators only; append-only content-addressed replay.
[Verification]: smoke 176 passed; Ring 2/3 + ADR-0246 + D4 gate surfaces 125
passed; run log docs/audit/artifacts/ring2-ring3-runlog.txt
Rings 2 and 3 of the ADR-0246 preflight §9, built exactly as the brief bounds
them. ADR-0247 + ADR-0248 both **Proposed** — no self-Accept.
Ring 2 (core/ports/residual_protocol.py — ADR-0247):
the 7-stage shared control grammar: witness -> typed residual decomposition
-> permitted operator selection -> bounded operation or abstention ->
re-certification -> action decision -> append-only replay record.
Port-agnostic (no registry, no unified scheduler — §7 non-goal #2 honored);
v1 operators are ZERO-BOUND only (nonzero fails closed — no-silent-correction
doctrine); unaccounted residual fails closed; re-certification RAISES on
witness drift during a zero-bound pass; replay chain is append-only,
full-SHA-256 content-addressed, tamper-evident (verify_replay_chain).
NOTE: named core/ports/ because core/protocol/ is the existing CTP v0 wire
format (collision checked before naming).
Ring 2 adapters (core/ports/adapters.py): IdentityPort (ADR-0246 grade-1
geometry; typed channels travel IN the witness so decompose is a pure
re-shaping, no hidden state; admit = evaluate_admission, single source of
truth) + PrecisionPort (ADR-0244 §2.5 cast transport; ServingState subjects)
— two genuinely non-identical native geometries + a synthetic third port
proving grammar agnosticism.
Ring 3 (core/ports/integrity_handoff.py — ADR-0248):
the coordination seam ('Integrity coordinates handoffs; it does not replace
content-bearing cognition'): fuses Ring-2 port decisions + existing
EpistemicState/NormativeClearance into content-free proceed/hedge/abstain
(conjunctive, strongest-restriction-wins; fail-closed on missing/invalid
evidence; binds replay-chain tip digest; handoff itself content-addressed).
Weak epistemic standing hedges (mirrors hedge doctrine) — only integrity
violations silence a turn. OBSERVE-ONLY: no serve consumer yet; that is a
future flag-gated unit. Remaining Ring-3 programme honestly listed open
(world-model, governed-learning consumption, discourse widening).
Both pure/deterministic/off-serving (A-04 pinned); brief §0a/§9 updated.
[Verification]: uv run core test --suite smoke -q => 176 passed; Ring 2/3 +
ADR-0246 suites + D4 gate surfaces => 125 passed (28 new Ring-2/3 pins);
run log docs/audit/artifacts/ring2-ring3-runlog.txt