From b87cd46cf759740b10fc44a18127ef14087399cb Mon Sep 17 00:00:00 2001 From: Shay Date: Sat, 18 Jul 2026 15:07:29 -0700 Subject: [PATCH] chore(adr-0250): ratify Accepted + fold in PR #74 review findings MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Ratification: ADR-0250 Status Proposed->Accepted; §10 ruling record stamped (Shay, citing the acceptance-evidence pack as-is; full-holdout 50/50 wrong=0 affirmed; next arc = seal dev-holdout-2, dev-1 pinned as regression floor). Review findings (Shay, PR #74, non-blocking): 1. Instrument module docstring updated from 'Three deterministic domains' + 'GSM8K NOT ingestible / no compiler' to the four-domain post-0250 reality (arithmetic-chain solves the full dev holdout; frontier surfaced in scope). 2. Restored per-case refusal reasons: _corridor_and_baseline now raises MultiRegisterError (not bare None); the domain loop records refusal.reason and distinguishes graph_parse_failed. 'Recorded, never silently dropped' stays literal (0 rows at refused=0, but the path is honest). 3. Tightened the Tier-1 try to wrap compile_turn_program only (execute outside), so a future typed execute raise surfaces as a Tier-1 failure instead of silently rerouting into Tier-2 and masking a regression; added the note that Tier-2 fails closed on non-convergence while Tier-1 records-only, gold comparison binding wrong=0 on both. 10/10 instrument pins green. [Verification]: uv run python -m pytest tests/test_adr_0249_arithmetic_lift.py tests/test_generalized_lift_instrument.py -q --- .../ADR-0250-tier2-multi-entity-arithmetic.md | 13 +++-- evals/generalized_lift_instrument.py | 55 ++++++++++++------- 2 files changed, 43 insertions(+), 25 deletions(-) diff --git a/docs/adr/ADR-0250-tier2-multi-entity-arithmetic.md b/docs/adr/ADR-0250-tier2-multi-entity-arithmetic.md index 55ca327a..d4131c17 100644 --- a/docs/adr/ADR-0250-tier2-multi-entity-arithmetic.md +++ b/docs/adr/ADR-0250-tier2-multi-entity-arithmetic.md @@ -1,6 +1,6 @@ # ADR-0250: Tier-2 Multi-Entity Arithmetic — the Multi-Register Extension -**Status**: **Proposed** +**Status**: **Accepted** **Date**: 2026-07-18 **Authors**: Joshua Shay + multi-model R&D (implemented Fable 5) **Depends on**: ADR-0249 (reader→Hamiltonian compiler — quantity kernel, relation compiler, turn programs), ADR-0243/0244/0245 (relaxation, content-addressing, f64) @@ -90,6 +90,11 @@ Local-first CI: smoke green in-worktree at every phase. ## 10. Ruling record (Shay) -_Awaiting ratification._ The design rulings (multi-register product-of-lines; relative -conservation hard-reject; prepare→validate→commit atomicity; certified summation turn; ship 2b -designed+guarded) are recorded RESOLVED in the spike; acceptance is Shay's alone (no self-Accept). +**RATIFIED 2026-07-18 by Joshua Shay — ADR-0250 ACCEPTED.** The design rulings (multi-register +product-of-lines; relative conservation hard-reject; prepare→validate→commit atomicity; certified +summation turn; ship 2b designed+guarded) are affirmed, citing the acceptance evidence pack +(`docs/handoff/ADR-0250-Acceptance-Evidence.md`) as-is. The full-holdout result — the +reader→Hamiltonian compiler solving the entire real GSM8K dev holdout 50/50 wrong=0 by chained +certified relaxation — is affirmed as the definitive outcome. Next arc: re-arm the measurement +(seal dev-holdout-2, a larger real GSM8K slice where the frontier kinds appear) and let the honest +coverage table pick the next capability tier; dev-1 stays pinned at 50/50 as a regression floor. diff --git a/evals/generalized_lift_instrument.py b/evals/generalized_lift_instrument.py index 4e9c61ae..9997990c 100644 --- a/evals/generalized_lift_instrument.py +++ b/evals/generalized_lift_instrument.py @@ -2,11 +2,11 @@ Instrument-first doctrine (ADR-0190 lesson): this module DECIDES the "meaningful, generalized lift without overfitting" question instead of -narrating it. Three deterministic domains, identical compiled problems for +narrating it. Four deterministic domains, identical compiled problems for both paths, independent gold, and an honest-NULL protocol: if the corridor adds no delta, the report says so. -Domains (corridor v1's honest ingestible surface): +Domains (corridor's honest ingestible surface): * ``propositional`` — entailment on an enumerated case family. Corridor: :func:`propositional_entails` (exact ground-energy verdicts). Baseline: the @@ -22,12 +22,16 @@ Domains (corridor v1's honest ingestible surface): * ``multimodal-completion`` — the sensorium corridor pattern (audio partial → full audio+vision percept), scored on whether articulation names BOTH constituent percepts; baseline articulates the raw partial ingress. +* ``arithmetic-chain`` — real GSM8K dev-holdout problems routed to the + reader→Hamiltonian compiler: Tier-1 single-accumulator (ADR-0249) and Tier-2 + multi-entity / transfer / certified summation (ADR-0250), each vs a symbolic + fold of the SAME compiled program. The full dev holdout solves wrong=0 + (PARITY — the field matches arithmetic, it does not beat it). -Scope limitations are RECORDED, never silently dropped (no-silent-caps): -GSM8K and other natural-language arithmetic are NOT ingestible by corridor -v1 — there is no reader→Hamiltonian compiler beyond the ≤5-atom -propositional and quadratic-well domains. That compiler is the real -composition frontier, and this instrument is the harness waiting for it. +Scope limitations are RECORDED, never silently dropped (no-silent-caps): the +reader→Hamiltonian compiler now solves the full dev holdout; the remaining +frontier (derived-operand transfers, non-affine kinds, >5-atom deduction) sits +at 0 cases on this holdout and is surfaced in the report's ``scope_limitations``. """ from __future__ import annotations @@ -512,19 +516,23 @@ def _symbolic_multi_register(program) -> float: def _corridor_and_baseline(graph): """Route a graph to Tier-1 (single-accumulator) or Tier-2 (multi-register), - returning (corridor_answer, baseline_answer) or None if genuinely un-ingestible. - Both paths consume the identical compiled program; the corridor relaxes, the - baseline folds arithmetically.""" + returning ``(corridor_answer, baseline_answer)``. Both paths consume the + identical compiled program; the corridor relaxes, the baseline folds + arithmetically. Raises ``MultiRegisterError`` when neither tier ingests. + + Only the Tier-1 COMPILE is wrapped for the fallthrough: a future typed raise + from ``execute_turn_program`` must surface as a Tier-1 execution failure, not + be silently rerouted into Tier-2 (which would mask a regression). The tiers + differ on non-convergence — Tier-2 fails closed (raises), Tier-1 records the + iterate — but the gold comparison binds wrong=0 on both regardless.""" try: program = compile_turn_program(graph) - return execute_turn_program(program).answer, _symbolic_fold(program.seed, program.steps) except TurnProgramError: - pass - try: - program = compile_multi_register_program(graph) - except MultiRegisterError: - return None - return execute_multi_register_program(program).answer, _symbolic_multi_register(program) + program = None + if program is not None: + return execute_turn_program(program).answer, _symbolic_fold(program.seed, program.steps) + mr_program = compile_multi_register_program(graph) # raises MultiRegisterError if un-ingestible + return execute_multi_register_program(mr_program).answer, _symbolic_multi_register(mr_program) def run_arithmetic_chain_domain() -> DomainOutcome: @@ -550,14 +558,19 @@ def run_arithmetic_chain_domain() -> DomainOutcome: graph = graph_from_dict(graph_dict) if graph_dict else None except (MathGraphError, KeyError, TypeError, ValueError): graph = None - answers = _corridor_and_baseline(graph) if graph is not None else None - if answers is None: + if graph is None: # reader/graph-parse failure — distinct from a compiler refusal corridor_refused += 1 baseline_refused += 1 - rows.append({"case_id": case_id, "ingested": False}) + rows.append({"case_id": case_id, "ingested": False, "refusal": "graph_parse_failed"}) + continue + try: + corridor_answer, baseline_answer = _corridor_and_baseline(graph) + except MultiRegisterError as refusal: # neither tier ingests — record the reason + corridor_refused += 1 + baseline_refused += 1 + rows.append({"case_id": case_id, "ingested": False, "refusal": refusal.reason}) continue - corridor_answer, baseline_answer = answers gold = None if gold_raw is None else float(gold_raw) corridor_ok = gold is not None and abs(corridor_answer - gold) < 1e-4 baseline_ok = gold is not None and abs(baseline_answer - gold) < 1e-4