From 353a52b833f15b938c5cf770d4323fc73d43a1cc Mon Sep 17 00:00:00 2001 From: Shay Date: Fri, 24 Jul 2026 14:46:55 -0700 Subject: [PATCH] =?UTF-8?q?docs(math-reader):=20Phase=204.1=20measured=20s?= =?UTF-8?q?tatus=20=E2=80=94=20seeding=20injection=20is=20built,=20convert?= =?UTF-8?q?s=200=20(null=20result)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Re-measured the frozen holdout_dev/v1 500: correct=5 wrong=0 refused=495 — exactly the tune-3 + measure-2 increment 1 recorded, no drift, wrong=0 held. Seeding-sentence injection is already in the tree (NEUTRAL_COUNT_VERBS in math_roundtrip.py, consumed by the seed matcher and the comparison anchor); its measured yield when it shipped was 0 conversions, verified four ways. Extending the allowlist further is the move ADR-0251 and the overfit inventory forbid — per-case pattern growth previously produced lift that committed WRONG answers on the real exam. Recommendation: mark Phase 4.1 complete-with-null-yield and re-point Phase 4 at the reader arc's own live recommendation — increment-2 CASE-FIRST on the already-identified closest cases (0000/0001/0148/0082). --- ...math-reader-phase-4-1-status-2026-07-24.md | 85 +++++++++++++++++++ 1 file changed, 85 insertions(+) create mode 100644 docs/research/math-reader-phase-4-1-status-2026-07-24.md diff --git a/docs/research/math-reader-phase-4-1-status-2026-07-24.md b/docs/research/math-reader-phase-4-1-status-2026-07-24.md new file mode 100644 index 00000000..17355166 --- /dev/null +++ b/docs/research/math-reader-phase-4-1-status-2026-07-24.md @@ -0,0 +1,85 @@ +# Math reader Phase 4.1 — measured status: the increment is already built, and it converts nothing + +**Date:** 2026-07-24 · **Arc:** generalization Phase 4.1 (Tier O, O4) · **Result:** null, by measurement + +## What was asked + +The generalization plan §2/Phase 4.1 and the Opus brief item O4 both name +"seeding-sentence injection (standing ruling #77 — the 75% 'no injection' +wall)" as the next math move, with wrong=0 on the full 500 as the gate and +parse-rate gain as the goal. + +## What is actually in the tree + +Seeding-sentence injection **is built and merged**. The reader arc's increment +1 shipped it as the #78 positive-polarity substrate: + +- `generate/math_roundtrip.py::NEUTRAL_COUNT_VERBS` — the curated + production/possession/acquisition allowlist, +- consumed by the seed matcher's `inject_discrete_count_statement` path and by + `generate/math_candidate_parser.py::_COMPARISON_ANCHOR_VERBS`, +- depletion/transfer verbs refuse by positive determination (an allowlist, not + a blocklist — the fail-open blocklist was reverted), +- documented in `docs/research/increment-1-band-plan-2026-07-19.md` §BUILD + OUTCOME. + +Its measured yield when it shipped: **PARSED unchanged — tune 3, measure 2, no +conversion.** The scoped band `{seed + forward-comparison + q:simple/summation}` +converted **0** tune cases, measured four independent ways (yield harness, band +map, band oracle, capability histogram). + +## Re-measured today, on the frozen split + +``` +uv run python evals/gsm8k_math/holdout_dev/v1/runner.py +holdout_dev: correct=5 wrong=0 refused=495 (n=500) +``` + +Five correct is exactly the tune-3 + measure-2 the increment-1 report recorded. +**No drift in five days of unrelated arc work, and wrong=0 still holds.** The +tune/measure split (`evals/gsm8k_math/holdout_dev/v1/split.py`) is frozen and +was not touched. + +## Why the honest move is to stop rather than extend + +The arc's own evidence says the wall is not where 4.1 aims: + +1. **The ~30-band bet was falsified.** In-band cases strand on capabilities + *outside* the band — 25 need multi-compound, 17 need compare-additive. + Adding more seed verbs cannot convert a case that also needs rate, + currency, and unit conversion in the same sentence. +2. **The taxonomy is not at bedrock.** The "smallest tractable set" + `{seed, q:simple}` dissolved on inspection: every one of its four cases + really needs rate / compare / currency / copula. Real GSM8K statements are + individually multi-capability — the conjunction recurs *within* a sentence, + so a capability list keeps fragmenting rather than converging. +3. **Growing the allowlist is the forbidden move.** ADR-0251 and + `docs/research/reader-arc-overfit-inventory-2026-07-19.md` exist because + per-case pattern growth previously produced overfit "lift" that committed + *wrong* answers on the real exam. An extension whose only evidence is "it + parses more tune cases" is the exact shape of that failure. + +## The live next move (unchanged, and it is not 4.1) + +The reader arc's own recommendation, carried to Shay with the increment-1 +foundations, is **case-first, not capability-first**: pick 2–3 specific closest +tune cases, hand-enumerate their *exact* end-to-end needs however many +capabilities that is, build precisely that, convert them, measure. The closest +cases are already identified — `0000`, `0001`, `0148`, `0082` — all needing +complex compare forms (chained / mass-noun entities / "than" + aggregate +reference) plus summation or difference. + +`multi-compound` is the single highest-leverage capability (on the critical +path for 70/106 = 66% of tractable needed-sets) but is insufficient alone. +`q:complex` remains the largest intractable wall (~35+ cases). + +## Recommendation for the plan + +Phase 4.1 should be marked **complete-with-null-yield** rather than pending, +and Phase 4 re-pointed at the increment-2 case-first work. The measurement +above is the evidence; nothing in this session's arc work moved it either way, +which is itself the useful signal — the math lane is genuinely orthogonal to +the reading/curriculum work, exactly as the plan's §3 dependency table assumed. + +Phase 4.2 (compare unblock) and 4.3 (q:complex decomposition study) are +untouched by this and remain queued behind the case-first increment.