From 4a5d0f460d562d722a18d41a6eeaeb5291ade3c1 Mon Sep 17 00:00:00 2001 From: Shay Date: Sat, 18 Jul 2026 17:25:26 -0700 Subject: [PATCH] docs(reader-arc): compare increment per-layer attrition funnel (diagnostic) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Instrument-first (PR #76 ruling): before touching the parser, measure where the 61 compare-marked TUNE cases die. Funnel: 57/61 die pre-enumeration; of those, **45 at the injection layer** ('recognizer matched but produced no injection'), 8 'no admissible candidate for statement', 3 'expected exactly one question sentence'; 3 solve today. Finding: the leverage is INJECTION (45/61), not the candidate regexes (which fire on clean clauses but the real multi-clause sentences never reach them) nor segmentation writ large. Injection is an EMITTER, so the sanctioned fix is making it inject compare candidates the recognizer already matches — never loosening the round-trip gate. Fix order by measured mass: injection (45) -> statement-admissibility (8) -> question-sentence (3). Tune-only development; single measure run at the end with refusal-by-reason + delta. --- .../compare-increment-funnel-2026-07-18.md | 51 +++++++++++++++++++ 1 file changed, 51 insertions(+) create mode 100644 docs/research/compare-increment-funnel-2026-07-18.md diff --git a/docs/research/compare-increment-funnel-2026-07-18.md b/docs/research/compare-increment-funnel-2026-07-18.md new file mode 100644 index 00000000..f72e615a --- /dev/null +++ b/docs/research/compare-increment-funnel-2026-07-18.md @@ -0,0 +1,51 @@ +# compare_multiplicative Increment — Per-Layer Attrition Funnel (diagnostic artifact) + +**Status**: DIAGNOSTIC — instrument-first, decides the fix order (PR #76 ruling) +**Date**: 2026-07-18 +**Scope**: the 61 compare-multiplicative-marked cases in the **tune** split of official `holdout_dev/v1` +(measure split untouched). Signals are the reader's own: `branches_enumerated`, `refusal_reason`, +`selected_graph`, then the corridor compiler. + +--- + +## The funnel + +Reader pipeline stages: recognizer-match → **injection** → enumeration → admissibility → graph → solve. + +| Layer | Cases (of 61) | +| :--- | ---: | +| Died **pre-enumeration** | **57** | +| Enumerated | 4 | +| …enumerated but no graph | 1 | +| …graph has a compare op | 3 | +| **Solved today** | **3** | + +Pre-enumeration refusal reasons (the 57): + +| Reason | Cases | +| :--- | ---: | +| **"recognizer matched but produced no injection"** | **45** | +| "no admissible candidate for statement" | 8 | +| "expected exactly one question sentence" | 3 | +| "no admissible candidate for question" | 1 | + +## The finding + +**The leverage is the injection layer: 45 / 61.** The recognizer already *matches* these compare +statements (e.g. 0000: "Aria has twice as many … as Emily, who has twice the number … as Spencer"); +the **injection** step then produces no candidate — the multi-clause / compare shapes don't inject. +This is where the mass is — not the candidate regexes (they fire on clean clauses but the real +sentences never reach them), and not segmentation writ large. + +## Consequence for the fix (emitter-fix, not gate-loosening) + +Injection is an **emitter** (it produces candidates), not a validator — so making it inject +compare candidates for statements the recognizer already matches is the sanctioned kind of change +(fix emitters; never loosen the round-trip gate). The remaining piles are secondary and sequenced +after: "no admissible candidate for statement" (8), the one-question-sentence constraint (3). + +Fix order, by measured mass: **injection (45) → statement-admissibility (8) → question-sentence (3)**. +Each fix is developed against tune only, to `wrong=0` on tune; the measure split gets its single run +at the end. Refusal-by-reason on the measure split (frequency / nested / anaphora / temporal / +inverse-undeterminable) is reported with the delta as the fail-closed evidence + the practice-lane +curriculum.