core/docs/research/math-phase-4-2-case-first-traces-2026-07-25.md
Claude 71bf04fb44
feat(provenance,teaching): close the lateral gaps the assessment actually found
Squashes the arc's work into one commit; the workflow-file edit it originally
carried is excluded (see the end of this message).

## Lane 1 — Workbench recorded a proved answer as ungrounded

With deduction_serving_enabled ratified ON (ADR-0256), workbench/api.py's live
chat route builds a bare ChatRuntime(), so the deduction composer decides
Workbench turns and stamps grounding_source="deduction" — but
_coerce_grounding_source carried a hand-copied whitelist of the six pre-arc
labels and silently rewrote anything else to "none". The runtime comment
reasoned this was inert because "REPL turns do not flow through Workbench's
CognitivePipelineRecord path". True, and irrelevant: the traffic flows the
other way. Stale since 2026-07-24.

Scope is one field. workbench/api.py:818 prefers TurnEvent.epistemic_state,
which read epistemic_state_needed — honest. So the UNregistered path degraded
honestly while the hand-copied whitelist asserted a falsehood; a second copy of
a closed enum was worse than no copy. Hence registration AND derivation:
GROUNDING_SOURCES exposes the Literal's members, and the coercion reads it.
workbench-ui badges/tokens/snapshot follow; enumCoverage.test.ts forces atomicity.

## Lane 2 — the ratification ceremony

The discovery loop was instrumented but not closed. teaching/ratification.py
turns a reviewed decision into a chain record, a corpus commit, and a receipt.

Its design turns on one observation: _ratified_rows DROPS unadmissible rows
silently — correct when serving, a trap when ratifying, because the file grows,
the commit lands, and the band count does not move. So the ceremony refuses to
call an append a ratification until it has re-read the curriculum through the
real loader and seen the chain arrive; a non-admitted append is rolled back.
Validation is a pre-flight courtesy, admission is the proof.

Arena queue entry and ledger reseal are deliberately NOT performed (bridge rule
1); the receipt names them. Front door: `core proposal-queue ratify`, a sibling
of `review` rather than a flag on it.

## Lane 3 — structural closures

- ADR-0263 gains rule 5: absence policy is DECLARED in CAPABILITY_LEDGERS, not
  passed at the call site. An AST-matched test fails if a serving path passes
  missing_ok again.
- Deductive suite added WHOLE to the pre-push gate: 285 tests in 29s against
  smoke's 216 in 62s, so no coverage trade was needed.
- Smoke/CI parity assertion made bidirectional. It was one-directional, and had
  drifted.
- test_prior_surface_deduction_binding.py pins correction binding on the
  deduction path. The review's diagnosis did NOT reproduce — hash_surface moves
  in lockstep — so it pins what is there. Mutation-checked.
- Domain-keyed ADR index over 312 flat-numbered files, explicitly partial.
- Arc-close brief template, plus this arc's own brief filled in against it.

## Lanes 4 and 5 — two premises falsified by measurement, one of them mine

Math 4.2: baseline reproduced (correct=5 wrong=0 refused=495); all four named
cases traced to one seam with each gap isolated by one-variable probes. Then the
number that changes the recommendation: the gap blocking case 0000 affects 1
case in 500, the 'than' gap blocking 0001 affects 2. ADR-0251's prohibition on
per-case growth now rests on a count. No reader change made.

CGA: versor_condition is 0.22% of a turn, not the "~10x proof latency" I claimed
— that multiplied an isolated microbenchmark by a call count and compared it to
a single verdict's latency. The real cost is geometric_product at 33,986
calls/turn (~73%) via cga_inner in search paths. The obvious closed form is NOT
bit-exact (954/4000 in f32); backend.vault_recall's serial fold IS (3000/3000,
worst-rel 0) and is the correct target. cargo test could not run —
static.crates.io is denied by the sandbox network policy — so the Rust parity
question stays open and the typestate lane is carried forward, not shipped
uncompiled.

## Not landed: three lines owed to .github/workflows/smoke.yml

The CI smoke gate is narrower than the local one —
test_pack_draft_serve_boundary.py (ADR-0253 INV-33) has been local-only, unseen
because the parity pin checked one direction. The edit was authored and rejected
at push for lacking the `workflow` OAuth scope, so it is recorded as a named,
dated PENDING_IN_CI exception rather than dropped: the assertion still fires on
any new divergence, and a second guard fires once the three land.

[Verification]: pre-push gates all green — smoke 236 passed, warmed_session 10
passed, deductive 285 passed. Ratification 14, ADR index 5, CLI suites 10.
Grounding/epistemic sweep 741 passed 1 skipped. workbench-ui 598 passed across
73 files, tsc -b clean. capability index 11 passed, digest unchanged. Math
holdout correct=5 wrong=0 refused=495. Committed chain corpora byte-unchanged
after the tests that write to them.
Environment caveat: the repo pins requires-python ==3.12.13, which uv cannot
fetch for linux-x86_64, so `uv sync --locked` fails. All Python runs used a
scratch venv on 3.12.11 with declared deps — not the locked universe, not the
full ~12k suite. The pin was left untouched. Re-run on a 3.12.13 host before
treating this as merge evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FduW6Krm3PPQv3P5iwBYtx
2026-07-25 04:51:15 +00:00

174 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Math Phase 4.2 — case-first traces for 0000 / 0001 / 0148 / 0082
**Baseline reproduced:** `holdout_dev: correct=5 wrong=0 refused=495 (n=500)`
unchanged from the 2026-07-24 measurement, five days and one arc later.
**Deliverable:** the per-case trace `math-reader-phase-4-1-status-2026-07-24.md`
asked for, plus a corpus-frequency measurement that **changes the
recommendation**. All four named cases are blocked by surface gaps that are
among the *rarest* in the 500. Building for them is the overfit move ADR-0251
prohibits — and this document is the first time that has been quantified rather
than argued.
**Nothing was changed in the reader.** See §5.
## 1. Where all four refuse — one seam
Every one of the four fails at the same place, with the same shape of message:
```
candidate_graph: recognizer matched but produced no injection for statement:
'<first statement>' (category=...)
```
Traced through `generate.recognizer_match.match(statement, registry)`:
| case | category assigned | `parsed_anchors` | injection |
|---|---|---|---|
| 0000 | `DISCRETE_COUNT_STATEMENT` | `()` | `()` |
| 0001 | `DISCRETE_COUNT_STATEMENT` | `()` | `()` |
| 0148 | `DESCRIPTIVE_SETUP_NO_QUANTITY` | `()` | `()` |
| 0082 | `DISCRETE_COUNT_STATEMENT` | `()` | `()` |
So the recognizer **classifies** each statement and then extracts **no grounded
anchor**. `_is_comparative_multiplicative_v1_surface` is `False` for all four,
so none of them ever reaches `inject_comparative_multiplicative` — the injector
that would handle "twice as many X as Y". The gate is
`_COMPARE_MULT_ANCHOR_RE` (`generate/math_candidate_parser.py:1229`), anchored
`^…$`, so it must match the **whole** statement:
```
^(?P<actor>ENTITY)\s+VERB\s+(?:a\s+)?(?P<anchor>twice|thrice|half|quarter|third)
\s+as\s+many\s+(?P<unit>\w+(?:\s+\w+)?)\s+as\s+(?P<reference>REF)\s*\.?$
```
## 2. Gap isolation — one variable at a time
Each gap isolated by mutating a known-good baseline in exactly one dimension.
This is the evidence, not the regex reading:
| probe | matches | isolates |
|---|---|---|
| `Aria has twice as many credits as Emily.` | **YES** | baseline |
| `… as many school credits as …` (2-word unit) | **YES** | unit slot admits 2 |
| `… as many high school credits as …` (3-word) | no | **unit slot caps at 2 words** |
| `… as many pounds of lobster as …` | no | **`of`-headed unit phrase** |
| `… as many credits than Emily.` | no | **`than` connector** |
| `… as Emily, who has twice … as Spencer.` | no | **trailing relative clause** |
| `At a station, Aria has twice …` | no | **leading adjunct** |
| `The number of people was twice the total number counted.` | no | **copula + bare `twice the N`** |
| `Mary has twice the amount of carrots as green beans.` | no | **`twice the amount of X as Y`** |
| `… as many credits as the two other students combined.` | no | **aggregate reference** |
## 3. Per-case traces
### Case 0000 — `expected 140.0`
> Aria has twice as many high school credits as Emily, who has twice the number
> of high school credits as Spencer. If Emily has 20 credits, what's twice the
> total number of high school credits the three have?
Blocked by **two** gaps in one sentence: a 3-word unit (`high school credits`)
and a trailing relative clause carrying a *second* comparative. Solving needs a
chained resolution (Emily=20 → Spencer=10, Aria=40), a summation (70), and a
final ×2 on the aggregate. **Four capabilities.**
### Case 0001 — `expected 480.0`
> Hooper Bay has twice as many pounds of lobster than the two other harbors
> combined. If the other two harbors have 80 pounds of lobster each, how many
> pounds of lobster are the three harbors holding?
Blocked by **four** gaps: `than` where the template requires `as`; an
`of`-headed unit (`pounds of lobster`); an aggregate reference (`the two other
harbors combined`); and a multi-token proper name —
`extract_proper_noun_subject` returns `'Hooper'`, not `'Hooper Bay'`, which
alone would fail the injector's narrow actor binding (`actor != actor_token`).
### Case 0148 — `expected 1500.0`
> At a people counting station, the number of people counted on the first day
> was twice the total number counted on the second day. If 500 people were
> counted on the second day, how many people were counted on the two days?
The only one classified `DESCRIPTIVE_SETUP_NO_QUANTITY` — the recognizer does
not see a count statement at all. Three gaps: a leading adjunct; a noun-phrase
quantity as subject (`the number of people counted on the first day`) rather
than an entity; and a copula + bare `twice the total number` form with no
`as many … as` at all.
### Case 0082 — `expected 2.0`
> Mary uses plastic grocery bags that can hold a maximum of twenty pounds. She
> buys 4 pounds of green beans, 6 pounds milk, and twice the amount of carrots
> as green beans. How many more pounds of groceries can Mary fit in that bag?
The refusing statement is not the comparative at all — it is the **capacity**
frame (`bags that can hold a maximum of twenty pounds`), a relative-clause
capacity assertion with a spelled-out cardinal. The comparative arrives in the
next sentence in the `twice the amount of X as Y` form. Then the question wants
capacity Σ(4, 6, 8). **Four capabilities**, none of which is the v1 template.
## 4. The measurement that changes the recommendation
Frequency of each blocking surface feature across all 500 holdout cases:
| surface feature | cases | % of 500 |
|---|---:|---:|
| aggregate reference (`combined` / `together` / `altogether`) | **53** | 10.6% |
| leading adjunct (`At/In/On/During …,`) | **28** | 5.6% |
| `twice/N times as many … as` (any unit width) | 29 | 5.8% |
| `twice the number/amount/total of` | **13** | 2.6% |
| copula + `was/is twice` | **13** | 2.6% |
| capacity frame (`hold` / `maximum of` / `capacity`) | 14 | 2.8% |
| — of the 29, unit is >2 words | **1** | **0.2%** |
| — of the 29, connector is `than` | **2** | **0.4%** |
**The gap blocking case 0000 affects one case in five hundred. The `than` gap
blocking case 0001 affects two.**
That is the whole finding. "Widen the unit slot from two words to three" and
"admit `than` alongside `as`" are precisely the per-case pattern growth ADR-0251
and `reader-arc-overfit-inventory-2026-07-19.md` exist to prohibit — and until
now the prohibition rested on a prior about overfitting. It now rests on a count.
The forms that actually recur are different ones: aggregate reference (53),
leading adjunct (28), `twice the number/amount of` (13), copula-twice (13).
**Honest limit on these numbers.** They are *surface-feature* counts, not
conversion estimates. A case containing `combined` is not thereby convertible —
§3 shows every one of these problems needs three or four capabilities at once,
so a case can carry a frequent feature and still strand on the other three.
The counts bound the opportunity; they do not predict yield. Treating them as
yield is the exact error the arc already made once, when tune-set parse-rate
gain was read as capability and shipped answers that were *wrong* on the exam.
## 5. What was deliberately not done
**No reader change was made.** The plan's Lane 4 exit criterion is the traces
plus `wrong=0` holding, not a shipped extension, and three things argue against
shipping one from this session:
1. §4 shows the named cases' gaps are the rarest in the corpus. Building them
is the prohibited move, now with a number attached.
2. The sealed 1,319-case test split is the final arbiter and must not be read.
A change justified only by movement on the open dev set has exactly the
evidentiary shape ADR-0251 was written against.
3. The four cases are each 34 capabilities deep. There is no "minimum reader
extension" that converts one of them — which is itself the strongest form of
the arc's own conclusion that *real GSM8K statements are individually
multi-capability*.
`wrong=0` still holds: `correct=5 wrong=0 refused=495`.
## 6. Recommendation
Phase 4.2 as scoped — case-first on 0000/0001/0148/0082 — should be recorded
as **traced and declined**, not attempted. The tracing was the right
instruction; its result is that these four are the wrong four.
If the reader arc resumes, the measured entry point is **aggregate reference**
(53 cases, 10.6%) — the single most common blocking surface in the corpus, and
the one whose semantics (`the two other harbors combined` → Σ over a named
complement set) is a genuine compositional capability rather than a template
widening. Scope it the same way this document scopes: isolate the gap, count
its reach, and require that the conversion evidence be a *capability* argument,
not a parse-rate delta.
Relates to [[project-generalization-arc]].