test_pack_measurements_phase2.py (ADR-0043) was reachable only under
`--suite full` and post-merge full-pytest.yml — outside the blocking PR
smoke gate. A regression flipping identity falsifiability (ratified packs
must diverge directionally), the pack-invariant grounding/refusal floor,
or the zero-fabrication invariant therefore cleared the PR gate and
surfaced only post-merge on main (the smoke blind spot for dedicated
test files).
Add the file to both the `smoke` suite tuple (core/cli.py) and the
smoke.yml CI gate so the falsifiability claim blocks-on-regression
rather than detect-after-merge. Adds ~4 min to the PR gate.
Verified: modified smoke gate green on the main base via the
CI-equivalent invocation (140 passed).
The recognizer/candidate-graph path is the single canonical reader.
Retires the flag-gated incremental-reader dispatch that admitted 0/50 on
train_sample and only added a dead fall-through:
- remove _try_comprehension_reader, _try_reader_for_question, _tokenize_sentence
and both dispatch blocks from generate/math_candidate_graph.py
- delete generate/comprehension/lifecycle_runtime_adapter.py (402 LOC,
used only by the question-reader dispatch)
- drop the comprehension_reader_questions config flag and the parse_and_solve
/ _score_one_candidate_graph config threading
- remove the --use-reader runner plumbing + flag-ON/OFF delta report from
the train_sample runner; refresh report.json (drops stale use_reader field
and a stale refusal-reason; verdicts unchanged at 3/47/0)
- remove the now-dead use_reader field from teaching/coverage.py
CoverageReport + the core teaching coverage CLI flag
- delete tests/test_reader_coexistence.py (flag-ON/OFF premise dissolved);
fix 3 ADR-0174 build_report calls and 2 subprocess invocations
lifecycle.py and audit.py are KEPT — they are load-bearing for the ADR-0172
math-contemplation teaching corridor (audit_problem -> teaching/math_*),
which a pre-deletion trace surfaced. The parent ADR's plan to delete
lifecycle.py was wrong; only its GSM8K scoring dispatch was inert.
Net -1,038 LOC (code + tests). Behavior-preserving:
- train_sample 3/47/0, byte-identical verdicts to pre-5a baseline
- determinism holds; smoke/packs/runtime/cognition/teaching lanes green
- contemplation corridor + lifecycle/audit tests pass
Pre-existing (NOT introduced here; reproduce on base with changes stashed):
5 out-of-curated-lane stale committed-artifact / stale-assertion failures
(test_math_evidence_e2e, test_adr_0126_runner_wiring, G3/coverage_probe
report-match, test_refusal_taxonomy_lane rebuild).
The repo is public. The thesis is *decoding, not generating* with
wrong=0 as the load-bearing invariant. The demo any visitor can run
to see the loop turn end-to-end on the canonical pack:
git clone https://github.com/AssetOverflow/core
cd core && uv pip install -e .
core demo flywheel
Four falsifiable scenes:
1. RATIFY — apply_composition_claim writes source JSONL; RAT-1
auto-compile regenerates compositions.jsonl + bumps
manifest.composition_checksum
2. LOAD — composition_registry picks up the new entry on the
next runtime turn
3. SOLVE — "Lilibeth fills 6 baskets where each basket holds
50 strawberries. How many strawberries does Lilibeth
have?" admits via matcher → injector → admission →
candidate-graph and produces answer=300
4. HAZARD — case 0050 (wrong=0 canary) remains refused; no SAFE
composition category can convert it
All four scenes byte-deterministic. The canonical pack is read-only
throughout; the demo mutates only a synthetic test pack in a
tempfile.TemporaryDirectory. One-time recognizer seed is idempotent
(same content_digest each run → no duplicate proposal log entries).
Exit code 0 iff all scenes pass; --json for CI integration.
Also adds:
- README "Watch the flywheel turn — one command" section pointing
to the demo + the coverage CLI (per-shape histogram + hazard pin)
- ProposalLog entry for the multiplicative_aggregate recognizer
with extract_values=True (one-time operator seed)
Files:
- evals/flywheel_demo/run_tour.py (new) — the four-scene tour
- evals/flywheel_demo/__init__.py (new)
- core/cli.py — `flywheel` added to `core demo` choices + dispatch
- README.md — new "Quick Start" subsection
- teaching/proposals/proposals.jsonl — seeded recognizer
Three review fixes:
1. Security: validate lane/split/version against ^[a-z0-9_]+$ before
building the runner module name. The runner_args list is passed to
subprocess.run without shell=True (no shell injection possible),
but defense-in-depth blocks arbitrary token characters from
reaching Python's -m module loader. Bad input now errors at the
CLI boundary with a clear message.
2. Bug-risk: _classify_refusal docstring referenced a
no_admissible_candidate bucket that the implementation never
emitted. Aligned docstring with actual buckets
(no_admissible_question / no_admissible_statement). Also made all
matching consistently case-insensitive (was mixed — some checks
used raw reason, one used .lower()).
3. Bug-risk: fetch_committed_baseline wrote to
.git/coverage_baseline_tmp.json. Replaced with tempfile.mkstemp in
the system temp dir — avoids (a) failures in non-git worktrees
where .git is a file pointer, (b) concurrent-access collisions
between simultaneous operators.
Tests (+3 new):
- test_classify_refusal_is_case_insensitive
- test_classify_docstring_matches_implementation_buckets
- test_fetch_committed_baseline_uses_system_temp
All 16 coverage tests green. Verified the validation:
core teaching coverage --lane 'evil; rm -rf /'
→ ERROR: lane='evil; rm -rf /' must match ^[a-z0-9_]+$
Brief D from PR #407. Closes the "flying blind on per-shape coverage"
gap identified in RAT-1's audit (finding 6).
After this PR, every operator can run a single command to see exactly
which refusal modes their work moved (or didn't), without re-eyeballing
report.json by hand.
Modules
-------
- teaching/coverage.py — pure aggregator:
- _classify_refusal — maps each per-case refusal reason to a
stable bucket (recognizer_empty_injection(<ShapeCategory>),
no_admissible_question, no_admissible_statement,
unexpected_question_count, other)
- build_coverage_report — reads a lane's report.json + emits a
CoverageReport with counts, refusal_taxonomy (sorted by count
desc), case_0050_verdict, optional delta vs baseline
- fetch_committed_baseline — uses `git show HEAD:<relpath>` to
pull the baseline report.json for delta computation
- core/cli.py:
- cmd_teaching_coverage — formats the report for terminal output
- core teaching coverage [--lane gsm8k_math] [--split train_sample]
[--version v1] [--use-reader] [--run] [--delta] [--json]
CLI output example
------------------
Lane: gsm8k_math/train_sample/v1 (use_reader=True)
Counts: correct=3 refused=47 wrong=0
Refusal taxonomy:
21 recognizer_empty_injection(discrete_count_statement)
6 no_admissible_statement
5 recognizer_empty_injection(multiplicative_aggregation)
4 no_admissible_question
4 recognizer_empty_injection(currency_amount)
3 recognizer_empty_injection(rate_with_currency)
2 recognizer_empty_injection(descriptive_setup_no_quantity)
2 recognizer_empty_injection(temporal_aggregation)
Wrong=0: ✓
Case 0050 hazard pin: refused ✓
Tests (13 new)
--------------
tests/test_teaching_coverage_cli.py — classification narrowness,
counts aggregation, case 0050 verdict capture, delta computation,
missing-baseline path, missing-report error, taxonomy sort order,
wrong=0 invariant visibility via as_dict.
Suite results
-------------
core test --suite teaching -q → 106 passed (93 → +13)
core test --suite runtime -q → 20 passed
core test --suite packs -q → 127 passed
core eval gsm8k_math --split public → 150/150, wrong=0
Note on Brief E (lexical auto-compile): the audit was WRONG. The
lexicon loader (generate/comprehension/lexicon.py::load_lexicon)
reads from the per-category source files directly; the compiled
lexicon.jsonl is only a manifest-checksum pin, not the source of
truth at runtime. apply_lexical_claim() writes a new entry → next
turn the loader sees it. Brief E is a non-issue; closing without a
code PR.
Verified by direct test: stage a clone of the math pack, write a
synthetic lemma to drain_token.jsonl, clear the lexicon cache, load
again → new entry present. So 3 of the 5 audit gaps closed (A, D,
E-as-correction); B and C remain as the next operator dispatch
targets.
Independent of PR #406 (RAT-1) and PR #408 (WAVE-A). Based on main.
Addresses 5 of 47 train_sample "recognizer matched but produced no
injection" refusals (the largest single failure-mode bucket
identified in RAT-1's audit).
Modules
-------
- generate/recognizer_match.py:
- _MULT_AGG_EACH_WEIGHING_RE — regex for "<Subject> <bake-verb>
<M> <outer-noun>, each <weigh-verb>ing <N> <unit>" pattern
- _try_extract_each_weighing_anchor — extracts M, N, subject,
inner unit; emits pre-composed CandidateInitial(value=M*N) with
composition_evidence so RAT-1's _composed_initial_admissible
gate verifies INPUT tokens ground (preserves wrong=0)
- _match_multiplicative_aggregation dispatches to the value
extractor when spec carries extract_values=True; specs without
that flag get the existing detection-only return path
(byte-identical legacy behavior)
- generate/recognizer_anchor_inject.py:
- inject_multiplicative_aggregation — new per-category injector;
narrow by anchor.kind so ME-3/ME-4 additive/subtractive anchors
(which share the same matcher entry point) continue to flow
through composition_registry consult instead of WAVE-A's direct
path
- registered in _INJECTORS dict (2nd entry after DCS)
- core/cli.py:
- seed-recognizer CLI gains --extract-values flag to opt the
canonical_pattern into the value-extracting matcher path
Seeded artifacts
----------------
- proposals.jsonl: rat1-seed-4dc30608fb783bc7 — multiplicative_
aggregation recognizer with anchor_kind=multiplicative_aggregate,
extract_values=True, observed_units covering ounces/strawberries/
questions/etc.
Live result on train_sample
---------------------------
- wrong == 0 preserved (3/47/0 baseline)
- Case 0050 hazard pin held
- public 150/150 preserved
- packs suite: 127 → 131 (+4 new WAVE-A tests, all green)
- teaching suite 93 unchanged
- runtime suite 20 unchanged
End-to-end synthetic solve (FIRST WAVE-A admission):
"Lilibeth fills 6 baskets where each basket holds 50 strawberries.
How many strawberries does Lilibeth have?" → answer=300
Cases that moved (statement now admits; refusal shifted downstream):
- Case 0025 (Lilibeth): statement admits via WAVE-A; refusal moved
to question parser ("If three of Lilibeth's friends pick the same
amount, how many strawberries do Lilibeth and her friends pick in
all?")
- Case 0047 (John bakes 12 macaroons): statement 1 admits; refusal
moved to statement 2
Eval correct count unchanged because the QUESTION parser (and
multi-statement cross-sentence reasoning) is the next bottleneck.
RAT-1's audit identified that gap; WAVE-A closes the injector half.
The remaining 3 multiplicative_aggregation refusals (0006, 0013,
0045) have different shape patterns the WAVE-A regex does not yet
cover; they're follow-up matcher extensions in the same architecture.
Tests
-----
- tests/test_wave_a_multiplicative_aggregation_injector.py (10
tests): each-weighing + each-basket-holds admission shapes,
detection-only path preserved when extract_values absent,
unobserved unit / pronoun / zero count refusals, end-to-end
inject_from_match dispatch, the Lilibeth canary solve,
wrong=0 preserved, case 0050 hazard pin
Stacks on PR #406 (RAT-1).
The user's question — "shouldn't we be running it multiple times so
it can learn? or is that part broken?" — exposed that the math
teaching loop's `ratify → admit` closure had been structurally
broken at the connector between operator ratification and runtime
visibility. The handlers wrote source files (compositions/, frames/)
that the runtime loader never read because no compile step
regenerated the runtime artifacts.
This PR fixes the gap end-to-end AND fires the first live composition
admission on the canonical pack.
Modules
-------
- language_packs/compile_pack.py — unified compile step that
regenerates frames.jsonl + compositions.jsonl + updates
manifest.{frame,composition}_checksum atomically. Idempotent.
- teaching/math_composition_ratification.py — apply_composition_claim
now calls compile_pack at end of successful ratification. Closes
the source-file→runtime-artifact gap.
- teaching/math_frame_ratification.py — same auto-compile wire for
apply_frame_claim.
- generate/math_candidate_parser.py — CandidateInitial gains optional
composition_evidence Mapping field. When populated, signals the
candidate was produced by a registry-gated composition (ADR-0169);
the value/unit/entity are DERIVED arithmetic over grounded inputs.
- generate/math_candidate_graph.py — new _composed_initial_admissible
predicate that branches on composition_evidence. Wrong=0 preserved
by requiring each composition INPUT token (count, amount) to ground
in source_span literally; the derived value is admitted because the
arithmetic over grounded inputs is deterministic.
- generate/math_candidate_graph.py — discourse-level prior_subject
tracking: capture proper-noun subjects from ALL statement sentences
(including ADR-0136.S.0 context-filler sentences that get filtered
out before the candidate loop). Without this, "John adopts a dog"
(no numbers) is dropped and the cross-sentence subject resolver for
case 0019 sees prior_subject=None.
- generate/recognizer_match.py — all four composition matchers
(ME-1 currency-per-unit same-sentence, ME-2 cross-sentence, ME-3
additive, ME-4 subtractive) now populate composition_evidence in
CandidateInitial. Also added standalone " each " / " apiece " to
_PER_UNIT_TOKENS so currency_amount detection-only matcher refuses
per-item costs instead of swallowing them.
CLIs
----
- core teaching compile-pack — explicit operator surface for
regenerating runtime artifacts. JSON output for CI integration.
- core teaching seed-recognizer — operator surface for seeding a
RatifiedRecognizer entry in the proposal log for a given
(shape_category, anchor_kind). Writes created + transition(accepted)
events directly via ProposalLog._append.
Seeded artifacts (the actual loop closure)
------------------------------------------
- proposals.jsonl: new rat1-seed-48dd2673d6ad673d RatifiedRecognizer
entry for shape_category=rate_with_currency,
anchor_kind=currency_per_unit_composition.
- compositions/multiplicative_composition.jsonl: ratified
"bound(count) × bound(unit_cost)" affirms entry sourced from
case 0019 evidence.
- compositions.jsonl + manifest.composition_checksum: compiled
runtime artifact + manifest pin (RAT-1 auto-compile).
Live result on train_sample
---------------------------
- wrong == 0 preserved (3 correct / 47 refused / 0 wrong)
- Case 0050 hazard pin holds (refused)
- public split 150/150 preserved
- Case 0019 sentence 1 ("requires 3 vet appointments, which cost
$400 each") NOW ADMITS via composition. Previously refused with
"recognizer matched but produced no injection". The refusal moved
downstream to sentence 2 (a different currency_amount detection
bottleneck that is its own follow-up).
This is the first time a composition ratification on the canonical
pack actually reaches the runtime. The flywheel turned one
revolution.
Tests
-----
- tests/test_rat1_end_to_end_admission.py — 4 new live tests:
composition statement admits on isolated synthetic problem, case
0019 cross-sentence admission, wrong=0 preserved on train_sample,
case 0050 hazard pin.
- tests/test_consumption_empty_registry_no_op.py — refactored to use
isolated synthetic packs (the canonical pack may now carry ratified
entries).
- tests/test_math_{frame,composition}_ratification.py — updated
"manifest checksum unchanged" tests to "lexicon checksum
preserved" semantics: RAT-1 auto-compile may add the new optional
checksum fields; pre-existing lexicon checksum stays untouched.
Suite results: teaching 93, packs 131 (+4), runtime 20. All green.
Final PR of the matcher-extension wave. Ships:
1. tests/test_me5_all_categories_integration.py — 4 new tests:
- test_all_three_canaries_admit_through_full_pipeline: stages a
pack with all three SAFE_COMPOSITION_CATEGORIES entries +
ratifies, runs Maria/Sam/Tom canaries through matcher →
inject_from_match, asserts admission for all three
- test_partial_pack_only_admits_present_categories: refusal-
preferring when only one category is ratified
- test_all_safe_categories_have_extension_admission: pins that
SAFE_COMPOSITION_CATEGORIES is exactly the three covered
categories (breaks if future ADR widens without matcher)
- test_falsifies_uniformly_suppresses_across_categories:
polarity discipline holds across all three matchers
2. docs/handoff/ME1-ME5-MILESTONE.md — wave milestone doc:
- architecture diagram (audit → ratify → compile → load →
match → consult → admit)
- SAFE_COMPOSITION_CATEGORIES coverage matrix
- invariants preserved across the entire stack
- scope boundary (what does NOT fire yet — RAT-1 follow-up)
- recommended next dispatch
3. Test registration in core/cli.py packs suite.
Across the full ME-1..ME-5 stack:
- 5 stacked PRs (#400/#401/#402/#403/#404)
- 1 foundation PR (#398 — consumption wiring)
- 114 new tests, all green
- packs suite 127 passed
- core eval gsm8k_math --split public → 150/150, wrong=0
- All three SAFE_COMPOSITION_CATEGORIES have matcher extensions
Anti-regression invariants preserved across the entire stack:
- wrong == 0 on public split
- Case 0050 hazard pin (parametrized over all three categories)
- ADR-0166 — no new eval lanes
- ADR-0167 partition — no cognition imports
- ADR-0169 mutation boundary — registry is a gate, not arithmetic
- All matcher detection paths byte-identical
- engine_state/* never committed
- SAFE_COMPOSITION_CATEGORIES enforced at write AND load
- polarity falsifies honored uniformly
Live train_sample admission requires operator-seeded ratifications
(RAT-1 follow-up). Wiring is end-to-end correct, verified by ME-5
integration tests.
Memory: milestone-me1-me5-matcher-extensions-complete saved.
Stacks on PR #403 (base: feat/matcher-extension-subtractive).
Extends _match_multiplicative_aggregation with a new branch keyed on
anchor_kind="additive_quantity_composition". When a statement carries
"<Subject> <verb> <N> <unit> and <M> <unit>" (same unit) shape, emits
a pre-composed CandidateInitial(N+M, unit) and publishes
composition_shape="bound(qty_a) + bound(qty_b)".
Subject binding under Option A (refuse on pronoun / determiner / no
proper-noun head). Cross-sentence subject support (mirroring ME-2)
is deferred — not needed for the v1 ME-3 canaries.
Verb whitelist: lost / gained / earned / saved / made / paid / spent /
bought / sold / added / removed / received. Verbs that route through
CandidateInitial.matched_anchor's existing post-init whitelist;
unmapped verbs fall back to "had".
Unit normalization: rstrip 's' for plural matching (pounds vs pound).
Cross-unit composition refused — no conversion table in v1.
Tests (15 new, all green):
- same-unit admission with sum
- pronoun subject refuses
- determiner subject refuses
- cross-unit refuses
- unobserved unit refuses
- zero count refuses
- plural normalization
- unknown verb refuses
- multiplicative_aggregate detection path unaffected
- wrong anchor_kind refuses
- anchor audit fields complete
- source_span substring invariant
- no match returns None
- end-to-end admission via composition_registry
- end-to-end falsifies suppresses
Registered in core/cli.py "packs" suite. core test --suite packs -q →
106 passed (91 existing + 15 new).
Anti-regression invariants preserved:
- wrong == 0 on gsm8k_math public 150/150
- Case 0050 hazard pin holds
- ADR-0166 — no new eval lanes
- ADR-0167 partition — no cognition imports
- Original multiplicative_aggregate detection path byte-identical
- ME-1 currency-per-unit path unaffected
- ME-2 cross-sentence path unaffected
- engine_state/* not committed
Live train_sample admission requires the same operator workflow as
ME-2: a RatifiedRecognizer for the new anchor_kind + composition_registry
entry for "bound(qty_a) + bound(qty_b)" under additive_composition.
Without those, the wiring is correctly positioned but dormant — no
regression in the live eval.
Stacks on PR #401 (base: feat/matcher-extension-cross-sentence-subject).
Admits case 0019's composition sentence via prior_subject resolved
from upstream sentences. Stacks on PR #400 (ME-1).
Modules
-------
- generate/recognizer_match.py:
- _CROSS_SENTENCE_COMPOSITION_RE — regex for "requires N noun, which
cost(s) $X each" (no subject prefix)
- try_extract_cross_sentence_composition_anchor(statement, spec,
prior_subject) — refuses on None / empty / pronoun prior_subject;
publishes the same composition_shape + composed_initial payload as
ME-1, sourced via prior_subject
- extract_proper_noun_subject(statement) — head proper-noun extractor
used by callers to track running prior_subject; rejects determiners,
sentence-initial connectors (After/How/Every/...), and pronouns
- match() dispatcher gains keyword-only prior_subject parameter;
when a per-category matcher returns None for a RATE_WITH_CURRENCY
recognizer with currency_per_unit_composition anchor_kind AND
prior_subject is supplied, the cross-sentence helper is tried as
a fallback
- generate/math_candidate_graph.py:
- tracks _prior_subject across statement_sentences iteration
- passes prior_subject to recognizer_match.match()
- updates _prior_subject from each sentence's head proper-noun
Tests (19 new, all green)
-------------------------
- test_me2_cross_sentence_subject.py (15 tests)
- subject extraction narrowness (proper noun / determiner / connector
/ pronoun / non-string)
- cross-sentence helper happy path + refusals (None, empty, pronoun,
unobserved currency / per_unit, wrong anchor_kind, zero count,
multi-match)
- source_span substring invariant
- kind label "currency_per_unit_composition_cross_sentence"
- test_me2_case_0019_admits.py (4 tests)
- case_0019_admits_with_prior_subject_john — the truth test
- case_0019_refuses_without_prior_subject — ME-1 Option A still holds
- case_0019_refuses_with_pronoun_prior — refusal-preferring
- maria_same_sentence_unaffected_by_prior_subject — ME-1 path intact
Registered in core/cli.py "packs" suite.
Suite results
-------------
core test --suite packs -q → 91 passed (existing + ME-1's 21 + 19 new)
core test --suite runtime -q → 20 passed
core eval gsm8k_math --split public → 150/150, wrong=0
Scope boundary
--------------
The wiring is load-bearing AND tested end-to-end via synthetic
recognizer registry (test_case_0019_admits_with_prior_subject_john
proves the full chain match → inject → admit).
For the LIVE train_sample case 0019 admission, two ratifications must
also be seeded (operator workflow outside this PR's code scope):
1. A RatifiedRecognizer in the proposal log with shape_category=
RATE_WITH_CURRENCY and canonical_pattern carrying
anchor_kind="currency_per_unit_composition"
2. A composition_registry entry for "bound(count) × bound(unit_cost)"
under multiplicative_composition with polarity=affirms
With both ratifications in place, case 0019 admits via the wiring
this PR ships. Without them, the live train_sample run remains at
the 3/47 baseline (preserved; no regression).
Anti-regression invariants preserved
------------------------------------
- wrong == 0 on gsm8k_math public
- Case 0050 hazard pin holds (no _COMPOSITION_SUBJECT_BUY_RE or
_CROSS_SENTENCE_COMPOSITION_RE match on case 0050's sentences)
- ADR-0166 — no new eval lanes
- ADR-0167 partition — no cognition imports
- ME-1 Maria same-sentence path byte-identical (test pins)
- Existing currency_per_unit_rate path unaffected (test pins)
- prior_subject is keyword-only on match() (additive; old callers
unaffected)
- engine_state/* not committed
Stacks on PR #400 (base: feat/matcher-extension-currency-per-unit-composition).
Closes the consumption-half of the math teaching loop for two of three
sub-types per docs/handoff/CONSUMPTION-WIRING-DISPATCH-PACK.md (PR #397).
Companion to the doctrinal brief in PR #396.
Modules
-------
- language_packs/compile_frames.py — byte-deterministic compile of
frames/*.jsonl → frames.jsonl (sorted by (frame_category, surface_form))
- language_packs/compile_compositions.py — same shape for
compositions/*.jsonl → compositions.jsonl
- generate/comprehension/frame_registry.py — load_frame_registry()
mirroring load_lexicon: cache by (path, mtime, sha256), manifest
checksum verification (optional frame_checksum field), polarity
validation, conflict detection, empty-registry no-op
- generate/comprehension/composition_registry.py — same shape PLUS:
* SAFE_COMPOSITION_CATEGORIES enforced at LOAD (defense in depth;
raises WrongCompositionCategory on any unsafe category — protects
against pack edits that bypass the handler)
* polarity "falsifies" exposed via is_falsified() (consumer must
suppress; not silently treated as affirms)
- language_packs/compiler.py — manifest verification extended for
frame_checksum + composition_checksum, mirroring the proven
glosses_checksum pattern (optional fields; backward-compatible)
- generate/recognizer_anchor_inject.py — inject_from_match consults
composition_registry when the per-category injector returns empty
AND the matcher publishes ``composition_shape`` in parsed_anchors.
Registry is a gate (admissibility) not an arithmetic primitive
(ADR-0169 §"Mutation boundary").
Tests (38 new, all green)
-------------------------
tests/test_frame_registry_load.py (11 tests)
tests/test_composition_registry_load.py (11 tests)
tests/test_composition_consult_in_injector.py ( 6 tests)
tests/test_consumption_case_0050_hazard_pin.py( 3 tests, parametrized
over allowlist)
tests/test_consumption_empty_registry_no_op.py( 4 tests)
tests/test_consumption_partition.py ( 3 tests)
Registered in core/cli.py "packs" suite.
Suite results
-------------
core test --suite teaching -q → 93 passed
core test --suite runtime -q → 20 passed
core test --suite packs -q → 51 passed
core eval gsm8k_math --split public → 150/150, wrong=0
Truth-test rows (6-row binding table in dispatch pack):
#1 Case 0019 admits ............. PARTIAL — see Scope Boundary below
#2 Case 0050 stays refused ....... PASS
#3 train_sample 3/47 → ≥4/46 ..... PARTIAL — same as #1#4 wrong == 0 preserved .......... PASS
#5 public split 150/150 .......... PASS
#6 Empty-registry no-op .......... PASS
Scope Boundary (honest finding)
-------------------------------
Rows #1 and #3 (case 0019 admission) require a matcher extension that
publishes ``composition_shape`` + a pre-composed CandidateInitial in
parsed_anchors. The existing currency_amount / multiplicative_aggregation
matchers in generate/recognizer_match.py are detection-only (return
empty parsed_anchors). This PR ships the consumption infrastructure
correctly but the runtime path remains dormant until a follow-up PR
extends the matcher. The dispatch pack's truth test #1/#3 cannot fire
without that extension.
The wiring is positioned correctly: inject_from_match → consult
composition_registry → admit on affirms-with-payload, suppress on
falsifies, refuse on absence. A synthetic recognizer match with
populated composition_shape + composed_initial DOES admit through the
new path (covered by 6 tests in test_composition_consult_in_injector.py).
A follow-up brief naming the matcher-extension work is the
recommended next step.
Anti-regression invariants verified
-----------------------------------
- wrong == 0 on core eval gsm8k_math (public 150/150)
- case 0050 stays refused (parametrized over allowlist categories)
- ADR-0166 — no new eval lanes
- ADR-0167 partition — no cognition imports in any new module
- Empty-registry runtime byte-identical to today (no-op test)
- SAFE_COMPOSITION_CATEGORIES enforced at write AND load
- polarity semantics (affirms vs falsifies) honored
- engine_state/* never committed
Bundles three post-Tier-1 follow-ups into one PR (no scope change, no
new ADR — implementation tightening on the already-shipped corridor).
(1) Standalone JSONL self-containment
teaching/math_contemplation_proposal.py
+ to_jsonl_record() — emits proposal_id + full evidence_pointers
(nested dicts including audit_row) + full reasoning_trace.steps
+ from_jsonl_record() — inverse; goes through build_proposal()
so all invariants are re-validated; raises on proposal_id mismatch
canonical_bytes() UNCHANGED (still the content-hash function;
trace_id/proposal_id stability preserved)
core/cli.py W3 lane now writes to_jsonl_record() output instead of
canonical_bytes() — same compact-JSON encoding (sort_keys=True,
ensure_ascii=False, separators=(",", ":"))
workbench/readers.py loads via self-contained record fields directly;
decompose_audit() re-run removed. read_math_proposal() now reads
reasoning_trace.steps and evidence_pointers from the JSONL record.
(2) Widened change_kind heuristic dispatch
teaching/math_contemplation.py
+ _CHANGE_KIND_BY_PAIR table on (refusal_reason, missing_operator):
(unexpected_category, pre_frame_filler_sentence) → matcher_extension
(unexpected_category, multi_subject_sentence) → frame_reclassification
(unexpected_category, fraction_percentage_literal) → matcher_extension
(unexpected_category, descriptive_frame_question) → frame_reclassification
(unresolved_pronoun, pronoun_resolution) → matcher_extension
Single-key fallback (lexicon_entry/narrowness_violation/
frame_unrecognized) retained for completeness.
hypothesis-step justification text updated to reflect new table.
Result on audit_brief_11.json:
3 matcher_extension (was 0)
2 frame_reclassification (was 0)
3 injector_sub_shape (was 8)
0 vocabulary_addition (no unknown_word group ≥2 in train sample)
(3) shape_category structural gap
MathReaderRefusalEvidence does not carry shape_category, so the
proposal cannot derive it. All proposals continue to emit
ShapeCategory.UNCATEGORIZED with a structural-gap comment. No
invented values — handler dispatch decision (per ADR-0167-FOLLOWUPS
§1) drives ratification routing today, not shape_category.
Tests
+ W1: 5 new tests (to_jsonl_record self-containment, round-trip,
byte stability, proposal_id mismatch rejection, canonical_bytes
unchanged invariant)
+ W2: 3 new pair-dispatch tests + real-audit change_kind distribution
test + shape_category-uncategorized test
+ W3: 2 new tests (records are self-contained, round-trip via
from_jsonl_record); existing byte-comparison test updated to use
proposal_id ordering instead of canonical_bytes
+ W4: existing 6 tests updated to build JSONL via to_jsonl_record;
+ 1 new decoupling test that drops teaching.math_contemplation from
sys.modules and verifies the workbench still loads + serves detail
Verification
- core eval math-contemplation produces the expected 3/2/3 distribution
- core test --suite teaching -q → 33 passed
- core test --suite runtime -q → 20 passed
- All 57 ADR-0172 W1-W4 tests pass (49 existing + 8 new)
Determinism / invariants preserved
- canonical_bytes() byte-stable (test pins this)
- to_jsonl_record() byte-stable via sort_keys=True + no floats
- wrong=0 invariant: proposals stay evidence-only; no auto-apply
- ChangeKind Literal unchanged (4 values; no new ones invented)
Add decompose_audit(audit_path) to teaching/math_contemplation.py.
Groups audit_brief_11.json refusal rows by
(refusal_reason, missing_operator), emits one
MathReaderRefusalShapeProposal per group of >=2 rows, each carrying a
4-step ReasoningTrace (observation -> grouping -> hypothesis ->
conclusion).
Determinism:
- Group iteration sorted by (refusal_reason, missing_operator).
- Evidence per group sorted by case_id.
- Output tuple sorted by proposal_id.
- 10x rerun -> byte-identical proposals + trace_ids.
Pure read-only: audit file is not mutated, no proposals written to
disk, no chat/field/generate/algebra imports.
Tests (tests/test_adr_0172_w2_decomposer.py): real-audit emission,
determinism (10x), evidence floor, change-kind dispatch over all four
heuristic branches, four-step trace, case_id sort, proposal_id sort,
empty input -> empty tuple, unmapped operator skip, missing file ->
FileNotFoundError, no-mutation contract.
Added to core test --suite teaching.
* feat(ADR-0167/W1-A): MathReaderRefusalEvidence schema + canonical-bytes
Foundation type for routing comprehension-reader refusals into the
teaching corridor. Frozen dataclass with sha256 evidence_hash computed
from deterministic canonical bytes (mirrors state.to_canonical_bytes
pattern). Includes SUB_TYPE_FOR_OPERATOR mapping table covering all 13
missing_operator values in the current audit artifact.
Wave 1 only — no runtime mutation, no teaching-store integration, no
admission path. Downstream W2-A/B/C/D type-import from this module.
* feat(ADR-0167/W2-C): domain discriminator + cross-domain audit
- Links to the audit doc: docs/handoff/ADR-0167-W2C-cross-domain-audit.md
- Inventory details: 5 construction sites, 8 consumption sites
- Verification: 0 cognition test files were modified; all tests are green
- Downstream partition work flagged: contemplation indexing (in teaching/contemplation.py) and replay gate (in teaching/proposals.py)
Adds two pre-gate checks to propose_from_candidate that fire after the
Step 2 capacity check and before the replay gate. No log entry is
written on either refusal — the append-only invariant holds.
Check order at function entry (ADR-0161 §3):
1. Capacity (Step 2) → RefusedAtCapacity
2. Duplicate → RefusedAsDuplicate
3. Dependent_on_pending → RefusedAsDependent
4. Replay gate → auto-reject on regression
New frozen dataclasses:
@dataclass(frozen=True, slots=True)
class RefusedAsDuplicate:
proposal_id: str
existing_state: str # covers all states: pending/accepted/rejected/withdrawn
reason: str = "duplicate"
@dataclass(frozen=True, slots=True)
class RefusedAsDependent:
candidate_id: str
dependent_on: tuple[str, ...] # pending proposal_ids that block
overlapping_lemmas: tuple[str, ...] # normalised lemmas that triggered
reason: str = "dependent_on_pending"
Lemma-overlap rule: case-insensitive exact-match on strip().lower().
Conservative — over-reject rather than admit-with-hidden-dependency.
False positives are recoverable (re-emit after blocker is ratified);
false negatives silently couple ratification choices.
CLI surfaces both outcomes in cmd_teaching_propose and
cmd_teaching_propose_from_exemplars (exit code 1).
Step 2 backpressure tests updated: made pre-populated candidates use
unique objects to avoid triggering the new dependency check, and
updated idempotency assertions to reflect the new RefusedAsDuplicate
return for re-submitted content.
Co-references: ADR-0161 §3, Step 1 PR #296, Step 2 PR #311,
ADR-0057, ADR-0151.
Phase C is the first phase where operator-authored exemplar corpora
become engine-derived recognizer proposals automatically. The math
thesis ("decodes, not generates") manifests in the math lane here.
Modules
- teaching/exemplar_ingest.py — pure-function loader for Phase B
exemplar JSONLs. ExemplarCorpus carries a sha256 digest over its
canonical (sorted-by-exemplar_id, sort-keyed) bytes.
- teaching/recognizer_synthesis.py — per-category synthesizers
(_synthesize_descriptive_setup_no_quantity / _temporal_aggregation /
_rate_with_currency) distil a corpus into one RecognizerSpec.
Determinism: same corpus -> byte-identical spec. Narrowness: the
spec records only observed sub-shapes; an out-of-corpus currency
symbol or window unit does not match. Phase B author_notes surface
in canonical_pattern.unresolved_notes — never silently dropped.
- teaching/contemplation.py — contemplate_exemplar_corpus(corpus)
returns a DiscoveryCandidate whose proposed_chain encodes the
RecognizerSpec as a synthetic four-field chain plus the full
recognizer_spec submap. Evidence cites every exemplar's case_id.
- teaching/replay.py — run_admissibility_replay_gate(spec, *,
active_corpus_path=None) runs cognition + G1..G5+S1 + GSM8K
train_sample. In-process baseline cache keyed on the active
corpus digest. WRONG-COUNT INVARIANT: if a candidate run lifts
the GSM8K train_sample wrong count, gate returns
replay_equivalent=False with
regressed_metrics=["gsm8k_train_sample_wrong_count"].
- teaching/source.py — ProposalKind widened with "exemplar_corpus";
exhaustive-match docs + tests updated.
CLI
- core teaching propose-from-exemplars <path> [--all] [--review-date]
[--log] [--json]. Routes the candidate through the existing
propose_from_candidate path with the admissibility gate substituted
for the cognition-only run_replay_equivalence. Never auto-accepts;
proposals land as pending for operator review.
Tests (38 new)
- tests/test_exemplar_ingest.py (12) — load, digest stability,
malformed-record rejection, file-name binding, read-only purity.
- tests/test_recognizer_synthesis.py (16) — determinism, purity,
per-category subsumption, narrowness (out-of-corpus seeds rejected),
author_notes surfaced.
- tests/test_admissibility_replay_gate.py (6) — happy path, cache
hit/invalidation, WRONG-COUNT INVARIANT regression, capability-axis
regression, cognition regression.
- tests/test_propose_from_exemplars_cli.py (4) — single corpus, --all,
determinism, read-only snapshot.
Acceptance evidence (dry run)
- All three Phase B corpora produce replay_equivalent=true,
wrong_count_delta=0. Proposal IDs:
descriptive_setup_no_quantity: 59223f13722f906a1cf9b65d9b01c990
rate_with_currency: 46ce297f797ff16da12db5de422ca3c9
temporal_aggregation: a3b892546977c5f0f64c578d6052adbd
- G1..G5+S1 wrong=0 unchanged; GSM8K train_sample 3/47/0 unchanged.
- core test --suite smoke -q: 67 passed.
- uv run core eval refusal_taxonomy: case_digest
d030f826cb0f4088771d90c52c8be2ff75054ab27c7d47eae8dbfe1225b2eea1
unchanged.
Cross-refs: ADR-0163 (Phase C), ADR-0057 (gating discipline),
ADR-0151 (auto-proposal), ADR-0152 (learning-arc), ADR-0149/0154
(recognizer pipeline), ADR-0094 (ProposalSource), Phase A PR #297,
Phase B PR #298.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* docs(math): ADR-0163 — path to GSM8K mastery via candidate-graph admissibility (proposed)
Audit reframes the math roadmap entirely.
State of main: every named math capability axis (G1..G5, S1) passes
at 100% with wrong=0 on its controlled lane. binding_graph,
math_versor_arithmetic, math_symbolic_equivalence, math_parser,
math_candidate_parser, math_solver, math_verifier, math_realizer,
math_problem_graph — all landed. The worktrees on disk are stale
forks.
State of GSM8K (50-case train sample): correct=0, refused=50, wrong=0.
Every refusal reason is identical: "candidate_graph: no admissible
candidate for statement: <STATEMENT>".
The reframe: the gap is NOT in operator algebra, NOT in binding graph
internals, NOT in symbolic equivalence. The gap is in
generate/math_candidate_graph.py — the admissibility surface that
turns a natural-language statement into a candidate the downstream
pipeline can consume. The capability axes pass at 100% because they
test statement shapes the candidate-graph already admits. GSM8K
refuses at 100% because its statements span shapes the candidate-graph
has never been taught.
Six-phase plan to lift GSM8K under the thesis "decodes, not generates":
A. Refusal taxonomy (measure before building)
B. Exemplar corpora per shape category (≤20 statements each, ≤3 per round)
C. Contemplation runner ingests exemplars; emits DerivedRecognizer
proposals
D. Operator ratifies through ADR-0161 HITL queue (no new surface)
E. Re-baseline GSM8K train sample. Round 1 exit: correct ≥ 10, wrong = 0.
Round 2: ≥ 25. Round 3: ≥ 35.
F. Scale to public/v1 (200 cases, target correct ≥ 100), then
holdout (measurement-only — never tune against).
Three non-negotiables:
- wrong = 0 at every phase. Auto-rejected by replay gate, not by
operator vigilance.
- No hand-rolled recognizers in generate/. Every recognizer lands
via contemplation → proposal → review corridor.
- Active corpus mutation only via accept_proposal.
Status: proposed. Implementation lands as three PRs starting with
Phase A scaffolding.
Scope discipline: docs-only. No code, no eval changes, no corpus
mutation.
* feat(ADR-0161.1): core teaching queue list|show — read-only queue projection
* fix(ADR-0161.1): restore gap-queue CLI + rename new commands to hitl-queue + R1..R5 refinements
ADR-0163 Phase A measurement. Reads the GSM8K train-sample refusal report
(50 cases, all refused on candidate-graph admissibility) and emits a
histogram of statement shapes. Read-only: no corpus, pack, or proposal
mutation; the categorizer is rules-only with no LLM, embedding, or
learned model.
Lane: evals/refusal_taxonomy/ (auto-discovered by evals.framework)
- shape_categories.py — ShapeCategory enum + deterministic categorizer
(9 ADR-mandated baseline categories + UNCATEGORIZED, first-match-wins)
- runner.py — pure run_lane(cases) -> LaneReport
- contract.md — purpose, doctrine, schema, ADR compatibility
- public/v1/cases.jsonl — 50 refused statements (sorted by case_id)
- v1/report.json — first run output (categorized_rate=72%)
CLI: core teaching refusal-taxonomy [--input PATH] [--json] [--save]
Accepts a cases JSONL or a raw GSM8K eval report.json directly.
Helper: scripts/build_refusal_taxonomy_cases.py rebuilds the v1 case set
from the GSM8K train-sample report deterministically.
Tests: tests/test_refusal_taxonomy_lane.py (21 passing) cover schema
integrity, lane auto-discovery, enum exhaustiveness, categorizer
determinism + purity + no-ML-imports, histogram correctness, replay
byte-identity, committed report match, helper extraction, and a
read-only invariant snapshot over teaching/, packs/, language_packs/data/.
v1 histogram (50-case sample):
17 descriptive_setup_no_quantity
14 uncategorized
4 temporal_aggregation
3 rate_with_currency
3 fractional_rate_of_change
3 indefinite_quantity
3 comparative_with_unit
2 nested_question_target
1 unit_partition
0 conditional_quantity
total=50 categorized_rate=72% uncategorized=28% (below 50% target)
Top three by count (Phase B candidates):
1. descriptive_setup_no_quantity (17)
2. temporal_aggregation (4)
3. tie at 3 — operator selects from {rate_with_currency,
fractional_rate_of_change, indefinite_quantity, comparative_with_unit}
Phase B is not started in this PR — the ADR explicitly requires the
operator to ratify the top-N selection before any exemplar corpus is
authored.
Invariants verified:
- tests/test_adr_0131_*.py: 224 passed, 0 wrong on G1..G5 + S1
- core test --suite smoke -q: 67 passed
- The refusal_taxonomy/__init__.py and runner do not import openai,
anthropic, transformers, torch, sklearn, sentence_transformers,
requests, or httpx — verified by test_categorizer_no_llm_or_ml_imports.
Cross-references: ADR-0163 (parent), ADR-0114a (capability obligations),
ADR-0149 (recognizer pipeline substrate that Phases C–E build on).
Refs: [[thesis-decoding-not-generating]] — the rules-only categorizer
honors the doctrine: the engine learns to find better shapes; this PR
does not stuff it with another found pattern.
Three follow-ups raised in the W-025 PR #286 review, completed together so
the lane reaches its full mastery-level contract.
1. ``core eval`` failure-printer is now gated on ``lane_name == "cognition"``.
Before this fix, every non-cognition lane that returned clean case_details
without ``intent_correct``/``versor_closure`` keys triggered a spurious
``failures (N): <case_id>: intent, versor=0.00e+00`` block at the end of
the human-readable output, even when every metric passed. This matched
the gating pattern already used for the workers preamble at the top of
``cmd_eval``.
2. EPILOG examples in ``core/cli.py`` now advertise
``core eval contemplation_quality`` and the ``--json --save`` form, so
the lane is discoverable from ``core --help`` and not only from
``core eval --list``.
3. Tightened the learning-arc demo's Scene 5 to thread the demo's
tempdir-scoped ``engine_state_dir`` into the second ``ChatRuntime``.
The previous default-constructed runtime checkpointed to the repo's
``engine_state/``, which contradicted ADR-0159's read-only claim.
ADR-0146/0150 still govern the runtime checkpoint path itself.
Tests:
- ``tests/test_contemplation_quality_lane.py`` (35 tests):
case-set integrity, lane discovery, ``evaluate_report`` purity over
well-formed / malformed / boundary-violating inputs, ``run_lane``
invocation-contract enforcement (single case, supported source enum),
and a read-only invariant snapshot on ``teaching/corpora``, ``packs/``,
and ``language_packs/data/``.
- ``tests/test_eval_cli_failure_printer.py`` (4 tests): pins the
cognition-only gating of the failure printer with stubbed
``evals.framework`` so the regression cannot return as a lane-blind
condition.
Validation:
uv run pytest tests/test_contemplation_quality_lane.py \
tests/test_eval_cli_failure_printer.py \
tests/test_learning_arc_demo.py -q # 50 passed
uv run core test --suite smoke -q # 67 passed
uv run core eval contemplation_quality # 9/9 passed, clean output
Two-session arc where engine derives connective+object from corpus
decomposition; operator ratifies rather than authors. Distinguishes
from learning-loop (operator-authored) and directly exercises W-018
checkpoint contemplation and W-017 auto-proposal provenance path.
Closes W-013 wiring debt. Per Phase 2 operator decision: wire
core.cognition.explain into the live core chat REPL.
Changes:
- core/cognition/explain.py: add explain_from_intent(intent, correction_text)
companion to explain() — same dispatch table, skips the full
CognitiveTurnResult round-trip. Callers with only a DialogueIntent can
use this directly.
- chat/runtime.py: add _last_intent and _last_input_text instance fields;
store intent on every classify_intent_from_input() call (pack-grounded
path and stub/empty-vault path); add explain_last_turn() -> str method
that calls explain_from_intent(_last_intent, correction_text=_last_input_text).
- core/cli.py: in cmd_chat REPL loop, handle "/explain" command — calls
runtime.explain_last_turn() and prints the canonical prompt restatement
(or a "no prior turn" message to stderr if no turn has run yet).
- tests/test_explain_repl.py: 11 tests pinning explain_from_intent dispatch
for all intent tags and the ChatRuntime.explain_last_turn() contract.
Per ADR-0017 (Responsive-with-Axiology): introspection is per-turn and
operator-invoked, never autonomous — the /explain command is correct
placement for this feature.
Final wire-up after all 10 ADR-0114a obligations + ADR-0131.4
composite gate landed. Composes:
- all 10 obligation verdicts (5 from new auditor modules,
5 from inline checks over existing infrastructure)
- ADR-0131.4 composite math gate verdict
- ADR-0092 reviewer-signed claim entry from docs/reviewers.yaml
into a single deterministic promotion verdict + canonical
signed/unsigned ``expert_claims_math_v1_signed.json`` artifact.
Empirical verdict on current main (first evaluation):
all_obligations_passed: True
composite_gate_passed: True
technical_pass: True
claim_digest: d164866975341d9b82503caf50c0404ee140eab21fd60f589536c6daf6e1d706
reviewer_signature_present: False
promote_admitted: False
refusal_reason: awaiting reviewer signature
Every technical gate passes. The PR ships in the architecturally-
correct "awaiting reviewer signature" state — the reviewer's
signature is the separate, auditable operator action that
consummates the promotion.
Operator workflow (post-merge):
1. Run `core capability math-expert-promote`, confirm verdict,
capture claim_digest.
2. Add entry to docs/reviewers.yaml under math_expert_claims:
- domain_id: mathematics_logic
signed_by: shay-j
claim_digest: "d164866975341d9b82503caf50c0404ee140eab21fd60f589536c6daf6e1d706"
3. Re-run — promote_admitted flips to True.
4. Separate ledger-flip PR (out of scope here) consumes the
signed artifact and writes the capability ledger.
Safety property: if the evidence bundle changes after signing
(B-lane re-run, pack edit, obligation report shift), the digest
changes and the existing signature stops matching. The verdict
reports the mismatch explicitly and the operator must re-inspect
and re-sign — a ledger flip can't survive a silent evidence change.
New files:
- core/capability/expert_promotion_math.py — the composer
- tests/test_adr_0120_math_expert_promotion.py — 18 tests
- docs/decisions/ADR-0120-math-expert-promotion-wireup.md — ADR
Modified:
- core/cli.py — new `core capability math-expert-promote` cmd
- docs/reviewers.yaml — added math_expert_claims: [] section
with documentation comment
Tests: 18/18 covering each inline obligation evaluator
(#1/#3/#4/#7/#9 pass + failure modes), composer integration
against current main, reviewer-signature path (matching → admitted;
mismatched → refused with explicit diagnostic), digest
reproducibility, artifact byte-equality. All pass in 0.49s.
Trust boundary: read-only access to 4 B-lane reports +
GSM8K probe + 5 obligation auditor reports (transitively) +
frontier dir + docs/reviewers.yaml; single deterministic write
to the artifact path; no dynamic imports, no shell, no network.
This is the last PR before the first mathematics_logic -> expert
ledger flip attempt. The actual flip is reserved for a separate
small PR that consumes the signed artifact.
35-case OOD set (ood-001..ood-035): surface-varied siblings of B3's 35
solved_correct public cases. Entity-name pool: Maya/Liam/Noah/Diana/Felix/
Priya/Omar/Rosa/Jun/Kai. Unit-noun pool: oranges/marbles/pencils/books/
stamps/coins/balls (all parser-allowed count nouns). Every case in-grammar
per ADR-0131.3 and parseable without error.
Auditor (core/capability/ood_ratio.py): reads B3 public report.json + OOD
report.json, computes ood_ratio = ood_accuracy / public_accuracy, enforces
two independent gates — ratio ≥ 0.95 and wrong == 0.
CLI: core capability ood-ratio (exit 0 iff both gates pass).
Measured: public 50/50=1.000, OOD 35/35=1.000, ratio=1.000. Obligation #10
and B3 public lane unchanged.
Implements the external auditor for ADR-0114a Obligation #6:
"depth_curve.py produces a per-bucket curve;
accuracy(N) >= accuracy(depth_1) * (1 - eps)^(N - 1) for eps = 0.05."
Mirrors PR #189's auditor pattern (re-runs lane via the candidate-
graph pipeline, aggregates over committed cases, emits deterministic
report). Uses len(trace.steps) as the authoritative depth — the
engine's actually-executed reasoning, not the case's declared depth.
New module core/capability/depth_curve.py:
- Bucket schema mirrors ADR-0119.6: depth_1, depth_2-3,
depth_4-5, depth_6-8. Depth > 8 raises rather than silently
extending. Depth == 0 (initial-only problems) skipped — nothing
to decay.
- representative_depth = min(bucket) — most permissive bound
convention; tightening requires an ADR amendment.
- epsilon = 0.05 pinned per ADR-0120 §Threshold rationale.
- Two-axis verdict: obligation_6_mechanism_wired (always true if
auditor ran), obligation_6_assertion_holds (every populated
bucket satisfies the decay bound), coverage_sufficient (>=2
buckets populated AND >=3 cases each — required for the
assertion to be statistically meaningful).
CLI: core capability depth-curve (added to core/cli.py).
Writes evals/obligation_6_depth_curve/<lane_id>.json.
Empirical verdict on current main:
lane: B3_bounded_grammar
cases_total: 50
cases_solved: 22
mechanism_wired: True
assertion_holds: True
coverage_sufficient: False
populated: [depth_1 (21/21=1.0000), depth_2-3 (1/1=1.0000)]
Both populated buckets satisfy the decay bound. Coverage gap is
honestly named in the refusal_reason: depth_2-3 has only 1 case,
depth_4-5 and depth_6-8 have none. This is B3-owner work (case
authoring under the existing grammar contract), not auditor work;
reserved as a B3 v1.1 follow-up PR.
Honest scope-limit: B3 only. B1 (algebra, no trace) and B2 (chain
validation, not problem-solving) need different metrics — separate
sub-ADRs.
Trust boundary: read-only access to B3 cases + transitive pack
reads via the pipeline; single deterministic write to artifact path.
Tests: 24/24 covering bucket schema closure (depth 1..8 + raise on
9+), decay bound math (epsilon pinned, formula correct, depth_1 has
no bound), coverage-sufficient policy (thresholds pinned), lane
evaluation (passes on real B3 + refuses on missing cases),
coverage-sufficient distinction (B3 today vs synthetic 5+5 fixture
showing both pass), determinism (report identical + artifact
byte-equal).
External auditor for ADR-0114a Obligation #8:
"adversarial/score.py reports wrong == 0 across all families;
>= 30 cases x >= 8 families."
Verdict on current main:
cases_total: 36
families_total: 9
cases_refused: 28
cases_solved: 8
cases_wrong: 0 <-- the gate
obligation_8_passed: True
New module core/capability/adversarial.py mirrors PR #189/#190/#191
auditor pattern. Pure function over the committed cases set; broad
exception capture (correctly classified as refused — engine
couldn't process the input) makes the auditor robust to upstream
typed-refusal gaps.
New dataset evals/obligation_8_adversarial/v1/cases.jsonl — 36
cases x 9 families, closed taxonomy:
- paraphrase (verb outside initial-anchor whitelist)
- unrecognized_unit (not in en_units_v1)
- conditional (if/would/suppose)
- pronoun_coref (cross-sentence he/she/they)
- hedged_quantity (about/almost/approximately)
- ordinal_confusion (the 5th/third in cardinal position)
- implicit_subject (no named entity)
- self_reference (actor as comparison ref or transfer target)
- distractor_noise (adjectival/temporal/irrelevant siblings)
CLI: core capability adversarial. Writes
evals/obligation_8_adversarial/<lane_id>.json. Exit 0 iff
obligation passes.
Honest disclosure — 8 of 36 cases solved rather than refused;
none produced wrong answers. Two parser-layer gaps surfaced:
Gap A (pronoun_coref, 4/4 solved): unbound sibling sentences
silently drop; engine returns last-asserted state. Faithful but
semantically poor. Reserved follow-up: tighten admissibility so
unbound sentences refuse the whole case.
Gap B (unrecognized_unit, 4/4 solved): _canonicalize_unit
falls back to '+s' plural rule when pack doesn't recognize
the unit. Reserved follow-up: opt-in strict mode behind a flag
(some B3 units aren't in en_units_v1 either; strict mode
requires parallel pack extension).
Bug caught: adv-self-reference-003 ("Sam gives 3 apples to
Sam.") raises uncaught MathGraphError from
Operation.__post_init__. Auditor catches it as
refused-via-exception; ~3-line follow-up in
_build_op_candidate fixes the parser side.
Trust boundary: read-only access to cases + transitive pack reads;
single deterministic write to artifact path.
Tests: 11/11 in tests/test_adr_0114a_8_adversarial.py covering
threshold pinning (>= 30 cases / >= 8 families), closed taxonomy
(every documented family has cases; no unknown families),
obligation-passes snapshot, per-family wrong=0 invariant, failure
modes (missing file, below-threshold count), determinism (report
identical + artifact byte-equal).
Implements the external auditor ADR-0114a Obligation #10 requires:
"Every SolutionTrace.steps[*].pack_lemma_id resolves to a real
lexicon entry in the domain's operator pack." The solver enforces
this at solve time; this PR audits it from outside.
New module core/capability/pack_provenance.py:
- _load_lexicon_lemmas(): independent re-read of pack lexicon
- _parse_lemma_id(): <pack_id>:<lemma> shape parser
- validate_lane(): re-runs candidate-graph pipeline on a B-lane's
cases, walks every solver step, validates pack_lemma_id parses
AND resolves to a lexicon entry. Per-case + per-lane verdict.
- emit_provenance_report(): deterministic artifact emission.
CLI: core capability pack-provenance (added to core/cli.py).
Writes evals/obligation_10_pack_provenance/<lane_id>.json.
Empirical verdict on current main (post-PR #186):
lane: B3_bounded_grammar
cases_total: 50
cases_validated: 25 (every expected-correct B3 case)
cases_skipped_unsolved: 25 (refusal-expected probes — by design)
cases_violated: 0
obligation_10_passed: True
5 distinct lemma_ids observed (add, subtract, transfer,
compare_additive, compare_multiplicative) — all resolve to
en_arithmetic_v1. The other 3 op kinds (multiply, divide,
apply_rate) ratify-at-solve-time via _resolve_pack_lemmas so the
obligation holds for them too if a future case exercises them.
Honest scope-limit: B3 only. B1 (symbolic equivalence) and B2
(teaching corpus) equivalents deferred to separate sub-ADRs —
B1 needs reframing (algebra normalization chain, not arithmetic
steps); B2 can use this same auditor signature once corpus
solver-trace exercise is confirmed case-by-case.
Composition with ADR-0131.4: orthogonal. Composite gate verdict
+ obligation #10 verdict + 4 other obligation auditors (when
they land) + reviewer signature → full ADR-0120 wire-up.
Trust boundary: read-only access to pack lexicon + B3 cases;
single deterministic write to artifact path. No dynamic imports,
no shell passthrough, no network. Pure deterministic auditor.
Tests: 19/19 in tests/test_adr_0114a_10_pack_provenance.py
covering lemma-id parser (well-formed + malformed), lexicon loader
(real pack + every failure mode), lane validator (passes on real
B3 + refuses on missing pack/cases + skips refusal-expected cases
without false violation), determinism (report identical across
calls + artifact byte-equal).
Implements ADR-0131's revision of the ADR-0120 expert-promotion
contract for mathematics_logic: replaces the single-benchmark
GSM8K-coverage check with a composite B1+B2+B3 requirement.
New module core/capability/composite_math_gate.py:
- evaluate_composite_math_gate(): pure function over already-
committed B-lane reports; handles heterogeneous report shapes
(B1/B2 counts vs B3 metrics); applies pinned thresholds
(correct_rate >= 0.95 AND wrong == 0); composes verdicts.
- Reproducible SHA-256 claim_digest over canonical evidence bundle.
- GSM8K honest-disclosure (admission/wrong/refused/substrate)
embedded in artifact but never gates per ADR-0131.
CLI: core capability math-expert-gate (added to core/cli.py).
Writes evals/math_expert_claims/v1/expert_claims_math_v1.json.
Empirical verdict on current main (post-PR #182/#183/#184/#185):
composite_gate_passed: True
B1_public: 185/185 wrong=0 rate=1.0000
B1_sealed: 14/14 wrong=0 rate=1.0000
B2_teaching_corpus: 40/40 wrong=0 rate=1.0000
B3_bounded_grammar: 50/50 wrong=0 rate=1.0000
GSM8K disclosure: 0/50 admission, wrong=0, substrate=candidate_graph
The math expert is gate-passing under ADR-0131's revised composite
contract. The architectural bet ADR-0131 placed has paid off.
Honest scope-limit: this implements only the ADR-0131-specific
revision (composite benchmark portion). The full ADR-0120 10-
obligation contract still requires substrate for 5 missing
obligations (OOD ratio, perturbation, depth curve, adversarial,
operation-provenance-via-pack). Those are sequencing-wise *after*
ADR-0131.4, not bundled. Reviewer signature via ADR-0092 registry
is also reserved.
Trust boundary: read-only access to 5 committed lane reports;
single deterministic write to the artifact path. No dynamic
imports, no recomputation of lane verdicts.
Tests: 12/12 in tests/test_adr_0131_4_composite_math_gate.py
covering threshold pinning, heterogeneous shape handling, gate
logic (passing + every failure mode), GSM8K honest disclosure
(never gates), determinism (claim_digest + artifact byte-equality),
and a snapshot test confirming current main satisfies the gate.
ADR-0131.4 module note: the parent ADR-0131 plan named
formation/ratify.py + formation/promote.py as the wire-up site —
that was a misidentification (those govern teaching-example
SPECULATIVE→COHERENT bridging per ADR-0021, not domain-tier
promotion). Correct site is core/capability/, where audit-passed
gate already lives.
Wraps existing math pipeline (parser -> solver -> verifier) against
PR #159's 50-case train sample. Emits deterministic report.json with
per-case verdicts. CLI exit code reflects exit criterion
(correct >= 10 AND wrong == 0).
Baseline against current parser: 0 correct / 0 wrong / 50 refused.
This baseline is the inner-loop gradient signal for ADR-0126's
candidate-graph parser (in flight on feat/adr-0126-candidate-graph).
Registers tests/test_adr_0126_train_sample_runner.py under
'core test --suite math' so the wrong == 0 invariant becomes a hard
CI gate per ADR-0114a Obligation #4 (refuse rather than confabulate).
Depends on PR #159 (gemini/adr-0126-train-sample). Rebase onto main
after #159 lands.
The word "expert" in the previous status name implied raw-capability parity
with frontier LLMs on the same benchmark — which the gate does NOT verify.
What the gate actually verifies is CORE *claim-shape compliance*:
* signed digest (replay-reproducible from on-disk lane results)
* replay determinism (same inputs → byte-equal trace_hash)
* typed refusal (fabrication refused, not paraphrased)
* exact recall (no ANN, no cosine, no attention bottleneck)
* grounding-source provenance
These are claim shapes a transformer LLM cannot structurally produce
regardless of raw accuracy. A frontier LLM might score higher on the
same benchmark but cannot pass this contract.
Rename scope (semantics only, per ADR-0113):
status string "expert-demo" → "audit-passed"
predicate key predicates.expert_demo → predicates.audit_passed
reason key expert_demo_reason → audit_passed_reason
YAML key expert_demo_claims → audit_passed_claims
CLI command core demo expert → core demo audit-passed
output dir evals/expert_demos/ → evals/audit_passed/
artifact filenames expert_demo.{json,html} → audit_passed.{json,html}
HTML title CORE Expert-Demo: X → CORE Audit-Passed: X
Internal Python identifiers (module/file/function/class names like
`expert_demo.py`, `evaluate_expert_demo`, `ExpertDemoClaim`,
`expert_demo_claim_for`) are deliberately kept to minimize churn. ADR
file titles (ADR-0106..0112) preserved as historical record.
`expert` namespace reserved for ADR-0114+: an actual capability tier
above `audit-passed` backed by a public benchmark with a stated
threshold. ADR-0114 proposes the first such target — GSM8K-math —
laying out a falsifiable 7-phase arc (parser → solver → verifier →
stepped-realizer → eval lane → first `expert` ledger tier promotion).
Tests: 184 directly-affected tests green (140 capability/expert-demo
suite + 34 demo/audit-tour + 10 correction-cue). Smoke suite 67/67.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Closes the asymmetry between the `expert-demo` ledger status (audit
artifact only) and the actual `core demo` surface (runnable
walkthroughs producing HTML + JSON). Until this commit the word
"demo" in `expert-demo` was aspirational; now it corresponds to
something a reader can open.
What it does
- Reads the signed expert_demo_claims entry from docs/reviewers.yaml
- Loads latest on-disk result files for each attached lane × split
- Re-derives the evidence-bundle digest and asserts byte-for-byte
match against the signed claim_digest — this is the load-bearing
audit step, now exercised at two independent enforcement points
(ledger gate + showcase)
- Runs each lane's metrics through the ADR-0109 lane-shape registry
and surfaces the verdict
- Picks the first three cases from each split verbatim (deterministic
by file order) and renders them as HTML for inspection
- Emits expert_demo.json (canonical bytes, deterministic) + expert_demo.html
Surface
core demo expert --domain mathematics_logic
core demo expert --domain physics
# → evals/expert_demos/<domain>/latest/expert_demo.{json,html}
Read-only by construction: cannot mutate docs/reviewers.yaml or any
lane result file. Tested. Unpromoted domains raise ValueError —
no silent fallback, no "preview" mode that fakes a showcase.
Generated artifacts are gitignored — the inputs they derive from are
already committed, so duplicating the renders would just churn the
tree.
Tests: 16 new cases pinning all five ADR-0112 invariants. Smoke suite
still 67/67 green.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Single 30-second artifact composing four CORE invariants
(determinism, honest unknown, reviewed learning, multi-hop with
trace) by delegating to existing DemoCommand adapters. **No new
mechanism** — every claim is backed by an already-shipped,
separately-tested adapter. Closes the 8-ADR scale-up slate.
- new core/demos/learning_loop_adapter.py: LearningLoopDemo wraps
ADR-0056 reviewed-teaching loop; _strip_volatile_paths drops
transient temp-dir paths from raw before serialization so the
adapter's report_sha256 is content-stable across runs
- new core/demos/showcase_adapters.py:
- FabricationControlPublicDemo: re-runs ADR-0096 public split,
produces 3 claims (refusal_recall_meets_threshold,
fabrication_rate_below_threshold, trace_evidence_present)
- MultiHopTraceDemo: runs 'Does light reveal truth?' with
transitive_surface=True + composed_surface=True against
cognition pack; surfaces a 3-hop walk light→truth→knowledge→
evidence; produces 3 claims (grounded_answer, depth_two_or_more,
walk_evidence_present)
- new core/demos/showcase.py: run_showcase() composes 4 scenes,
emits showcase.json + per-scene artifacts; render_html() produces
presentation-only static HTML with no JS injection vector;
ShowcaseScene dataclass; MAX_RUNTIME_SECONDS=30 hard ceiling
with DemoContractError if exceeded
- CLI: 'showcase' added to demo target choices; --output-dir flag
added; cmd_demo dispatch branch writes showcase.json + showcase.html
- new evals/public_demo/ lane with 4 cases:
- all_claims_supported (each scene + composite)
- determinism_run_to_run_byte_equality (two runs identical after
stripping volatile keys: total_runtime_ms, json_path,
transient_corpus)
- runtime_under_budget (≤30s)
- pure_composition_no_new_mechanism (grep gate over showcase
imports — must come from core/chat/generate/language_packs/
teaching/evals or allowed stdlib only)
- lane is itself byte-identical across runs (sha256 5707db8efc6a..);
runtime case omits exact runtime_ms (it varies near bucket
boundaries) but still asserts ≤ budget
- 8 unit tests with module-scoped fixture (showcase runs once,
~13s total) covering payload shape, scene order, runtime budget,
HTML render absence of <script>, and the pure-composition import
gate independently of the lane
- ADR-0099 measured: total_runtime_ms ~12.8s, well under 30s budget
- smoke 67/67, cognition eval byte-identical 100/100/100/100;
all 6 ADR-0092..0099 lanes byte-identical:
reviewer_registry 681a2aab..
miner_loop_closure 9f071733..
domain_contract_validation f9c06cde..
fabrication_control sum 01e1b6b7..
demo_composition 27d83824..
public_demo 5707db8e..