core/evals
Shay cfda71dc1c feat(evals): grammar round-trip instrument (Phase 1)
The measurement foundation for docs/plans/grammar-unification-2026-07-26.md.

WHY: evals/deterministic_fluency reports 1.00 on all six predicates and
still passes "banana does the.", "wet ground rains the is." and
"is is is is." — it checks terminal punctuation, presence of a verb-shaped
token, and two anti-shape regexes. Heuristic predicates will always have
that failure mode, because grammaticality cannot be measured without a
grammar. So this lane measures agreement between the two halves of CORE
that already encode grammar, and requires the measurement to FAIL on salad.

Two directions, reported separately because they fail for different
reasons and have different remedies:

  G-round-trip  graph -> realize_target -> surface -> comprehend -> graph
  S-round-trip  surface -> comprehend -> graph -> categorical renderer -> surface

v1 baseline on main @ 9696443a:

  graph_cases            280      surface_cases            8
  g_write_rate         1.000      s_read_rate          1.000
  g_read_rate          0.000      s_renderable_rate    0.625
  g_exact_rate         0.000      s_surface_match_rate 0.000
                                  negative_cases          16
                                  reject_rate          1.000

g_read_rate and s_surface_match_rate are pins on measured DEFECTS, not
goals; they may be revised upward only. s_surface_match_rate = 0 is the
§1.7 categorical render defect caught by construction — the lane found it
without being told to look.

g_args_rate and g_predicates_rate are deliberately separate: high argument
agreement with low predicate agreement would mean the grammars align and
only the vocabulary is split, a materially different remedy from both
being low. That distinction decides the arc's direction (plan §6).

Design notes:
- The committed cases.jsonl is the SINGLE source for authored surfaces —
  no in-module duplicate, since a second copy of a corpus is the defect
  this arc exists to remove. Negative shuffles are DERIVED at run time so
  they cannot drift from the positives.
- The shuffles are lexically identical to positives (same vocabulary, same
  length, order destroyed) so the lane cannot pass by vocabulary-checking.
- Fixed rotation, not a PRNG, so reject_rate is byte-reproducible.
- The lane keeps a local copy of the reader's quantifier map ON PURPOSE so
  it never becomes a consumer of what it measures;
  test_quantifier_map_matches_reader fails loudly if the reader changes.
- _render_categorical deliberately reaches a private serving function: a
  lane measuring a private copy would measure what users never see.

Every guarantee is paired with a mutation test. The load-bearing one is
test_reject_rate_goes_red_when_the_reader_accepts_everything: an
accept-everything reader must drive reject_rate to 0.0. Without it,
reject_rate == 1.0 would be unfalsifiable — precisely the defect that
makes the existing fluency lane decoration.

Also documents plainly what round-trip does NOT prove: it measures mutual
intelligibility, not English quality. english_fluency_ood accepts "river
flows valley" and round-trip would be happy with it. No metric here may be
cited as evidence of prose quality.

scripts/measure_grammar_seam.py reproduces every number in the plan's §1
so a reader can check them instead of trusting them.

[Verification]: in-worktree on CPython 3.12.13, uv sync --locked —
smoke 621 (unchanged), deductive 364 (349 + 15 new). Lane SHA pins
verified separately. No serving code touched; new files plus one suite
registration line.
2026-07-26 16:30:49 -07:00
..
adr_0242_v2_energy_compare feat: close ADR-0241/0242 post-Accept backlog (local-first mastery) 2026-07-16 15:25:43 -07:00
adr_0243_cognitive_lifecycle feat(adr-0243): Phase 4 falsifiability benchmark — metrics eval + CLI dispatcher 2026-07-17 12:59:24 -07:00
adr_0244_gamma_calibration fix(tests): clear the last reds — a machine-specific path and a stale drift pin 2026-07-25 16:31:06 +00:00
adr_0244_identity_gate fix(cognition): pre-stage land repairs after L2/PASSTHROUGH excision 2026-07-20 13:34:59 -07:00
adr_0244_qtop_vacuity feat(adr-0244): prove Q_top vacuity — D4 hollow-gate evidence + annotation upgrade 2026-07-17 14:39:52 -07:00
adr_0246_discrimination feat(adr-0246): Opus audit + §3.7 admit surface + serve wiring + §6.3 discrimination 2026-07-17 22:51:46 -07:00
adr_0246_geometric_suite draft(adr-0246): slice-1 scaffold — §6.1/§6.2 eval suite + malformed-F guard 2026-07-17 22:09:32 -07:00
adr_0246_grounding_feasibility feat(adr-0246): completion — §4.1 records, path serve integration, §11 feasibility (honest NULL), ADR body Proposed 2026-07-17 23:36:41 -07:00
adr_0246_mismatch_diagnostic feat(adr-0246): §3 induced-action primitives (pure, off-serving) 2026-07-17 21:36:57 -07:00
adversarial_identity
analogical_transfer feat(adr-0240/0243): seam S2 — chiral-composed, harness-driven biography write-path 2026-07-18 10:15:38 -07:00
anchor_lens_tour
anti_regression feat: strengthen visibility and measurement of CLOSE flywheel proposal review/ratification side (#794) 2026-06-16 19:33:25 -07:00
articulation
articulation_of_status
audio_sensorium feat(adr-0181-p4): audio compiler eval gate lane (sensorium/audio) (#470) 2026-05-29 11:20:31 -07:00
audit_tour chore: Refactor CLI and Governance Anchors (#926) 2026-07-03 12:34:56 -07:00
calibration chore: Refactor CLI and Governance Anchors (#926) 2026-07-03 12:34:56 -07:00
capability_index chore(generalization): Tier S housekeeping — smoke promotion, capability-index entries, promotion sweep (S5) 2026-07-24 17:13:19 -07:00
classical_literature_ood
close_derived_climb chore: Refactor CLI and Governance Anchors (#926) 2026-07-03 12:34:56 -07:00
cognition fix(cognition): expose hash_surface; point the register-matrix pin at it 2026-07-25 14:06:36 +00:00
cold_start_grounding fix(quarantine): drain all 60 quarantined tests — QUARANTINE=∅ (#267) 2026-05-25 11:22:12 -07:00
combined_rate_oracle docs(cmb): lookback review + fix single-agent-attribution hygiene hazard (H3) 2026-06-08 13:27:17 -07:00
compositionality
compound_intent_decomposition
comprehension feat: one-hop sound relational entailment (inverse/symmetric) + capability-index lane (#775) 2026-06-15 11:39:41 -07:00
constraint_oracle feat(verified): P1-C — split bound_slots_digest into separate obligation 2026-06-08 17:17:31 -07:00
contemplation_quality feat(W-025): contemplation quality eval lane (ADR-0159) (#286) 2026-05-25 20:38:52 -07:00
contradiction_detection
conversation
conversational_thread_coherence
cross_domain_transfer
curriculum_loop_closure feat(deduction-serve): Band v6-EX — existential witnesses, decided (ADR-0261) 2026-07-24 14:14:37 -07:00
curriculum_serve feat(curriculum): taught negatives — polarity is read (ADR-0264 R1-R4, R8) 2026-07-26 12:19:19 -07:00
deduction_serve fix(deduction-serve): resolve assert_corpus_sound name collision 2026-07-25 16:47:01 -07:00
deductive_logic test(l10): add independent-gold adversarial logic fixtures 2026-06-05 09:08:23 -07:00
demo_composition fix(ci): path-stable env deltas so lane SHA pins stop thrashing 2026-07-14 17:41:04 -07:00
determination_closure feat(determine): idle deductive consolidation — the loop learns from determined facts (Step D) 2026-06-06 12:28:09 -07:00
determination_estimation Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
deterministic_fluency
dimensional feat: dimensional-reasoning lane — 3rd diversity-panel domain 2026-06-04 16:38:56 -07:00
discourse_paragraph
domain_contract_validation chore(ci): re-pin drifted lane SHAs + refresh canonical reports (#229) 2026-05-24 14:25:11 -07:00
edge_budget chore: Refactor CLI and Governance Anchors (#926) 2026-07-03 12:34:56 -07:00
elementary_mathematics_ood
english_fluency_ood
environment_falsification Add offline witness log importer 2026-06-06 12:37:57 -07:00
event_vision_sensorium Add event vision sensorium lane 2026-06-06 12:37:57 -07:00
fabrication_control chore(ci): re-pin drifted lane SHAs + refresh canonical reports (#229) 2026-05-24 14:25:11 -07:00
field_incidence work done towards furthering comprehension, contemplation, learning, and also began working on fixing some failing tests 2026-06-16 12:27:59 -07:00
flywheel_demo Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
forward_semantic_control Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
foundational_biology_ood
foundational_physics_ood
frame_verdict_text_cwa feat(frame-verdict): closed-world FrameVerdict substrate — PR-1..4 + hardening (ADR-0222 B4) (#787) 2026-06-16 06:23:03 -07:00
frontier_compare
generalization feat(evals): implement safe metadata and fail-closed evaluator check for GSM1K 2026-06-23 07:52:32 -07:00
grammar_roundtrip feat(evals): grammar round-trip instrument (Phase 1) 2026-07-26 16:30:49 -07:00
grammatical_coverage
gsm8k_math test(reader-arc): pin deterministic tune/measure split BEFORE grammar work 2026-07-18 16:39:04 -07:00
gsm8k_parser_dev feat: add ADR-0125 perturbation suite 2026-05-22 17:12:33 -07:00
hebrew_fluency
identity_divergence
industry_demos Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
inference_closure
introspection
koine_greek_fluency
l10_always_on feat(identity): split engine identity from build provenance (ADR-0220 PR C) (#774) 2026-06-15 11:38:04 -07:00
l10_continuity test(l10): W2-R arbitrary-interruption recovery harness (ADR-0219) 2026-06-15 02:45:29 -07:00
lab
learning_arc feat(workbench): add UI catch-up evidence scaffolding (#891) 2026-06-23 10:09:45 -07:00
learning_loop fix(ci): re-pin public_demo lane SHA after showcase content drift (#807) 2026-06-17 15:42:40 -07:00
long_context_cost
math_bounded_grammar/v1 feat(ADR-0131.3): bounded-grammar word-problem benchmark — lane PASSED 50/50 (#180) 2026-05-23 11:27:04 -07:00
math_capability_axes chore(evals): refresh stale committed reports 2026-06-03 01:01:50 -07:00
math_expert_claims/v1 docs(claims): ADR-0200 reconciliation — expert claim to audit-passed truth 2026-06-02 10:06:16 -07:00
math_symbolic_equivalence feat(ADR-0131.1.F): frontier-baseline comparison harness for B1 (#178) 2026-05-23 12:14:06 -07:00
math_teaching_corpus/v1 feat(ADR-0131.2.B): B2 teaching-corpus enrichment — load-bearing gate (#177) 2026-05-23 11:29:48 -07:00
miner_loop_closure
monotonic_learning
multi_agent_composition
multi_sentence_response
multi_step_reasoning
obligation_2_ood_ratio feat(ADR-0114a.2): OOD-ratio auditor — Obligation #2 wired for B3, ratio=1.00 (#193) 2026-05-23 16:25:28 -07:00
obligation_5_perturbation feat(ADR-0114a.5): reasoning-isolation perturbation suite — Obligation #5 wired for B3, PASSING 130/130 preserving, 68/68 breaking (#191) 2026-05-23 16:07:59 -07:00
obligation_6_depth_curve feat(ADR-0114a.6): depth-curve auditor — Obligation #6 wired for B3 (assertion holds, coverage gap named) (#190) 2026-05-23 16:19:58 -07:00
obligation_8_adversarial feat(ADR-0114a.8): adversarial auditor — Obligation #8 wired, PASSING; surfaces 2 known parser-layer gaps (#192) 2026-05-23 16:11:37 -07:00
obligation_10_pack_provenance feat(ADR-0114a.10): pack-provenance auditor — Obligation #10 wired for B3, PASSING 2026-05-23 15:44:53 -07:00
orthogonality_tour
prompt_diversity
proof_carrying_promotion_demo feat(demos): Implement ADR-0218 PR D proof-carrying promotion demo (#927) 2026-07-03 13:01:49 -07:00
proofwriter_owa feat(eval): ProofWriter-OWA refusal-floor lane (B1) — independent oracle, measure-only (#779) 2026-06-15 12:28:20 -07:00
propositional_logic feat(comprehend): propositional-logic comprehension (4th domain, flagship oracle) 2026-06-05 23:24:54 -07:00
provenance
public_demo fix(ci): re-pin public_demo lane — real content drift, not the env flake (S2) 2026-07-24 16:58:05 -07:00
rate_oracle test(rate-oracle): port _canonical_outcome non-vacuous validation to R3 (R3-vac) 2026-06-08 07:38:20 -07:00
realizer_guard
refusal_calibration feat(identity): L11 identity continuity — same identity across reboot, not just same bytes 2026-06-05 13:52:57 -07:00
refusal_taxonomy fix: green test-fast suite, consolidate ADR graph under docs/adr, and complete governance cohesion anchors 2026-06-30 17:56:12 -07:00
register_diagnostics
register_tour fix(ci): re-pin public_demo lane SHA after showcase content drift (#807) 2026-06-17 15:42:40 -07:00
relational feat(reader): add overlaps_event finite-verb reader surface (B3) (#783) 2026-06-15 14:56:09 -07:00
relational_inference/v1 feat(reader): add overlaps_event finite-verb reader surface (B3) (#783) 2026-06-15 14:56:09 -07:00
relational_metric feat(oracle): narrow reverse-solve contract — base of one more/fewer_than (PR-7a) 2026-06-07 06:22:36 -07:00
relational_operator_ablation feat(eval): sealed relational operator ablation lane v1 2026-07-19 21:00:15 -07:00
relational_transitive feat(determine): add transitive strict-order relational inference (#781) 2026-06-15 14:20:30 -07:00
reports Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
results
reviewer_registry
sample_efficiency
self_consistency_over_time
sensorimotor_sensorium Add sensorium eval and governance runway 2026-06-03 20:53:05 -07:00
sensorium Add event vision sensorium lane 2026-06-06 12:37:57 -07:00
set_membership evals: add staged independent gold lanes 2026-06-05 16:03:26 -07:00
setup_oracle feat(comprehension): inverse reader frame — base of a more/fewer-than (PR-7b / R2 C0) 2026-06-07 07:06:26 -07:00
structure_mapping feat(trackb): S1–S4 symbolic SME, selector, pure-S1 coverage gain 2026-07-19 19:49:07 -07:00
syllogism feat(comprehend): multi-word NP chunking under a canonicalization contract 2026-06-05 22:47:34 -07:00
symbolic_logic
teaching_injection_resistance
total_ordering feat(comprehend): multi-word NP chunking under a canonicalization contract 2026-06-05 22:47:34 -07:00
vision_sensorium Add vision evidence and sensorimotor contracts 2026-06-03 20:27:46 -07:00
walkthrough_chain
warmed_session_consistency
zero_code_domain_acquisition
__init__.py
_parallel.py
baseline_runner.py
CLAIMS.md Lane 4: Registry Consolidation (language_packs to packs) 2026-07-04 15:11:28 -07:00
cognition_cases.jsonl
framework.py
generalized_lift_instrument.py chore(adr-0250): ratify Accepted + fold in PR #74 review findings 2026-07-18 15:07:29 -07:00
holdout_runner.py
lift_evidence_handoff.py feat(evals): seams S4+S5 — generalized-lift instrument + ports/handoff evidence 2026-07-18 10:34:26 -07:00
logic_cnf_compiler.py feat(adr-0249): P3 structural formula→CNF converter (deduction leg) 2026-07-18 12:40:31 -07:00
metrics.py fix(evals): carry hash_surface on the CaseResult the register matrix actually uses 2026-07-25 14:25:31 +00:00
multi_register_program.py feat(reader-arc): compare compiler-side guards for inverse + chained shapes 2026-07-18 17:14:16 -07:00
parallel.py
run_cognition_eval.py fix(evals): carry hash_surface on the CaseResult the register matrix actually uses 2026-07-25 14:25:31 +00:00
turn_program.py feat(adr-0249): P4 turn-program compiler + chained-relaxation executor 2026-07-18 12:55:31 -07:00
vocab_trigger_instrument.py feat(generalization): vocab-trigger instrument — mechanism-vs-coverage refusal histogram (S3) 2026-07-24 16:54:13 -07:00