The measurement foundation for docs/plans/grammar-unification-2026-07-26.md.
WHY: evals/deterministic_fluency reports 1.00 on all six predicates and
still passes "banana does the.", "wet ground rains the is." and
"is is is is." — it checks terminal punctuation, presence of a verb-shaped
token, and two anti-shape regexes. Heuristic predicates will always have
that failure mode, because grammaticality cannot be measured without a
grammar. So this lane measures agreement between the two halves of CORE
that already encode grammar, and requires the measurement to FAIL on salad.
Two directions, reported separately because they fail for different
reasons and have different remedies:
G-round-trip graph -> realize_target -> surface -> comprehend -> graph
S-round-trip surface -> comprehend -> graph -> categorical renderer -> surface
v1 baseline on main @ 9696443a:
graph_cases 280 surface_cases 8
g_write_rate 1.000 s_read_rate 1.000
g_read_rate 0.000 s_renderable_rate 0.625
g_exact_rate 0.000 s_surface_match_rate 0.000
negative_cases 16
reject_rate 1.000
g_read_rate and s_surface_match_rate are pins on measured DEFECTS, not
goals; they may be revised upward only. s_surface_match_rate = 0 is the
§1.7 categorical render defect caught by construction — the lane found it
without being told to look.
g_args_rate and g_predicates_rate are deliberately separate: high argument
agreement with low predicate agreement would mean the grammars align and
only the vocabulary is split, a materially different remedy from both
being low. That distinction decides the arc's direction (plan §6).
Design notes:
- The committed cases.jsonl is the SINGLE source for authored surfaces —
no in-module duplicate, since a second copy of a corpus is the defect
this arc exists to remove. Negative shuffles are DERIVED at run time so
they cannot drift from the positives.
- The shuffles are lexically identical to positives (same vocabulary, same
length, order destroyed) so the lane cannot pass by vocabulary-checking.
- Fixed rotation, not a PRNG, so reject_rate is byte-reproducible.
- The lane keeps a local copy of the reader's quantifier map ON PURPOSE so
it never becomes a consumer of what it measures;
test_quantifier_map_matches_reader fails loudly if the reader changes.
- _render_categorical deliberately reaches a private serving function: a
lane measuring a private copy would measure what users never see.
Every guarantee is paired with a mutation test. The load-bearing one is
test_reject_rate_goes_red_when_the_reader_accepts_everything: an
accept-everything reader must drive reject_rate to 0.0. Without it,
reject_rate == 1.0 would be unfalsifiable — precisely the defect that
makes the existing fluency lane decoration.
Also documents plainly what round-trip does NOT prove: it measures mutual
intelligibility, not English quality. english_fluency_ood accepts "river
flows valley" and round-trip would be happy with it. No metric here may be
cited as evidence of prose quality.
scripts/measure_grammar_seam.py reproduces every number in the plan's §1
so a reader can check them instead of trusting them.
[Verification]: in-worktree on CPython 3.12.13, uv sync --locked —
smoke 621 (unchanged), deductive 364 (349 + 15 new). Lane SHA pins
verified separately. No serving code touched; new files plus one suite
registration line.
276 lines
10 KiB
Python
276 lines
10 KiB
Python
"""Tests for the grammar round-trip lane.
|
|
|
|
The load-bearing tests here are the *mutation* tests. This lane exists because
|
|
``evals/deterministic_fluency`` reports 1.00 on all six predicates while
|
|
accepting ``"banana does the."`` — a pin that cannot fail is worthless. So
|
|
every guarantee this lane makes is paired with a test proving the guarantee can
|
|
be broken.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import pytest
|
|
|
|
from evals.grammar_roundtrip import runner as rt_runner
|
|
from evals.grammar_roundtrip.corpora import (
|
|
negative_surface_cases,
|
|
positive_graph_cases,
|
|
positive_surface_cases,
|
|
)
|
|
from evals.grammar_roundtrip.projection import (
|
|
CanonicalProposition,
|
|
compare,
|
|
from_meaning_graph,
|
|
from_proposition_graph,
|
|
quantifier_predicate,
|
|
)
|
|
from evals.grammar_roundtrip.runner import run_lane
|
|
from generate.graph_planner import (
|
|
ArticulationStep,
|
|
ArticulationTarget,
|
|
GraphNode,
|
|
PropositionGraph,
|
|
RhetoricalMove,
|
|
)
|
|
from generate.intent import IntentTag
|
|
from generate.meaning_graph.model import (
|
|
Entity,
|
|
MeaningGraph,
|
|
MeaningSpan,
|
|
Relation,
|
|
)
|
|
from generate.meaning_graph.reader import Comprehension, Refusal, comprehend
|
|
|
|
|
|
# The three strings that pass every content predicate of
|
|
# evals/deterministic_fluency. This lane's rejection of them is the concrete,
|
|
# recorded improvement over that lane.
|
|
_FLUENCY_LANE_FALSE_POSITIVES = (
|
|
"banana does the.",
|
|
"wet ground rains the is.",
|
|
"is is is is.",
|
|
)
|
|
|
|
|
|
@pytest.fixture(scope="module")
|
|
def report():
|
|
return run_lane()
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# The negative corpus: the guarantee, and proof it can fail
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_negative_corpus_is_fully_rejected(report):
|
|
"""Word salad must never be comprehended into propositions."""
|
|
assert report.metrics["reject_rate"] == 1.0
|
|
assert report.metrics["negative_cases"] >= 16
|
|
|
|
|
|
def test_the_fluency_lanes_false_positives_are_rejected():
|
|
"""The exact strings deterministic_fluency passes must be refused here."""
|
|
for surface in _FLUENCY_LANE_FALSE_POSITIVES:
|
|
result = comprehend(surface, source_id="t")
|
|
if isinstance(result, Refusal):
|
|
continue
|
|
assert not from_meaning_graph(result.meaning_graph), (
|
|
f"{surface!r} was comprehended into propositions; the negative "
|
|
"corpus guarantee is broken"
|
|
)
|
|
|
|
|
|
def test_reject_rate_goes_red_when_the_reader_accepts_everything(monkeypatch):
|
|
"""MUTATION: an accept-everything reader must drive reject_rate to 0.
|
|
|
|
Without this test ``reject_rate == 1.0`` would be unfalsifiable — exactly
|
|
the defect that makes the existing fluency lane decoration.
|
|
"""
|
|
span = MeaningSpan(source_id="t", start=0, end=1, text="x")
|
|
graph = MeaningGraph(
|
|
entities=(
|
|
Entity(entity_id="a", name="a", span=span),
|
|
Entity(entity_id="b", name="b", span=span),
|
|
),
|
|
relations=(Relation(predicate="subset", arguments=("a", "b"), span=span),),
|
|
)
|
|
monkeypatch.setattr(
|
|
rt_runner, "comprehend", lambda text, source_id="input": Comprehension(meaning_graph=graph)
|
|
)
|
|
mutated = run_lane()
|
|
assert mutated.metrics["reject_rate"] == 0.0, (
|
|
"an accept-everything reader still scored a perfect reject_rate — the "
|
|
"negative corpus is not actually gating"
|
|
)
|
|
|
|
|
|
def test_shuffled_negatives_reuse_positive_vocabulary():
|
|
"""Shuffles must be lexically identical to a positive, order destroyed.
|
|
|
|
A lane that rejects only hand-authored salad could be checking vocabulary
|
|
rather than grammar. The shuffles remove that escape.
|
|
"""
|
|
positives = {
|
|
frozenset(c.surface.rstrip(".").lower().split()) for c in positive_surface_cases()
|
|
}
|
|
shuffles = [c for c in negative_surface_cases() if c.case_id.startswith("neg-shuf-")]
|
|
assert shuffles, "no shuffled negatives were generated"
|
|
for case in shuffles:
|
|
tokens = frozenset(case.surface.rstrip(".").lower().split())
|
|
assert tokens in positives, (
|
|
f"{case.case_id} does not reuse a positive case's exact vocabulary"
|
|
)
|
|
|
|
|
|
def test_shuffles_are_byte_stable_across_calls():
|
|
"""The negative corpus must be reproducible — no PRNG, no clock."""
|
|
first = tuple(c.surface for c in negative_surface_cases())
|
|
second = tuple(c.surface for c in negative_surface_cases())
|
|
assert first == second
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Baselines — these pin DEFECTS and must be revised upward when fixed
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_g_roundtrip_baseline_is_zero(report):
|
|
"""BASELINE PIN (a defect, not a goal): CORE reads 0% of what it writes.
|
|
|
|
Measured on main @ 9696443a. When the grammar is unified this must be
|
|
revised **upward**; it must never be revised downward to accommodate a
|
|
regression.
|
|
"""
|
|
assert report.metrics["graph_cases"] == 280
|
|
assert report.metrics["g_write_rate"] == 1.0
|
|
assert report.metrics["g_read_rate"] == 0.0
|
|
assert report.metrics["g_exact_rate"] == 0.0
|
|
|
|
|
|
def test_s_roundtrip_pins_the_categorical_render_defect(report):
|
|
"""BASELINE PIN (a defect): nothing CORE reads renders back to its input.
|
|
|
|
The reader comprehends all 8 positive surfaces, and the serving categorical
|
|
renderer reproduces none of them, because it interpolates singularized
|
|
entity ids into plural templates. Phase 2B fixes this and must raise this
|
|
number.
|
|
"""
|
|
assert report.metrics["s_read_rate"] == 1.0
|
|
assert report.metrics["s_surface_match_rate"] == 0.0
|
|
|
|
|
|
def test_the_categorical_render_defect_concretely():
|
|
"""The specific defect, stated as an example rather than a rate."""
|
|
result = comprehend("All dogs are animals.", source_id="t")
|
|
assert isinstance(result, Comprehension)
|
|
props = sorted(from_meaning_graph(result.meaning_graph))
|
|
assert props == [CanonicalProposition("subset", "dog", "animal", False)]
|
|
rendered = rt_runner._render_categorical("subset", "dog", "animal")
|
|
# "all dog are animal" — plural template, singular entity ids.
|
|
assert rendered == "all dog are animal"
|
|
assert rendered != "all dogs are animals"
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# The projection itself
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_quantifier_map_matches_reader():
|
|
"""The lane's local quantifier map must track the reader's.
|
|
|
|
The copy is deliberate — the lane must not import the thing it measures —
|
|
so this test is what keeps the yardstick honest if the reader changes.
|
|
"""
|
|
from generate.meaning_graph.reader import _QUANTIFIER_PREDICATE as reader_map
|
|
|
|
for quantifier, predicate in reader_map.items():
|
|
assert quantifier_predicate(quantifier) == predicate
|
|
assert quantifier_predicate(None) is None
|
|
assert quantifier_predicate("nonsense") is None
|
|
|
|
|
|
def test_proposition_graph_projection_folds_quantifier():
|
|
"""A writer graph's ``quantifier`` slot folds into the reading predicate."""
|
|
node = GraphNode(
|
|
node_id="n1", subject="dog", predicate="is_a", obj="animal",
|
|
source_intent=IntentTag.UNKNOWN,
|
|
)
|
|
step = ArticulationStep(
|
|
node_id="n1", move=RhetoricalMove.ASSERT, predicate="is_a",
|
|
subject="dog", quantifier="all",
|
|
)
|
|
projected = from_proposition_graph(
|
|
PropositionGraph(nodes=(node,)),
|
|
ArticulationTarget(steps=(step,), source_intent=IntentTag.UNKNOWN),
|
|
)
|
|
assert projected == frozenset({CanonicalProposition("subset", "dog", "animal", False)})
|
|
|
|
|
|
def test_projection_drops_non_binary_relations():
|
|
"""Arity != 2 is dropped — the writing side cannot express it."""
|
|
span = MeaningSpan(source_id="t", start=0, end=1, text="x")
|
|
graph = MeaningGraph(
|
|
entities=(
|
|
Entity(entity_id="a", name="a", span=span),
|
|
Entity(entity_id="b", name="b", span=span),
|
|
Entity(entity_id="c", name="c", span=span),
|
|
),
|
|
relations=(
|
|
Relation(predicate="between", arguments=("a", "b", "c"), span=span),
|
|
Relation(predicate="subset", arguments=("a", "b"), span=span),
|
|
),
|
|
)
|
|
assert from_meaning_graph(graph) == frozenset(
|
|
{CanonicalProposition("subset", "a", "b", False)}
|
|
)
|
|
|
|
|
|
def test_compare_decomposes_argument_and_predicate_agreement():
|
|
"""A predicate name only counts when it agrees on the SAME arguments."""
|
|
expected = frozenset({CanonicalProposition("subset", "dog", "animal")})
|
|
# right arguments, wrong predicate name
|
|
args_only = frozenset({CanonicalProposition("member", "dog", "animal")})
|
|
agreement = compare(expected, args_only)
|
|
assert agreement.args_match == 1
|
|
assert agreement.predicates_match == 0
|
|
assert agreement.exact_match == 0
|
|
|
|
# right predicate name, wrong arguments — must NOT count
|
|
pred_elsewhere = frozenset({CanonicalProposition("subset", "cat", "plant")})
|
|
agreement = compare(expected, pred_elsewhere)
|
|
assert agreement.args_match == 0
|
|
assert agreement.predicates_match == 0
|
|
|
|
|
|
def test_compare_exact_agreement():
|
|
prop = CanonicalProposition("subset", "dog", "animal")
|
|
agreement = compare(frozenset({prop}), frozenset({prop}))
|
|
assert agreement.exact_rate == 1.0
|
|
assert agreement.args_rate == 1.0
|
|
assert agreement.predicates_rate == 1.0
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# Corpus integrity
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_positive_surfaces_are_all_inside_the_reader_envelope():
|
|
"""Every positive surface must actually be comprehended.
|
|
|
|
A refused positive would silently measure the reader's coverage instead of
|
|
the round-trip, and would make s_surface_match_rate uninterpretable.
|
|
"""
|
|
for case in positive_surface_cases():
|
|
result = comprehend(case.surface, source_id="t")
|
|
assert isinstance(result, Comprehension), (
|
|
f"{case.case_id} ({case.surface!r}) is refused; it does not belong "
|
|
"in the positive corpus"
|
|
)
|
|
|
|
|
|
def test_graph_corpus_is_harvested_from_committed_case_files():
|
|
cases = positive_graph_cases()
|
|
assert len(cases) == 280
|
|
assert all(c.nodes for c in cases)
|