Merge pull request 'feat(evals): grammar round-trip instrument — CORE cannot read its own writing (0/280)' (#130) from feat/grammar-roundtrip-instrument into main
This commit is contained in:
commit
0948f7cdb8
9 changed files with 1436 additions and 0 deletions
|
|
@ -244,6 +244,7 @@ TEST_SUITES: dict[str, tuple[str, ...]] = {
|
|||
"tests/test_ledger_reseal.py",
|
||||
"tests/test_ratified_ledger_bridge.py",
|
||||
"tests/test_vocab_trigger_instrument.py",
|
||||
"tests/test_grammar_roundtrip.py",
|
||||
),
|
||||
"full": ("tests/",),
|
||||
}
|
||||
|
|
|
|||
1
evals/grammar_roundtrip/__init__.py
Normal file
1
evals/grammar_roundtrip/__init__.py
Normal file
|
|
@ -0,0 +1 @@
|
|||
"""Grammar round-trip eval lane (see runner.py)."""
|
||||
143
evals/grammar_roundtrip/contract.md
Normal file
143
evals/grammar_roundtrip/contract.md
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
# Grammar Round-Trip Eval Lane — Contract
|
||||
|
||||
**Lane:** `grammar_roundtrip`
|
||||
**Version:** v1
|
||||
**Created:** 2026-07-26
|
||||
**Plan:** `docs/plans/grammar-unification-2026-07-26.md` (Phase 1)
|
||||
|
||||
## What this lane measures
|
||||
|
||||
Whether CORE can **read what it writes** and **write what it reads** — and,
|
||||
critically, whether it *refuses* word salad.
|
||||
|
||||
## Why it exists
|
||||
|
||||
`evals/deterministic_fluency` reports **1.00 on all six of its predicates**
|
||||
and still passes every one of these:
|
||||
|
||||
| candidate | passes deterministic_fluency? |
|
||||
|---|---|
|
||||
| `"banana does the."` | yes |
|
||||
| `"wet ground rains the is."` | yes |
|
||||
| `"is is is is."` | yes |
|
||||
|
||||
It checks terminal punctuation, presence of a verb-shaped token, and two
|
||||
anti-shape regexes. It cannot distinguish English from word salad, so its 100%
|
||||
carries no information about fluency.
|
||||
|
||||
Heuristic predicates will always have that failure mode, because
|
||||
**grammaticality cannot be measured without a grammar.** This lane therefore
|
||||
measures agreement between the two halves of CORE that already encode grammar —
|
||||
the reader and the writer — and requires the measurement to *fail* on salad.
|
||||
Round-trip is falsifiable with no judge, no embedding, and no gold aesthetic.
|
||||
|
||||
## The two directions
|
||||
|
||||
| direction | pipeline | what a failure means |
|
||||
|---|---|---|
|
||||
| **G-round-trip** | graph → `realize_target` → surface → `comprehend` → graph | CORE cannot read its own writing |
|
||||
| **S-round-trip** | surface → `comprehend` → graph → categorical renderer → surface | CORE cannot reproduce what it just understood |
|
||||
|
||||
Both are reported because they fail for different reasons and have different
|
||||
remedies.
|
||||
|
||||
## Metrics
|
||||
|
||||
| metric | definition |
|
||||
|---|---|
|
||||
| `g_write_rate` | fraction of graph cases the writer produced any surface for |
|
||||
| `g_read_rate` | fraction of written surfaces the reader comprehended |
|
||||
| `g_args_rate` | fraction of expected propositions whose **argument pair** was recovered |
|
||||
| `g_predicates_rate` | fraction whose **predicate name** was recovered *on the same arguments* |
|
||||
| `g_exact_rate` | fraction recovered exactly (predicate + arguments + polarity) |
|
||||
| `s_read_rate` | fraction of positive surfaces comprehended |
|
||||
| `s_renderable_rate` | fraction whose projection the categorical renderer can express at all |
|
||||
| `s_surface_match_rate` | fraction that render back to the input surface |
|
||||
| `reject_rate` | fraction of the **negative** corpus refused or reduced to zero propositions |
|
||||
|
||||
`g_args_rate` and `g_predicates_rate` are reported separately on purpose. High
|
||||
argument agreement with low predicate agreement means the grammars align and
|
||||
only the *vocabulary* is split — a materially different remedy from both being
|
||||
low. Collapsing them into one boolean would hide the distinction that decides
|
||||
this arc's direction (plan §6).
|
||||
|
||||
## Corpora
|
||||
|
||||
| corpus | source | size |
|
||||
|---|---|---|
|
||||
| positive graphs | committed `english_fluency_ood` + `grammatical_coverage` case files | 280 |
|
||||
| positive surfaces | in-module, each verified comprehensible by probe | 8 |
|
||||
| negative surfaces | hand-authored salad + deterministic token shuffles of the positives | 16 |
|
||||
|
||||
The shuffles are load-bearing: they are **lexically identical** to positive
|
||||
cases — same vocabulary, same length, order destroyed. A lane that rejects
|
||||
hand-authored salad but accepts the shuffles is doing vocabulary checking, not
|
||||
grammar checking. Shuffling uses a fixed rotation rather than a PRNG so
|
||||
`reject_rate` is byte-reproducible.
|
||||
|
||||
## Thresholds
|
||||
|
||||
| metric | v1 requirement | rationale |
|
||||
|---|---|---|
|
||||
| `reject_rate` | **1.00, always** | the guarantee; regression is a hard failure |
|
||||
| `g_read_rate` | recorded baseline, revised **upward only** | currently a measured defect |
|
||||
| `s_surface_match_rate` | recorded baseline, revised **upward only** | currently a measured defect |
|
||||
|
||||
## v1 baseline — measured on `main` @ `9696443a`
|
||||
|
||||
```
|
||||
graph_cases 280
|
||||
g_write_rate 1.000
|
||||
g_read_rate 0.000 <-- CORE reads 0% of what it writes
|
||||
g_propositions_expected 370
|
||||
g_args_rate 0.000
|
||||
g_predicates_rate 0.000
|
||||
g_exact_rate 0.000
|
||||
|
||||
surface_cases 8
|
||||
s_read_rate 1.000
|
||||
s_renderable_rate 0.625
|
||||
s_surface_match_rate 0.000 <-- nothing renders back to its input
|
||||
|
||||
negative_cases 16
|
||||
reject_rate 1.000 <-- the guarantee holds
|
||||
```
|
||||
|
||||
Two of these are **pins on known defects**, not goals. They are expected to be
|
||||
revised upward by plan Phases 2B and 5, and must never be revised downward to
|
||||
accommodate a regression.
|
||||
|
||||
`s_renderable_rate = 0.625` is not a defect: 3 of the 8 positive surfaces
|
||||
project to `member` / `less` predicates, which have no `all X are Y` categorical
|
||||
surface at all. Those are reported as unrenderable rather than as match
|
||||
failures, because "cannot write this" and "wrote this wrongly" need different
|
||||
fixes.
|
||||
|
||||
## What this lane does NOT prove
|
||||
|
||||
Round-trip is **necessary, not sufficient** for fluency. It measures *mutual
|
||||
intelligibility* — both halves agreeing about a surface. Two halves can agree
|
||||
on an impoverished construction: `english_fluency_ood` accepts `"river flows
|
||||
valley"` as a correct surface, and round-trip would be perfectly happy with it.
|
||||
|
||||
So a high round-trip rate licenses **"CORE means what it says"**, never
|
||||
**"CORE writes well"**. The negative corpus is what prevents the first failure
|
||||
mode. Nothing in this lane prevents the second, and no metric here should ever
|
||||
be cited as evidence of prose quality.
|
||||
|
||||
## Mutation guarantees
|
||||
|
||||
Every guarantee is paired with a test proving it can break
|
||||
(`tests/test_grammar_roundtrip.py`):
|
||||
|
||||
- `test_reject_rate_goes_red_when_the_reader_accepts_everything` — an
|
||||
accept-everything reader must drive `reject_rate` to **0.0**. Without this,
|
||||
`reject_rate == 1.0` would be unfalsifiable, which is precisely the defect
|
||||
that makes the existing fluency lane decoration.
|
||||
- `test_shuffled_negatives_reuse_positive_vocabulary` — closes the
|
||||
vocabulary-checking escape hatch.
|
||||
- `test_positive_surfaces_are_all_inside_the_reader_envelope` — a refused
|
||||
positive would silently measure reader coverage instead of round-trip.
|
||||
- `test_quantifier_map_matches_reader` — the lane keeps a deliberate local copy
|
||||
of the reader's quantifier map so it never becomes a consumer of the thing it
|
||||
measures; this test fails loudly if the reader's map changes.
|
||||
168
evals/grammar_roundtrip/corpora.py
Normal file
168
evals/grammar_roundtrip/corpora.py
Normal file
|
|
@ -0,0 +1,168 @@
|
|||
"""Corpora for the grammar round-trip lane.
|
||||
|
||||
Three corpora, and all three are load-bearing:
|
||||
|
||||
* :func:`positive_graph_cases` — writing-side graphs, for the **G-round-trip**
|
||||
``read(write(g))``. Sourced from the committed ``english_fluency_ood`` and
|
||||
``grammatical_coverage`` case files so the lane measures the same graphs the
|
||||
existing fluency lanes score.
|
||||
* :func:`positive_surface_cases` — reading-side surfaces inside the reader's
|
||||
demonstrated envelope, for the **S-round-trip** ``write(read(s))``. This is
|
||||
the direction that produces a non-zero signal today, which is what makes the
|
||||
lane diagnostic rather than a flat zero.
|
||||
* :func:`negative_surface_cases` — word salad that MUST be rejected. Without
|
||||
this corpus the lane is decoration: ``evals/deterministic_fluency`` reports
|
||||
1.00 on all six predicates and still passes ``"banana does the."``. A lane
|
||||
that scores only positives cannot tell "understands English" from "accepts
|
||||
anything".
|
||||
|
||||
The negative corpus is deliberately built two ways — hand-authored salad plus
|
||||
mechanical token shuffles of the positive surfaces — so it cannot be satisfied
|
||||
by pattern-matching a fixed list of bad strings.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
_REPO_ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
_GRAPH_CASE_GLOBS = (
|
||||
"evals/english_fluency_ood/public/v1/cases.jsonl",
|
||||
"evals/english_fluency_ood/holdouts/v1/cases.jsonl",
|
||||
"evals/grammatical_coverage/public/v1/cases.jsonl",
|
||||
"evals/grammatical_coverage/public/v2/cases.jsonl",
|
||||
"evals/grammatical_coverage/holdouts/v1/cases.jsonl",
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class GraphCase:
|
||||
"""One writing-side case: an id plus the raw proposition-graph dict."""
|
||||
|
||||
case_id: str
|
||||
nodes: tuple[dict[str, object], ...]
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class SurfaceCase:
|
||||
"""One reading-side case: a surface plus why it is in the corpus."""
|
||||
|
||||
case_id: str
|
||||
surface: str
|
||||
note: str = ""
|
||||
|
||||
|
||||
def positive_graph_cases() -> tuple[GraphCase, ...]:
|
||||
"""Writing-side graphs harvested from the committed fluency case files."""
|
||||
out: list[GraphCase] = []
|
||||
for rel in _GRAPH_CASE_GLOBS:
|
||||
path = _REPO_ROOT / rel
|
||||
if not path.exists():
|
||||
continue
|
||||
for line in path.read_text().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
case = json.loads(line)
|
||||
nodes = case.get("proposition_graph", {}).get("nodes", [])
|
||||
if not nodes:
|
||||
continue
|
||||
out.append(
|
||||
GraphCase(case_id=str(case.get("id", "?")), nodes=tuple(nodes))
|
||||
)
|
||||
return tuple(out)
|
||||
|
||||
|
||||
#: The lane's own committed corpus. Authored surfaces live HERE and only here
|
||||
#: — not duplicated in module tables — because a second copy of a corpus is the
|
||||
#: same defect this arc exists to remove. The derived shuffles below are
|
||||
#: generated at run time rather than committed, so they cannot drift away from
|
||||
#: the positives they are derived from.
|
||||
_SURFACE_CASES_PATH = Path(__file__).resolve().parent / "public" / "v1" / "cases.jsonl"
|
||||
|
||||
_KIND_POSITIVE = "positive_surface"
|
||||
_KIND_NEGATIVE = "negative_surface"
|
||||
|
||||
|
||||
def _load_surface_cases(kind: str) -> tuple[SurfaceCase, ...]:
|
||||
if not _SURFACE_CASES_PATH.exists():
|
||||
raise FileNotFoundError(
|
||||
f"grammar_roundtrip corpus missing: {_SURFACE_CASES_PATH}"
|
||||
)
|
||||
out: list[SurfaceCase] = []
|
||||
for line in _SURFACE_CASES_PATH.read_text().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
record = json.loads(line)
|
||||
if record.get("kind") != kind:
|
||||
continue
|
||||
out.append(
|
||||
SurfaceCase(
|
||||
case_id=str(record["id"]),
|
||||
surface=str(record["surface"]),
|
||||
note=str(record.get("note", "")),
|
||||
)
|
||||
)
|
||||
return tuple(out)
|
||||
|
||||
|
||||
def positive_surface_cases() -> tuple[SurfaceCase, ...]:
|
||||
"""Reading-side surfaces inside the reader's demonstrated envelope.
|
||||
|
||||
Each was verified comprehensible by probe before being committed — a case
|
||||
the reader refuses belongs in a coverage note, not here, because a refusal
|
||||
would measure the reader's envelope rather than the round-trip.
|
||||
``tests/test_grammar_roundtrip.py`` enforces that.
|
||||
"""
|
||||
return _load_surface_cases(_KIND_POSITIVE)
|
||||
|
||||
|
||||
def _shuffle_tokens(surface: str, seed: int) -> str:
|
||||
"""Deterministically permute *surface*'s tokens into a non-identity order.
|
||||
|
||||
Uses a fixed rotation rather than a PRNG so the corpus is byte-stable
|
||||
across runs and machines — a shuffled negative case that varies per run
|
||||
would make the lane's ``reject_rate`` irreproducible.
|
||||
"""
|
||||
body = surface.rstrip(".!?")
|
||||
tokens = body.split()
|
||||
if len(tokens) < 3:
|
||||
return surface
|
||||
shift = 1 + (seed % (len(tokens) - 1))
|
||||
rotated = tokens[shift:] + tokens[:shift]
|
||||
return " ".join(rotated) + "."
|
||||
|
||||
|
||||
def negative_surface_cases() -> tuple[SurfaceCase, ...]:
|
||||
"""Word salad: hand-authored plus mechanical shuffles of the positives.
|
||||
|
||||
The shuffles matter because they are *lexically identical* to positive
|
||||
cases — same vocabulary, same length, only the order destroyed. A lane
|
||||
that rejects the hand-authored salad but accepts the shuffles is doing
|
||||
vocabulary checking, not grammar checking.
|
||||
"""
|
||||
positives = positive_surface_cases()
|
||||
out = list(_load_surface_cases(_KIND_NEGATIVE))
|
||||
for i, case in enumerate(positives):
|
||||
shuffled = _shuffle_tokens(case.surface, i)
|
||||
if shuffled.rstrip(".!?").lower() == case.surface.rstrip(".!?").lower():
|
||||
continue
|
||||
out.append(
|
||||
SurfaceCase(
|
||||
case_id=f"neg-shuf-{case.case_id}",
|
||||
surface=shuffled,
|
||||
note=f"token shuffle of {case.case_id}",
|
||||
)
|
||||
)
|
||||
return tuple(out)
|
||||
|
||||
|
||||
__all__ = (
|
||||
"GraphCase",
|
||||
"SurfaceCase",
|
||||
"negative_surface_cases",
|
||||
"positive_graph_cases",
|
||||
"positive_surface_cases",
|
||||
)
|
||||
219
evals/grammar_roundtrip/projection.py
Normal file
219
evals/grammar_roundtrip/projection.py
Normal file
|
|
@ -0,0 +1,219 @@
|
|||
"""The MeaningGraph <-> PropositionGraph projection this lane measures against.
|
||||
|
||||
CORE has two graph models and they are not type-compatible:
|
||||
|
||||
* ``MeaningGraph`` (reading) — ``Entity(entity_id, name, span, kind)`` plus
|
||||
``Relation(predicate, arguments: tuple[str, ...], span, negated)``: n-ary
|
||||
predicate/argument structure carrying source spans.
|
||||
* ``PropositionGraph`` (writing) — ``GraphNode(node_id, subject, predicate,
|
||||
obj, ...)`` plus ``GraphEdge(source, target, relation)``: fixed
|
||||
subject-predicate-object triples with rhetorical discourse edges.
|
||||
|
||||
A literal ``read(write(g)) == g`` is therefore not expressible. This module
|
||||
defines the projection both sides are compared *through*, and states plainly
|
||||
what the projection discards. Everything discarded here is a thing the lane
|
||||
CANNOT see — that list is the honest limit of the measurement.
|
||||
|
||||
Deliberately discarded
|
||||
----------------------
|
||||
* **Discourse structure** — rhetorical moves, ``GraphEdge`` relations, and
|
||||
sentence order. Two graphs whose propositions match but whose connectives
|
||||
differ project identically.
|
||||
* **Spans** — ``MeaningSpan`` has no counterpart on the writing side.
|
||||
* **Entity ``kind``** — the reading side infers ``individual``/``class``;
|
||||
the writing side has no slot for it.
|
||||
* **Tense and aspect** — carried on ``ArticulationStep``, absent from
|
||||
``Relation``.
|
||||
* **Quantifier identity** — the reading side folds quantification INTO the
|
||||
predicate (``all`` -> ``subset``, ``no`` -> ``disjoint``, ``some`` ->
|
||||
``intersects``); the writing side carries ``quantifier`` as a separate slot
|
||||
on the step. The projection keeps the reading side's convention, so a
|
||||
writer graph's ``quantifier`` is visible only through
|
||||
:func:`quantifier_predicate`.
|
||||
* **Relations of arity != 2** — the writing side cannot express them at all.
|
||||
|
||||
What it keeps
|
||||
-------------
|
||||
A frozenset of :class:`CanonicalProposition` — ``(predicate, subject, obj,
|
||||
negated)``, lowercased and stripped. Because the two sides do not share a
|
||||
predicate vocabulary (the writer says ``flows``/``is_defined_as``; the reader
|
||||
says ``subset``/``member``/``less``), the lane reports argument agreement and
|
||||
predicate agreement **separately**. A single match/no-match boolean would
|
||||
hide exactly the distinction that decides this arc's next direction.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
|
||||
#: Reading-side folding of a quantifier into a predicate name. Mirrors
|
||||
#: ``generate.meaning_graph.reader._QUANTIFIER_PREDICATE``; kept as a local
|
||||
#: copy ON PURPOSE — this lane must not become a consumer of the thing it
|
||||
#: measures, or a change to the reader would silently move the yardstick.
|
||||
#: The duplication is asserted-equal by
|
||||
#: ``tests/test_grammar_roundtrip.py::test_quantifier_map_matches_reader``,
|
||||
#: which fails loudly if the reader's map changes.
|
||||
_QUANTIFIER_PREDICATE: dict[str, str] = {
|
||||
"all": "subset",
|
||||
"no": "disjoint",
|
||||
"some": "intersects",
|
||||
}
|
||||
|
||||
|
||||
def quantifier_predicate(quantifier: str | None) -> str | None:
|
||||
"""Reading-side predicate name for a writing-side *quantifier* slot."""
|
||||
if quantifier is None:
|
||||
return None
|
||||
return _QUANTIFIER_PREDICATE.get(quantifier.strip().lower())
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True, order=True)
|
||||
class CanonicalProposition:
|
||||
"""One binary proposition, normalized so both sides are comparable."""
|
||||
|
||||
predicate: str
|
||||
subject: str
|
||||
obj: str
|
||||
negated: bool = False
|
||||
|
||||
@property
|
||||
def args(self) -> tuple[str, str]:
|
||||
"""The argument pair alone — predicate-vocabulary independent."""
|
||||
return (self.subject, self.obj)
|
||||
|
||||
|
||||
def _norm(value: object) -> str:
|
||||
return str(value).strip().lower()
|
||||
|
||||
|
||||
def from_meaning_graph(meaning_graph: Any) -> frozenset[CanonicalProposition]:
|
||||
"""Project a ``MeaningGraph`` into canonical propositions.
|
||||
|
||||
Relations of arity != 2 are dropped — the writing side cannot express
|
||||
them, so keeping them would make round-trip unachievable for a reason
|
||||
that has nothing to do with grammar.
|
||||
"""
|
||||
out: set[CanonicalProposition] = set()
|
||||
for relation in getattr(meaning_graph, "relations", ()) or ():
|
||||
arguments = tuple(getattr(relation, "arguments", ()) or ())
|
||||
if len(arguments) != 2:
|
||||
continue
|
||||
out.add(
|
||||
CanonicalProposition(
|
||||
predicate=_norm(relation.predicate),
|
||||
subject=_norm(arguments[0]),
|
||||
obj=_norm(arguments[1]),
|
||||
negated=bool(getattr(relation, "negated", False)),
|
||||
)
|
||||
)
|
||||
return frozenset(out)
|
||||
|
||||
|
||||
def from_proposition_graph(
|
||||
graph: Any, target: Any = None
|
||||
) -> frozenset[CanonicalProposition]:
|
||||
"""Project a ``PropositionGraph`` (+ optional target) into canonicals.
|
||||
|
||||
``target`` supplies the per-step ``negated`` / ``quantifier`` slots, which
|
||||
live on ``ArticulationStep`` rather than on ``GraphNode``. When a step
|
||||
carries a quantifier that the reading side folds into a predicate, the
|
||||
folded name is used so the two sides are comparable at all.
|
||||
"""
|
||||
steps_by_id: dict[str, Any] = {}
|
||||
for step in getattr(target, "steps", ()) or ():
|
||||
steps_by_id[step.node_id] = step
|
||||
|
||||
out: set[CanonicalProposition] = set()
|
||||
for node in getattr(graph, "nodes", ()) or ():
|
||||
step = steps_by_id.get(node.node_id)
|
||||
negated = bool(getattr(step, "negated", False)) if step else False
|
||||
quantifier = getattr(step, "quantifier", None) if step else None
|
||||
folded = quantifier_predicate(quantifier)
|
||||
out.add(
|
||||
CanonicalProposition(
|
||||
predicate=folded or _norm(node.predicate),
|
||||
subject=_norm(node.subject),
|
||||
obj=_norm(node.obj),
|
||||
negated=negated,
|
||||
)
|
||||
)
|
||||
return frozenset(out)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class Agreement:
|
||||
"""How well two canonical-proposition sets agree, decomposed.
|
||||
|
||||
The decomposition is the point. ``args_match`` says the *structure*
|
||||
survived the round-trip; ``predicates_match`` says the two sides also
|
||||
agree on what to CALL the relation. A high ``args_match`` with a low
|
||||
``predicates_match`` means the grammars are aligned and only the
|
||||
vocabulary is split — a very different remedy from both being low.
|
||||
"""
|
||||
|
||||
expected: int
|
||||
actual: int
|
||||
args_match: int
|
||||
predicates_match: int
|
||||
exact_match: int
|
||||
|
||||
@property
|
||||
def args_rate(self) -> float:
|
||||
return self.args_match / self.expected if self.expected else 0.0
|
||||
|
||||
@property
|
||||
def predicates_rate(self) -> float:
|
||||
return self.predicates_match / self.expected if self.expected else 0.0
|
||||
|
||||
@property
|
||||
def exact_rate(self) -> float:
|
||||
return self.exact_match / self.expected if self.expected else 0.0
|
||||
|
||||
def as_dict(self) -> dict[str, object]:
|
||||
return {
|
||||
"expected": self.expected,
|
||||
"actual": self.actual,
|
||||
"args_match": self.args_match,
|
||||
"predicates_match": self.predicates_match,
|
||||
"exact_match": self.exact_match,
|
||||
"args_rate": round(self.args_rate, 6),
|
||||
"predicates_rate": round(self.predicates_rate, 6),
|
||||
"exact_rate": round(self.exact_rate, 6),
|
||||
}
|
||||
|
||||
|
||||
def compare(
|
||||
expected: frozenset[CanonicalProposition],
|
||||
actual: frozenset[CanonicalProposition],
|
||||
) -> Agreement:
|
||||
"""Compare two canonical sets, decomposed by argument vs predicate."""
|
||||
actual_args = {p.args for p in actual}
|
||||
args_match = sum(1 for p in expected if p.args in actual_args)
|
||||
# A predicate only counts as agreeing when it agrees ON THE SAME
|
||||
# arguments — a shared predicate name over different arguments is not
|
||||
# recovery, and counting it would inflate the rate.
|
||||
predicates_match = sum(
|
||||
1
|
||||
for p in expected
|
||||
if any(a.predicate == p.predicate and a.args == p.args for a in actual)
|
||||
)
|
||||
exact_match = len(expected & actual)
|
||||
return Agreement(
|
||||
expected=len(expected),
|
||||
actual=len(actual),
|
||||
args_match=args_match,
|
||||
predicates_match=predicates_match,
|
||||
exact_match=exact_match,
|
||||
)
|
||||
|
||||
|
||||
__all__ = (
|
||||
"Agreement",
|
||||
"CanonicalProposition",
|
||||
"compare",
|
||||
"from_meaning_graph",
|
||||
"from_proposition_graph",
|
||||
"quantifier_predicate",
|
||||
)
|
||||
16
evals/grammar_roundtrip/public/v1/cases.jsonl
Normal file
16
evals/grammar_roundtrip/public/v1/cases.jsonl
Normal file
|
|
@ -0,0 +1,16 @@
|
|||
{"id": "pos-cat-A-regular", "kind": "positive_surface", "surface": "All dogs are animals.", "note": "A-form, regular plural"}
|
||||
{"id": "pos-cat-A-plural2", "kind": "positive_surface", "surface": "All wolves are mammals.", "note": "A-form, -ves plural (reader singularizes to 'wolve')"}
|
||||
{"id": "pos-cat-A-irreg", "kind": "positive_surface", "surface": "All men are mortal men.", "note": "A-form, irregular plural"}
|
||||
{"id": "pos-cat-E", "kind": "positive_surface", "surface": "No dogs are cats.", "note": "E-form"}
|
||||
{"id": "pos-cat-I", "kind": "positive_surface", "surface": "Some dogs are pets.", "note": "I-form"}
|
||||
{"id": "pos-mem", "kind": "positive_surface", "surface": "Socrates is a man.", "note": "singular membership; no categorical surface"}
|
||||
{"id": "pos-cmp-taller", "kind": "positive_surface", "surface": "Alice is taller than Bob.", "note": "comparative; no categorical surface"}
|
||||
{"id": "pos-cmp-above", "kind": "positive_surface", "surface": "Alice is above Bob.", "note": "spatial comparative; no categorical surface"}
|
||||
{"id": "neg-salad-banana", "kind": "negative_surface", "surface": "banana does the.", "note": "passes every content predicate of deterministic_fluency"}
|
||||
{"id": "neg-salad-wetground", "kind": "negative_surface", "surface": "wet ground rains the is.", "note": "passes every content predicate of deterministic_fluency"}
|
||||
{"id": "neg-salad-isisis", "kind": "negative_surface", "surface": "is is is is.", "note": "passes every content predicate of deterministic_fluency"}
|
||||
{"id": "neg-salad-colorless", "kind": "negative_surface", "surface": "colorless green ideas sleep furiously.", "note": "syntactically well-formed, semantically empty"}
|
||||
{"id": "neg-salad-funcwords", "kind": "negative_surface", "surface": "the of and to a in that.", "note": "function words only"}
|
||||
{"id": "neg-salad-quantonly", "kind": "negative_surface", "surface": "all some no every.", "note": "quantifiers only"}
|
||||
{"id": "neg-salad-nonverb", "kind": "negative_surface", "surface": "dog cat mammal animal.", "note": "nouns only, no predicate"}
|
||||
{"id": "neg-salad-reversed", "kind": "negative_surface", "surface": "animals are dogs all.", "note": "reversed A-form"}
|
||||
277
evals/grammar_roundtrip/runner.py
Normal file
277
evals/grammar_roundtrip/runner.py
Normal file
|
|
@ -0,0 +1,277 @@
|
|||
"""Grammar round-trip eval lane.
|
||||
|
||||
Measures whether CORE can read what it writes, and write what it reads, in
|
||||
both directions, with a negative corpus that must be rejected.
|
||||
|
||||
Why this lane exists
|
||||
--------------------
|
||||
``evals/deterministic_fluency`` reports 1.00 on all six of its predicates and
|
||||
still passes ``"banana does the."``, ``"wet ground rains the is."`` and
|
||||
``"is is is is."``. It checks terminal punctuation, the presence of a
|
||||
verb-shaped token, and two anti-shape regexes; it cannot distinguish English
|
||||
from word salad. Heuristic predicates will always have that failure mode,
|
||||
because **grammaticality cannot be measured without a grammar**.
|
||||
|
||||
So this lane measures agreement between the two halves of CORE that already
|
||||
encode grammar — the reader and the writer — and requires the measurement to
|
||||
fail on salad. Round-trip is falsifiable without a judge, an embedding, or a
|
||||
gold aesthetic.
|
||||
|
||||
What round-trip does NOT prove
|
||||
------------------------------
|
||||
Round-trip is **necessary, not sufficient** for fluency. It measures *mutual
|
||||
intelligibility*: both halves agreeing about a surface. Two halves can agree
|
||||
on an impoverished construction — ``english_fluency_ood`` accepts ``"river
|
||||
flows valley"`` as correct — and round-trip would be perfectly happy. A high
|
||||
round-trip rate therefore licenses the claim "CORE means what it says", not
|
||||
"CORE writes well". The negative corpus is what stops the lane from drifting
|
||||
into the first failure mode; nothing here stops the second, and no metric in
|
||||
this lane should ever be cited as evidence of prose quality.
|
||||
|
||||
Framework contract: ``run_lane(cases, config=None) -> LaneReport``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any
|
||||
|
||||
from generate.graph_planner import (
|
||||
ArticulationStep,
|
||||
ArticulationTarget,
|
||||
GraphNode,
|
||||
PropositionGraph,
|
||||
RhetoricalMove,
|
||||
)
|
||||
from generate.intent import IntentTag
|
||||
from generate.meaning_graph.reader import Refusal, comprehend
|
||||
from generate.realizer import realize_target
|
||||
|
||||
from evals.grammar_roundtrip.corpora import (
|
||||
GraphCase,
|
||||
SurfaceCase,
|
||||
negative_surface_cases,
|
||||
positive_graph_cases,
|
||||
positive_surface_cases,
|
||||
)
|
||||
from evals.grammar_roundtrip.projection import (
|
||||
Agreement,
|
||||
compare,
|
||||
from_meaning_graph,
|
||||
from_proposition_graph,
|
||||
)
|
||||
|
||||
#: Reading-side predicate -> categorical form letter, for the S-round-trip
|
||||
#: re-render. Only the three quantified forms have a categorical surface;
|
||||
#: ``member`` / ``less`` and friends have no ``all X are Y`` shape, and cases
|
||||
#: carrying them are reported as ``no_renderer`` rather than as a match
|
||||
#: failure — conflating "we cannot write this at all" with "we wrote it wrong"
|
||||
#: would hide which of the two is actually blocking.
|
||||
_PREDICATE_TO_FORM: dict[str, str] = {
|
||||
"subset": "A",
|
||||
"disjoint": "E",
|
||||
"intersects": "I",
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class LaneReport:
|
||||
metrics: dict[str, Any] = field(default_factory=dict)
|
||||
case_details: list[dict[str, Any]] = field(default_factory=list)
|
||||
|
||||
|
||||
def _normalize_surface(surface: str) -> str:
|
||||
"""Lowercase and drop terminal punctuation for surface comparison.
|
||||
|
||||
Case and the sentence-final period are the renderer's caller's business,
|
||||
not the clause renderer's, so comparing them would report a failure the
|
||||
grammar is not responsible for.
|
||||
"""
|
||||
return surface.strip().rstrip(".!?").strip().lower()
|
||||
|
||||
|
||||
def _build(case: GraphCase) -> tuple[ArticulationTarget, PropositionGraph]:
|
||||
nodes = tuple(
|
||||
GraphNode(
|
||||
node_id=str(n["node_id"]),
|
||||
subject=str(n["subject"]),
|
||||
predicate=str(n["predicate"]),
|
||||
obj=str(n["obj"]),
|
||||
source_intent=IntentTag.UNKNOWN,
|
||||
)
|
||||
for n in case.nodes
|
||||
)
|
||||
steps = tuple(
|
||||
ArticulationStep(
|
||||
node_id=str(n["node_id"]),
|
||||
move=RhetoricalMove.ASSERT,
|
||||
predicate=str(n["predicate"]),
|
||||
subject=str(n["subject"]),
|
||||
negated=bool(n.get("negated", False)),
|
||||
quantifier=n.get("quantifier"), # type: ignore[arg-type]
|
||||
tense=n.get("tense"), # type: ignore[arg-type]
|
||||
aspect=n.get("aspect"), # type: ignore[arg-type]
|
||||
)
|
||||
for n in case.nodes
|
||||
)
|
||||
return (
|
||||
ArticulationTarget(steps=steps, source_intent=IntentTag.UNKNOWN),
|
||||
PropositionGraph(nodes=nodes),
|
||||
)
|
||||
|
||||
|
||||
def _render_categorical(predicate: str, subject: str, obj: str) -> str | None:
|
||||
"""Re-render one canonical proposition through the SERVING renderer.
|
||||
|
||||
Deliberately reaches ``generate.proof_chain.render._categorical_clause``
|
||||
even though it is private: that function is what band v1b actually serves,
|
||||
and a lane that measured a private copy would measure something users
|
||||
never see.
|
||||
"""
|
||||
from generate.proof_chain.render import _categorical_clause
|
||||
|
||||
form = _PREDICATE_TO_FORM.get(predicate)
|
||||
if form is None:
|
||||
return None
|
||||
return _categorical_clause({"form": form, "subject": subject, "predicate": obj})
|
||||
|
||||
|
||||
def _run_graph_case(case: GraphCase) -> dict[str, Any]:
|
||||
"""G-round-trip: graph -> surface -> graph."""
|
||||
target, graph = _build(case)
|
||||
expected = from_proposition_graph(graph, target)
|
||||
surface = realize_target(target, graph).surface
|
||||
row: dict[str, Any] = {
|
||||
"case_id": case.case_id,
|
||||
"direction": "graph",
|
||||
"surface": surface,
|
||||
"wrote": bool(surface),
|
||||
"read": False,
|
||||
"refusal_reason": None,
|
||||
}
|
||||
if not surface:
|
||||
row["agreement"] = Agreement(len(expected), 0, 0, 0, 0).as_dict()
|
||||
return row
|
||||
result = comprehend(surface, source_id="grammar_roundtrip")
|
||||
if isinstance(result, Refusal):
|
||||
row["refusal_reason"] = result.reason
|
||||
row["agreement"] = Agreement(len(expected), 0, 0, 0, 0).as_dict()
|
||||
return row
|
||||
row["read"] = True
|
||||
row["agreement"] = compare(expected, from_meaning_graph(result.meaning_graph)).as_dict()
|
||||
return row
|
||||
|
||||
|
||||
def _run_surface_case(case: SurfaceCase) -> dict[str, Any]:
|
||||
"""S-round-trip: surface -> graph -> surface."""
|
||||
row: dict[str, Any] = {
|
||||
"case_id": case.case_id,
|
||||
"direction": "surface",
|
||||
"surface": case.surface,
|
||||
"note": case.note,
|
||||
"read": False,
|
||||
"rendered": None,
|
||||
"renderable": False,
|
||||
"surface_match": False,
|
||||
"refusal_reason": None,
|
||||
}
|
||||
result = comprehend(case.surface, source_id="grammar_roundtrip")
|
||||
if isinstance(result, Refusal):
|
||||
row["refusal_reason"] = result.reason
|
||||
return row
|
||||
row["read"] = True
|
||||
props = sorted(from_meaning_graph(result.meaning_graph))
|
||||
if not props:
|
||||
row["refusal_reason"] = "empty_projection"
|
||||
return row
|
||||
rendered = [_render_categorical(p.predicate, p.subject, p.obj) for p in props]
|
||||
if any(r is None for r in rendered):
|
||||
row["rendered"] = None
|
||||
return row
|
||||
row["renderable"] = True
|
||||
joined = "; ".join(r for r in rendered if r is not None)
|
||||
row["rendered"] = joined
|
||||
row["surface_match"] = _normalize_surface(joined) == _normalize_surface(case.surface)
|
||||
return row
|
||||
|
||||
|
||||
def _run_negative_case(case: SurfaceCase) -> dict[str, Any]:
|
||||
"""A negative case passes when the reader REFUSES or recovers nothing."""
|
||||
result = comprehend(case.surface, source_id="grammar_roundtrip")
|
||||
if isinstance(result, Refusal):
|
||||
return {
|
||||
"case_id": case.case_id,
|
||||
"direction": "negative",
|
||||
"surface": case.surface,
|
||||
"note": case.note,
|
||||
"rejected": True,
|
||||
"refusal_reason": result.reason,
|
||||
}
|
||||
props = from_meaning_graph(result.meaning_graph)
|
||||
return {
|
||||
"case_id": case.case_id,
|
||||
"direction": "negative",
|
||||
"surface": case.surface,
|
||||
"note": case.note,
|
||||
# Comprehending salad into ZERO propositions still counts as rejected:
|
||||
# the reader committed to no meaning. Comprehending it into one or more
|
||||
# propositions is a real failure — the lane must go red for that.
|
||||
"rejected": not props,
|
||||
"recovered": sorted(f"{p.predicate}({p.subject},{p.obj})" for p in props),
|
||||
"refusal_reason": None,
|
||||
}
|
||||
|
||||
|
||||
def _rate(numerator: int, denominator: int) -> float:
|
||||
return round(numerator / denominator, 6) if denominator else 0.0
|
||||
|
||||
|
||||
def run_lane(cases: list[dict[str, Any]] | None = None, config: Any = None) -> LaneReport: # noqa: ARG001
|
||||
"""Run the grammar round-trip lane over its own committed corpora.
|
||||
|
||||
``cases`` is accepted for framework-contract compatibility and ignored:
|
||||
the corpora are derived from committed case files and in-module tables so
|
||||
the lane's numbers are reproducible without an external split.
|
||||
"""
|
||||
graph_rows = [_run_graph_case(c) for c in positive_graph_cases()]
|
||||
surface_rows = [_run_surface_case(c) for c in positive_surface_cases()]
|
||||
negative_rows = [_run_negative_case(c) for c in negative_surface_cases()]
|
||||
|
||||
n_graph = len(graph_rows)
|
||||
wrote = sum(1 for r in graph_rows if r["wrote"])
|
||||
g_read = sum(1 for r in graph_rows if r["read"])
|
||||
expected_total = sum(int(r["agreement"]["expected"]) for r in graph_rows)
|
||||
args_total = sum(int(r["agreement"]["args_match"]) for r in graph_rows)
|
||||
preds_total = sum(int(r["agreement"]["predicates_match"]) for r in graph_rows)
|
||||
exact_total = sum(int(r["agreement"]["exact_match"]) for r in graph_rows)
|
||||
|
||||
n_surface = len(surface_rows)
|
||||
s_read = sum(1 for r in surface_rows if r["read"])
|
||||
s_renderable = sum(1 for r in surface_rows if r["renderable"])
|
||||
s_match = sum(1 for r in surface_rows if r["surface_match"])
|
||||
|
||||
n_neg = len(negative_rows)
|
||||
rejected = sum(1 for r in negative_rows if r["rejected"])
|
||||
|
||||
metrics: dict[str, Any] = {
|
||||
# --- G-round-trip: graph -> surface -> graph ---
|
||||
"graph_cases": n_graph,
|
||||
"g_write_rate": _rate(wrote, n_graph),
|
||||
"g_read_rate": _rate(g_read, n_graph),
|
||||
"g_propositions_expected": expected_total,
|
||||
"g_args_rate": _rate(args_total, expected_total),
|
||||
"g_predicates_rate": _rate(preds_total, expected_total),
|
||||
"g_exact_rate": _rate(exact_total, expected_total),
|
||||
# --- S-round-trip: surface -> graph -> surface ---
|
||||
"surface_cases": n_surface,
|
||||
"s_read_rate": _rate(s_read, n_surface),
|
||||
"s_renderable_rate": _rate(s_renderable, n_surface),
|
||||
"s_surface_match_rate": _rate(s_match, n_surface),
|
||||
# --- negative corpus: salad must be rejected ---
|
||||
"negative_cases": n_neg,
|
||||
"reject_rate": _rate(rejected, n_neg),
|
||||
}
|
||||
return LaneReport(metrics=metrics, case_details=[*graph_rows, *surface_rows, *negative_rows])
|
||||
|
||||
|
||||
__all__ = ("LaneReport", "run_lane")
|
||||
335
scripts/measure_grammar_seam.py
Normal file
335
scripts/measure_grammar_seam.py
Normal file
|
|
@ -0,0 +1,335 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Reproduce every measurement in the grammar-unification plan's §1.
|
||||
|
||||
Usage:
|
||||
uv run python scripts/measure_grammar_seam.py # all sections
|
||||
uv run python scripts/measure_grammar_seam.py --section tables
|
||||
|
||||
Each section corresponds to a numbered subsection of
|
||||
``docs/plans/grammar-unification-2026-07-26.md``. The plan quotes these
|
||||
numbers; this script is how a reader checks them instead of trusting them.
|
||||
|
||||
Read-only: imports and probes, no mutation, no writes.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import ast
|
||||
import collections
|
||||
import importlib
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
|
||||
_REPO_ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
if str(_REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(_REPO_ROOT))
|
||||
|
||||
# --- §1.3 inventory inputs ------------------------------------------------- #
|
||||
|
||||
_READ_PATH_FILES = (
|
||||
"generate/meaning_graph/reader.py",
|
||||
"generate/meaning_graph/projectors.py",
|
||||
"generate/meaning_graph/relational.py",
|
||||
"generate/proof_chain/english.py",
|
||||
"generate/proof_chain/member.py",
|
||||
"generate/proof_chain/cond_member.py",
|
||||
"generate/proof_chain/verb.py",
|
||||
"generate/proof_chain/exist.py",
|
||||
"generate/proof_chain/categorical.py",
|
||||
)
|
||||
_WRITE_PATH_FILES = (
|
||||
"generate/semantic_templates.py",
|
||||
"generate/templates.py",
|
||||
"generate/morphology.py",
|
||||
"generate/realizer.py",
|
||||
"generate/discourse_planner.py",
|
||||
"chat/register_variation.py",
|
||||
)
|
||||
|
||||
#: Groups of tables that encode the SAME linguistic fact and should agree.
|
||||
_SHOULD_AGREE = {
|
||||
"irregular plurals": (
|
||||
("generate.proof_chain.member", "_IRREGULAR_PLURALS"),
|
||||
("generate.meaning_graph.reader", "_IRREGULAR_PLURALS"),
|
||||
("generate.templates", "_IRREGULAR_PLURALS"),
|
||||
),
|
||||
"quantifier tokens": (
|
||||
("generate.proof_chain.english", "_QUANTIFIER_LEAD"),
|
||||
("generate.proof_chain.member", "_QUANTIFIER_TOKENS"),
|
||||
("generate.templates", "_PLURAL_QUANTIFIERS"),
|
||||
),
|
||||
"connective tokens": (
|
||||
("generate.proof_chain.english", "_STRUCTURAL"),
|
||||
("generate.proof_chain.member", "_CONNECTIVES"),
|
||||
("generate.proof_chain.verb", "_CONNECTIVES"),
|
||||
("generate.proof_chain.cond_member", "_CONNECTIVE_TOKENS"),
|
||||
),
|
||||
"negation-bearing tokens": (
|
||||
("generate.proof_chain.english", "_NEGATION_BEARING"),
|
||||
("generate.proof_chain.member", "_NEGATION_BEARING"),
|
||||
),
|
||||
"predicate display": (
|
||||
("generate.semantic_templates", "_PREDICATE_HUMANIZE"),
|
||||
("generate.templates", "_PREDICATE_DISPLAY"),
|
||||
),
|
||||
}
|
||||
|
||||
#: Plural-subject agreement oracle for §1.4 — hand-written English, not derived
|
||||
#: from the code under test (deriving it from the code would make it agree by
|
||||
#: construction and measure nothing).
|
||||
_AUX_PLURAL = {"is": "are", "are": "are", "has": "have", "have": "have", "belongs": "belong"}
|
||||
_BARE_BASE = {
|
||||
"causes": "cause", "evidences": "evidence", "means": "mean",
|
||||
"addresses": "address", "answers": "answer", "corrects": "correct",
|
||||
"defines": "define", "entails": "entail", "follows": "follow",
|
||||
"grounds": "ground", "implies": "imply", "orders": "order",
|
||||
"precedes": "precede", "recalls": "recall", "requires": "require",
|
||||
"reveals": "reveal", "supports": "support", "verifies": "verify",
|
||||
"contrasts with": "contrast with",
|
||||
}
|
||||
|
||||
|
||||
def _word_tables(path: str) -> dict[str, frozenset[str]]:
|
||||
"""Module-level assignments holding >=3 lowercase english-ish literals."""
|
||||
source = (_REPO_ROOT / path).read_text()
|
||||
out: dict[str, frozenset[str]] = {}
|
||||
for node in ast.parse(source).body:
|
||||
if not isinstance(node, (ast.Assign, ast.AnnAssign)):
|
||||
continue
|
||||
targets = node.targets if isinstance(node, ast.Assign) else [node.target]
|
||||
names = [t.id for t in targets if isinstance(t, ast.Name)]
|
||||
if not names:
|
||||
continue
|
||||
words = {
|
||||
sub.value.strip()
|
||||
for sub in ast.walk(node)
|
||||
if isinstance(sub, ast.Constant)
|
||||
and isinstance(sub.value, str)
|
||||
and sub.value.strip()
|
||||
and sub.value.strip().replace("_", "").replace(" ", "").replace("-", "").isalpha()
|
||||
and sub.value.strip().islower()
|
||||
}
|
||||
if len(words) >= 3:
|
||||
out[names[0]] = frozenset(words)
|
||||
return out
|
||||
|
||||
|
||||
def section_tables() -> None:
|
||||
"""§1.3 — duplication and Jaccard across the reading/writing paths."""
|
||||
print("=" * 74)
|
||||
print("§1.3 Linguistic knowledge duplicated across reading vs writing")
|
||||
print("=" * 74)
|
||||
read_tables: dict[str, frozenset[str]] = {}
|
||||
write_tables: dict[str, frozenset[str]] = {}
|
||||
for rel in _READ_PATH_FILES:
|
||||
if (_REPO_ROOT / rel).exists():
|
||||
for name, words in _word_tables(rel).items():
|
||||
read_tables[f"{pathlib.Path(rel).name}::{name}"] = words
|
||||
for rel in _WRITE_PATH_FILES:
|
||||
if (_REPO_ROOT / rel).exists():
|
||||
for name, words in _word_tables(rel).items():
|
||||
write_tables[f"{pathlib.Path(rel).name}::{name}"] = words
|
||||
|
||||
read_words = frozenset().union(*read_tables.values()) if read_tables else frozenset()
|
||||
write_words = frozenset().union(*write_tables.values()) if write_tables else frozenset()
|
||||
union = read_words | write_words
|
||||
shared = read_words & write_words
|
||||
|
||||
print(f" reading path: {len(read_tables):3d} word-tables, {len(read_words):3d} distinct words")
|
||||
print(f" writing path: {len(write_tables):3d} word-tables, {len(write_words):3d} distinct words")
|
||||
print(f" shared : {len(shared):3d} of {len(union)} union")
|
||||
print(f" JACCARD : {len(shared) / len(union):.3f}" if union else " JACCARD: n/a")
|
||||
|
||||
print("\n tables that encode the same fact:")
|
||||
for label, refs in _SHOULD_AGREE.items():
|
||||
loaded = []
|
||||
for module_name, attr in refs:
|
||||
try:
|
||||
value = getattr(importlib.import_module(module_name), attr)
|
||||
except (ImportError, AttributeError):
|
||||
continue
|
||||
keys = frozenset(value) if not isinstance(value, tuple) else frozenset(value)
|
||||
loaded.append((f"{module_name.split('.')[-1]}.{attr}", keys))
|
||||
if len(loaded) < 2:
|
||||
continue
|
||||
all_equal = all(k == loaded[0][1] for _, k in loaded)
|
||||
verdict = "IDENTICAL" if all_equal else "DIVERGE"
|
||||
sizes = ", ".join(f"{n}={len(k)}" for n, k in loaded)
|
||||
print(f" {verdict:9s} {label:26s} ({len(loaded)} copies: {sizes})")
|
||||
|
||||
|
||||
def section_agreement() -> None:
|
||||
"""§1.4 — plural-subject agreement over the predicate-display table."""
|
||||
print("\n" + "=" * 74)
|
||||
print("§1.4 Plural-subject agreement in the (eval-only) realizer")
|
||||
print("=" * 74)
|
||||
from generate.templates import _PREDICATE_DISPLAY, _inflect_predicate
|
||||
|
||||
def expected(display: str) -> str:
|
||||
head = display.split(" ", 1)[0]
|
||||
if head in _AUX_PLURAL:
|
||||
rest = display.split(" ", 1)[1] if " " in display else ""
|
||||
return (_AUX_PLURAL[head] + (" " + rest if rest else "")).strip()
|
||||
return _BARE_BASE.get(display, display)
|
||||
|
||||
wrong = []
|
||||
for key, display in sorted(_PREDICATE_DISPLAY.items()):
|
||||
got, want = _inflect_predicate(display, plural_subject=True), expected(display)
|
||||
if got != want:
|
||||
wrong.append((key, display, got, want))
|
||||
|
||||
total = len(_PREDICATE_DISPLAY)
|
||||
print(f" correct: {total - len(wrong)}/{total} WRONG: {len(wrong)}/{total}")
|
||||
for key, display, got, want in wrong:
|
||||
print(f" {key:24s} {display!r:28s} -> {got!r:26s} want {want!r}")
|
||||
|
||||
|
||||
def section_readers() -> None:
|
||||
"""§1.7 — reader-vs-reader singularization disagreement."""
|
||||
print("\n" + "=" * 74)
|
||||
print("§1.7 Do CORE's two readers agree on singularization?")
|
||||
print("=" * 74)
|
||||
from generate.meaning_graph.reader import _singularize
|
||||
from generate.proof_chain.member import _IRREGULAR_PLURALS as member_table
|
||||
from generate.templates import pluralize
|
||||
|
||||
pairs = sorted({p: s for p, s in member_table.items() if p != s}.items())
|
||||
disagree = [(p, s, _singularize(p)) for p, s in pairs if _singularize(p) != s]
|
||||
print(f" plurals the v3-MEM band reader knows: {len(pairs)}")
|
||||
print(f" MeaningGraph reader AGREES : {len(pairs) - len(disagree)}/{len(pairs)}")
|
||||
print(f" MeaningGraph reader DISAGREES : {len(disagree)}/{len(pairs)}")
|
||||
silent = [(p, s, g) for p, s, g in disagree if g is not None]
|
||||
refused = [(p, s) for p, s, g in disagree if g is None]
|
||||
print(f"\n SILENTLY WRONG singular ({len(silent)}):")
|
||||
for plural, singular, got in silent:
|
||||
print(f" {plural:10s} v3-MEM={singular:10s} MeaningGraph={got!r}")
|
||||
print(f"\n REFUSED by MeaningGraph reader ({len(refused)}):")
|
||||
print(f" {', '.join(p for p, _ in refused)}")
|
||||
|
||||
bad_write = [(p, s) for p, s in pairs if pluralize(s) != p]
|
||||
print(f"\n writer cannot reproduce {len(bad_write)}/{len(pairs)} of those plurals:")
|
||||
for plural, singular in bad_write:
|
||||
print(f" pluralize({singular!r}) -> {pluralize(singular)!r}, want {plural!r}")
|
||||
|
||||
|
||||
def section_served() -> None:
|
||||
"""§1.7 — malformed categorical surfaces reaching served output."""
|
||||
print("\n" + "=" * 74)
|
||||
print("§1.7 Malformed surfaces served by the ratified v1b categorical band")
|
||||
print("=" * 74)
|
||||
from chat.deduction_surface import deduction_grounded_surface
|
||||
|
||||
probes = ["dogs", "cats", "students", "wolves", "children", "men"]
|
||||
print(" synthetic probes:")
|
||||
for noun in probes:
|
||||
text = (
|
||||
f"All {noun} are mammals. All mammals are animals. "
|
||||
f"Therefore all {noun} are animals."
|
||||
)
|
||||
print(f" in : all {noun} are mammals ...")
|
||||
print(f" out: {deduction_grounded_surface(text)}")
|
||||
|
||||
print("\n ratified corpus cases that serve a malformed categorical clause:")
|
||||
flagged = 0
|
||||
served_categorical = 0
|
||||
for path in sorted((_REPO_ROOT / "evals/deduction_serve").glob("*/cases.jsonl")):
|
||||
for line in path.read_text().splitlines():
|
||||
if not line.strip():
|
||||
continue
|
||||
case = json.loads(line)
|
||||
surface = deduction_grounded_surface(case["text"])
|
||||
if not surface:
|
||||
continue
|
||||
import re
|
||||
|
||||
matches = re.findall(r"\b(all|no|some)\s+([a-z_]+)\s+(are|is)\b", surface)
|
||||
if not matches:
|
||||
continue
|
||||
served_categorical += 1
|
||||
for quant, noun, _copula in matches:
|
||||
if re.search(rf"\b{quant}\s+{noun}s?\b", case["text"], re.I) and not re.search(
|
||||
rf"\b{quant}\s+{noun}\b(?!s)", case["text"], re.I
|
||||
):
|
||||
flagged += 1
|
||||
print(f" [{case['id']}] {surface}")
|
||||
break
|
||||
print(f"\n {flagged} malformed of {served_categorical} serving a categorical clause")
|
||||
|
||||
|
||||
def section_roundtrip() -> None:
|
||||
"""§1.6 — the round-trip baseline and the six-reader refusal table."""
|
||||
print("\n" + "=" * 74)
|
||||
print("§1.6 Round-trip baseline, and every reader on writer output")
|
||||
print("=" * 74)
|
||||
from evals.grammar_roundtrip.runner import run_lane
|
||||
|
||||
metrics = run_lane().metrics
|
||||
for key in (
|
||||
"graph_cases", "g_write_rate", "g_read_rate", "g_args_rate",
|
||||
"g_predicates_rate", "g_exact_rate", "surface_cases", "s_read_rate",
|
||||
"s_renderable_rate", "s_surface_match_rate", "negative_cases", "reject_rate",
|
||||
):
|
||||
print(f" {key:24s} {metrics[key]}")
|
||||
|
||||
print("\n every reader against writer output and live served surfaces:")
|
||||
from generate.meaning_graph.reader import comprehend
|
||||
from generate.proof_chain.cond_member import read_cond_member_argument
|
||||
from generate.proof_chain.english import read_english_argument
|
||||
from generate.proof_chain.exist import read_exist_argument
|
||||
from generate.proof_chain.member import read_member_argument
|
||||
from generate.proof_chain.verb import read_verb_argument
|
||||
|
||||
readers = {
|
||||
"meaning_graph": lambda t: comprehend(t, source_id="probe"),
|
||||
"english": read_english_argument,
|
||||
"member": read_member_argument,
|
||||
"cond_member": read_cond_member_argument,
|
||||
"verb": read_verb_argument,
|
||||
"exist": read_exist_argument,
|
||||
}
|
||||
targets = (
|
||||
"Molecule binds enzyme.",
|
||||
"all wolves is defined a mammal",
|
||||
"Knowledge is what a person knows from truth and evidence.",
|
||||
)
|
||||
outcomes: collections.Counter[str] = collections.Counter()
|
||||
for text in targets:
|
||||
print(f"\n {text!r}")
|
||||
for name, fn in readers.items():
|
||||
try:
|
||||
result = fn(text)
|
||||
reason = getattr(result, "reason", "") or ""
|
||||
kind = type(result).__name__
|
||||
outcomes["refused" if "Refus" in kind else "read"] += 1
|
||||
print(f" {name:14s} {kind:22s} {reason}")
|
||||
except Exception as exc: # noqa: BLE001 - a raising reader is a result
|
||||
outcomes["raised"] += 1
|
||||
print(f" {name:14s} raised {type(exc).__name__}")
|
||||
print(f"\n totals: {dict(outcomes)}")
|
||||
|
||||
|
||||
_SECTIONS = {
|
||||
"tables": section_tables,
|
||||
"agreement": section_agreement,
|
||||
"readers": section_readers,
|
||||
"served": section_served,
|
||||
"roundtrip": section_roundtrip,
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument(
|
||||
"--section", choices=sorted(_SECTIONS), action="append",
|
||||
help="run only this section (repeatable); default is all",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
for name in args.section or list(_SECTIONS):
|
||||
_SECTIONS[name]()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
276
tests/test_grammar_roundtrip.py
Normal file
276
tests/test_grammar_roundtrip.py
Normal file
|
|
@ -0,0 +1,276 @@
|
|||
"""Tests for the grammar round-trip lane.
|
||||
|
||||
The load-bearing tests here are the *mutation* tests. This lane exists because
|
||||
``evals/deterministic_fluency`` reports 1.00 on all six predicates while
|
||||
accepting ``"banana does the."`` — a pin that cannot fail is worthless. So
|
||||
every guarantee this lane makes is paired with a test proving the guarantee can
|
||||
be broken.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from evals.grammar_roundtrip import runner as rt_runner
|
||||
from evals.grammar_roundtrip.corpora import (
|
||||
negative_surface_cases,
|
||||
positive_graph_cases,
|
||||
positive_surface_cases,
|
||||
)
|
||||
from evals.grammar_roundtrip.projection import (
|
||||
CanonicalProposition,
|
||||
compare,
|
||||
from_meaning_graph,
|
||||
from_proposition_graph,
|
||||
quantifier_predicate,
|
||||
)
|
||||
from evals.grammar_roundtrip.runner import run_lane
|
||||
from generate.graph_planner import (
|
||||
ArticulationStep,
|
||||
ArticulationTarget,
|
||||
GraphNode,
|
||||
PropositionGraph,
|
||||
RhetoricalMove,
|
||||
)
|
||||
from generate.intent import IntentTag
|
||||
from generate.meaning_graph.model import (
|
||||
Entity,
|
||||
MeaningGraph,
|
||||
MeaningSpan,
|
||||
Relation,
|
||||
)
|
||||
from generate.meaning_graph.reader import Comprehension, Refusal, comprehend
|
||||
|
||||
|
||||
# The three strings that pass every content predicate of
|
||||
# evals/deterministic_fluency. This lane's rejection of them is the concrete,
|
||||
# recorded improvement over that lane.
|
||||
_FLUENCY_LANE_FALSE_POSITIVES = (
|
||||
"banana does the.",
|
||||
"wet ground rains the is.",
|
||||
"is is is is.",
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def report():
|
||||
return run_lane()
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# The negative corpus: the guarantee, and proof it can fail
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_negative_corpus_is_fully_rejected(report):
|
||||
"""Word salad must never be comprehended into propositions."""
|
||||
assert report.metrics["reject_rate"] == 1.0
|
||||
assert report.metrics["negative_cases"] >= 16
|
||||
|
||||
|
||||
def test_the_fluency_lanes_false_positives_are_rejected():
|
||||
"""The exact strings deterministic_fluency passes must be refused here."""
|
||||
for surface in _FLUENCY_LANE_FALSE_POSITIVES:
|
||||
result = comprehend(surface, source_id="t")
|
||||
if isinstance(result, Refusal):
|
||||
continue
|
||||
assert not from_meaning_graph(result.meaning_graph), (
|
||||
f"{surface!r} was comprehended into propositions; the negative "
|
||||
"corpus guarantee is broken"
|
||||
)
|
||||
|
||||
|
||||
def test_reject_rate_goes_red_when_the_reader_accepts_everything(monkeypatch):
|
||||
"""MUTATION: an accept-everything reader must drive reject_rate to 0.
|
||||
|
||||
Without this test ``reject_rate == 1.0`` would be unfalsifiable — exactly
|
||||
the defect that makes the existing fluency lane decoration.
|
||||
"""
|
||||
span = MeaningSpan(source_id="t", start=0, end=1, text="x")
|
||||
graph = MeaningGraph(
|
||||
entities=(
|
||||
Entity(entity_id="a", name="a", span=span),
|
||||
Entity(entity_id="b", name="b", span=span),
|
||||
),
|
||||
relations=(Relation(predicate="subset", arguments=("a", "b"), span=span),),
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
rt_runner, "comprehend", lambda text, source_id="input": Comprehension(meaning_graph=graph)
|
||||
)
|
||||
mutated = run_lane()
|
||||
assert mutated.metrics["reject_rate"] == 0.0, (
|
||||
"an accept-everything reader still scored a perfect reject_rate — the "
|
||||
"negative corpus is not actually gating"
|
||||
)
|
||||
|
||||
|
||||
def test_shuffled_negatives_reuse_positive_vocabulary():
|
||||
"""Shuffles must be lexically identical to a positive, order destroyed.
|
||||
|
||||
A lane that rejects only hand-authored salad could be checking vocabulary
|
||||
rather than grammar. The shuffles remove that escape.
|
||||
"""
|
||||
positives = {
|
||||
frozenset(c.surface.rstrip(".").lower().split()) for c in positive_surface_cases()
|
||||
}
|
||||
shuffles = [c for c in negative_surface_cases() if c.case_id.startswith("neg-shuf-")]
|
||||
assert shuffles, "no shuffled negatives were generated"
|
||||
for case in shuffles:
|
||||
tokens = frozenset(case.surface.rstrip(".").lower().split())
|
||||
assert tokens in positives, (
|
||||
f"{case.case_id} does not reuse a positive case's exact vocabulary"
|
||||
)
|
||||
|
||||
|
||||
def test_shuffles_are_byte_stable_across_calls():
|
||||
"""The negative corpus must be reproducible — no PRNG, no clock."""
|
||||
first = tuple(c.surface for c in negative_surface_cases())
|
||||
second = tuple(c.surface for c in negative_surface_cases())
|
||||
assert first == second
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Baselines — these pin DEFECTS and must be revised upward when fixed
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_g_roundtrip_baseline_is_zero(report):
|
||||
"""BASELINE PIN (a defect, not a goal): CORE reads 0% of what it writes.
|
||||
|
||||
Measured on main @ 9696443a. When the grammar is unified this must be
|
||||
revised **upward**; it must never be revised downward to accommodate a
|
||||
regression.
|
||||
"""
|
||||
assert report.metrics["graph_cases"] == 280
|
||||
assert report.metrics["g_write_rate"] == 1.0
|
||||
assert report.metrics["g_read_rate"] == 0.0
|
||||
assert report.metrics["g_exact_rate"] == 0.0
|
||||
|
||||
|
||||
def test_s_roundtrip_pins_the_categorical_render_defect(report):
|
||||
"""BASELINE PIN (a defect): nothing CORE reads renders back to its input.
|
||||
|
||||
The reader comprehends all 8 positive surfaces, and the serving categorical
|
||||
renderer reproduces none of them, because it interpolates singularized
|
||||
entity ids into plural templates. Phase 2B fixes this and must raise this
|
||||
number.
|
||||
"""
|
||||
assert report.metrics["s_read_rate"] == 1.0
|
||||
assert report.metrics["s_surface_match_rate"] == 0.0
|
||||
|
||||
|
||||
def test_the_categorical_render_defect_concretely():
|
||||
"""The specific defect, stated as an example rather than a rate."""
|
||||
result = comprehend("All dogs are animals.", source_id="t")
|
||||
assert isinstance(result, Comprehension)
|
||||
props = sorted(from_meaning_graph(result.meaning_graph))
|
||||
assert props == [CanonicalProposition("subset", "dog", "animal", False)]
|
||||
rendered = rt_runner._render_categorical("subset", "dog", "animal")
|
||||
# "all dog are animal" — plural template, singular entity ids.
|
||||
assert rendered == "all dog are animal"
|
||||
assert rendered != "all dogs are animals"
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# The projection itself
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_quantifier_map_matches_reader():
|
||||
"""The lane's local quantifier map must track the reader's.
|
||||
|
||||
The copy is deliberate — the lane must not import the thing it measures —
|
||||
so this test is what keeps the yardstick honest if the reader changes.
|
||||
"""
|
||||
from generate.meaning_graph.reader import _QUANTIFIER_PREDICATE as reader_map
|
||||
|
||||
for quantifier, predicate in reader_map.items():
|
||||
assert quantifier_predicate(quantifier) == predicate
|
||||
assert quantifier_predicate(None) is None
|
||||
assert quantifier_predicate("nonsense") is None
|
||||
|
||||
|
||||
def test_proposition_graph_projection_folds_quantifier():
|
||||
"""A writer graph's ``quantifier`` slot folds into the reading predicate."""
|
||||
node = GraphNode(
|
||||
node_id="n1", subject="dog", predicate="is_a", obj="animal",
|
||||
source_intent=IntentTag.UNKNOWN,
|
||||
)
|
||||
step = ArticulationStep(
|
||||
node_id="n1", move=RhetoricalMove.ASSERT, predicate="is_a",
|
||||
subject="dog", quantifier="all",
|
||||
)
|
||||
projected = from_proposition_graph(
|
||||
PropositionGraph(nodes=(node,)),
|
||||
ArticulationTarget(steps=(step,), source_intent=IntentTag.UNKNOWN),
|
||||
)
|
||||
assert projected == frozenset({CanonicalProposition("subset", "dog", "animal", False)})
|
||||
|
||||
|
||||
def test_projection_drops_non_binary_relations():
|
||||
"""Arity != 2 is dropped — the writing side cannot express it."""
|
||||
span = MeaningSpan(source_id="t", start=0, end=1, text="x")
|
||||
graph = MeaningGraph(
|
||||
entities=(
|
||||
Entity(entity_id="a", name="a", span=span),
|
||||
Entity(entity_id="b", name="b", span=span),
|
||||
Entity(entity_id="c", name="c", span=span),
|
||||
),
|
||||
relations=(
|
||||
Relation(predicate="between", arguments=("a", "b", "c"), span=span),
|
||||
Relation(predicate="subset", arguments=("a", "b"), span=span),
|
||||
),
|
||||
)
|
||||
assert from_meaning_graph(graph) == frozenset(
|
||||
{CanonicalProposition("subset", "a", "b", False)}
|
||||
)
|
||||
|
||||
|
||||
def test_compare_decomposes_argument_and_predicate_agreement():
|
||||
"""A predicate name only counts when it agrees on the SAME arguments."""
|
||||
expected = frozenset({CanonicalProposition("subset", "dog", "animal")})
|
||||
# right arguments, wrong predicate name
|
||||
args_only = frozenset({CanonicalProposition("member", "dog", "animal")})
|
||||
agreement = compare(expected, args_only)
|
||||
assert agreement.args_match == 1
|
||||
assert agreement.predicates_match == 0
|
||||
assert agreement.exact_match == 0
|
||||
|
||||
# right predicate name, wrong arguments — must NOT count
|
||||
pred_elsewhere = frozenset({CanonicalProposition("subset", "cat", "plant")})
|
||||
agreement = compare(expected, pred_elsewhere)
|
||||
assert agreement.args_match == 0
|
||||
assert agreement.predicates_match == 0
|
||||
|
||||
|
||||
def test_compare_exact_agreement():
|
||||
prop = CanonicalProposition("subset", "dog", "animal")
|
||||
agreement = compare(frozenset({prop}), frozenset({prop}))
|
||||
assert agreement.exact_rate == 1.0
|
||||
assert agreement.args_rate == 1.0
|
||||
assert agreement.predicates_rate == 1.0
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Corpus integrity
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_positive_surfaces_are_all_inside_the_reader_envelope():
|
||||
"""Every positive surface must actually be comprehended.
|
||||
|
||||
A refused positive would silently measure the reader's coverage instead of
|
||||
the round-trip, and would make s_surface_match_rate uninterpretable.
|
||||
"""
|
||||
for case in positive_surface_cases():
|
||||
result = comprehend(case.surface, source_id="t")
|
||||
assert isinstance(result, Comprehension), (
|
||||
f"{case.case_id} ({case.surface!r}) is refused; it does not belong "
|
||||
"in the positive corpus"
|
||||
)
|
||||
|
||||
|
||||
def test_graph_corpus_is_harvested_from_committed_case_files():
|
||||
cases = positive_graph_cases()
|
||||
assert len(cases) == 280
|
||||
assert all(c.nodes for c in cases)
|
||||
Loading…
Reference in a new issue