fix(generate): inflect the head verb on every branch, not just the plural two

Phase 3 fixed two of the twelve branches of `_inflect_predicate` by hand and
pinned them with a hand-written oracle. The other ten kept handing whole
predicate phrases to single-verb functions, so the same root cause was still
live on **57 of 96** (branch x multi-word predicate) pairs, 9 of 12 branches:

    "belongs to"      --perfective-->    "has belongs toed"
    "belongs to"      --imperfective-->  "is belongs toing"
    "belongs to"      --past-->          "belongs toed"
    "is defined as"   --future-->        "will is defined a"
    "contrasts with"  --negated-->       "does not contrasts with"

Every branch now routes through `morphology.inflect_phrase_head`, which applies
a single-verb inflection to the finite verb and carries tokens 2..n through
byte-identically. Two closed tables were needed because `_base_form` is a
suffix stripper and the head of every copular predicate is a form of BE:
`base_form("is")` returned "i" and `present_participle("is")` returned "iing".

Also fixes predicate-nominal object agreement. `render_step` pluralized the
subject and never the object, so it wrote "all dogs are a mammal". The object
agrees only for a predicate nominal -- a closed set of two -- because a
prepositional object carries its own number and a productive rule would write
"all claims are grounded in evidences".

Measured
--------
    tail-mangling pairs          57/96  ->  0/96
    g_read_rate (round-trip)     0.0    ->  0.003413   first non-zero in the arc
    grammatical_coverage v1      49/49  (2 cases corrected, see below)
    english_fluency_ood          117/117 + 39/39 + 13/13   unchanged
    discourse_paragraph          12/12 + 6/6 + 5/5 + 1/1   unchanged
    zero_code_domain_acquisition 18/18 + 30/30 + 21/21     unchanged

`gram_C14_p01` and `gram_C14_p10` expected "...defined as compound", which is
not English. Corrected, and the old string moved into `reject_surfaces` so the
case now actively rejects what it used to accept. The correction is not my
judgment: the reader independently READS "all molecules are defined as
compounds" and REFUSES the singular form, and that is precisely what took
g_read_rate off zero.

Why the invariant instead of a bigger oracle
--------------------------------------------
A per-branch oracle has to be extended by hand for each new branch, and a
branch added without one is invisible -- which is how eight survived Phase 3.
The pin is structural instead: English marks tense, number and aspect on the
finite verb, so inflection must leave tokens 2..n byte-identical. It needs no
oracle and covers branches nobody has written yet.

It is necessary, not sufficient: "does not contrasts with" preserves its tail
perfectly and is still wrong. `test_do_support_puts_the_head_in_the_bare_
infinitive` is the sufficiency half.

Mutation
--------
    baseline                                                84 pass
    inflect_phrase_head applied to the whole phrase          39 FAIL
    _IRREGULAR_BASE emptied                                   2 FAIL
    _IRREGULAR_PARTICIPLE auxiliaries removed                 2 FAIL
    PREDICATIVE_NOMINAL emptied                               3 FAIL
    PREDICATIVE_NOMINAL widened to every copular predicate    2 FAIL

The last row is the one that matters: the plausible-but-wrong fix -- "pluralize
the object under a plural subject" -- is caught by the mass-noun control.

Not fixed here, deliberately: "has not the following steps" is archaic rather
than wrong, and modernizing it to do-support is a separate judgment call.

[Verification]: smoke 621, deductive 406, lane pins 11/11 unchanged
(grammatical_coverage is not a pinned lane; no pin was edited).
This commit is contained in:
Shay 2026-07-27 11:17:38 -07:00
parent bb82029e8d
commit 7c2b753d1d
7 changed files with 354 additions and 34 deletions

View file

@ -34,7 +34,7 @@
{"id": "gram_C13_p01", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wisdom", "predicate": "follows", "obj": "knowledge", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["wisdom is following knowledge"], "constraints": {"must_contain": ["wisdom", "is", "following", "knowledge"], "word_order": ["wisdom", "is", "following", "knowledge"], "max_words": 8}}
{"id": "gram_C13_p02", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "evidence", "predicate": "supports", "obj": "truth", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["evidence is supporting truth"], "constraints": {"must_contain": ["evidence", "is", "supporting", "truth"], "word_order": ["evidence", "is", "supporting", "truth"], "max_words": 8}}
{"id": "gram_C13_p03", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "dawn", "predicate": "precedes", "obj": "day", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["dawn is preceding day"], "constraints": {"must_contain": ["dawn", "is", "preceding", "day"], "word_order": ["dawn", "is", "preceding", "day"], "max_words": 8}}
{"id": "gram_C14_p01", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules are defined as compound"], "constraints": {"must_contain": ["all", "molecules", "are", "defined", "as", "compound"], "word_order": ["all", "molecules", "are", "defined", "as", "compound"], "max_words": 8}, "reject_surfaces": ["all molecules is defined a compound"]}
{"id": "gram_C14_p01", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules are defined as compounds"], "constraints": {"must_contain": ["all", "molecules", "are", "defined", "as", "compounds"], "word_order": ["all", "molecules", "are", "defined", "as", "compounds"], "max_words": 8}, "reject_surfaces": ["all molecules is defined a compound", "all molecules are defined as compound"]}
{"id": "gram_C14_p02", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "theory", "predicate": "is_caused_by", "obj": "observation", "quantifier": "some"}], "edges": []}, "accept_surfaces": ["some theories are caused by observation"], "constraints": {"must_contain": ["some", "theories", "are", "caused", "by", "observation"], "word_order": ["some", "theories", "are", "caused", "by", "observation"], "max_words": 8}, "reject_surfaces": ["some theories is caused by observation"]}
{"id": "gram_C14_p03", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wolf", "predicate": "belongs_to", "obj": "pack", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all wolves belong to pack"], "constraints": {"must_contain": ["all", "wolves", "belong", "to", "pack"], "word_order": ["all", "wolves", "belong", "to", "pack"], "max_words": 7}, "reject_surfaces": ["all wolfs belongs to pack"]}
{"id": "gram_C14_p04", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "claim", "predicate": "is_grounded_in", "obj": "evidence", "quantifier": "most"}], "edges": []}, "accept_surfaces": ["most claims are grounded in evidence"], "constraints": {"must_contain": ["most", "claims", "are", "grounded", "in", "evidence"], "word_order": ["most", "claims", "are", "grounded", "in", "evidence"], "max_words": 8}, "reject_surfaces": ["most claims is grounded in evidence"]}
@ -43,7 +43,7 @@
{"id": "gram_C14_p07", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "axiom", "predicate": "is_distinguished_from", "obj": "theorem", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all axioms are distinguished from theorem"], "constraints": {"must_contain": ["all", "axioms", "are", "distinguished", "from", "theorem"], "word_order": ["all", "axioms", "are", "distinguished", "from", "theorem"], "max_words": 8}, "reject_surfaces": ["all axioms is distinguished from theorem"]}
{"id": "gram_C14_p08", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "evidence", "predicate": "is_grounded_in", "obj": "truth", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all evidence is grounded in truth"], "constraints": {"must_contain": ["all", "evidence", "is", "grounded", "in", "truth"], "word_order": ["all", "evidence", "is", "grounded", "in", "truth"], "max_words": 8}}
{"id": "gram_C14_p09", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "knowledge", "predicate": "is_defined_as", "obj": "justified", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all knowledge is defined as justified"], "constraints": {"must_contain": ["all", "knowledge", "is", "defined", "as", "justified"], "word_order": ["all", "knowledge", "is", "defined", "as", "justified"], "max_words": 8}}
{"id": "gram_C14_p10", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all molecules are not defined as compound"], "constraints": {"must_contain": ["all", "molecules", "are", "not", "defined", "as", "compound"], "word_order": ["all", "molecules", "are", "not", "defined", "as", "compound"], "max_words": 9}, "reject_surfaces": ["all molecules do not is defined a compound"]}
{"id": "gram_C14_p10", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all molecules are not defined as compounds"], "constraints": {"must_contain": ["all", "molecules", "are", "not", "defined", "as", "compounds"], "word_order": ["all", "molecules", "are", "not", "defined", "as", "compounds"], "max_words": 9}, "reject_surfaces": ["all molecules do not is defined a compound", "all molecules are not defined as compound"]}
{"id": "gram_C14_p11", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wolf", "predicate": "belongs_to", "obj": "pack", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all wolves do not belong to pack"], "constraints": {"must_contain": ["all", "wolves", "do", "not", "belong", "to", "pack"], "word_order": ["all", "wolves", "do", "not", "belong", "to", "pack"], "max_words": 9}, "reject_surfaces": ["all wolfs do not belongs to pack"]}
{"id": "gram_C14_p12", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "causes", "obj": "reaction", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules cause reaction"], "constraints": {"must_contain": ["all", "molecules", "cause", "reaction"], "word_order": ["all", "molecules", "cause", "reaction"], "max_words": 6}, "reject_surfaces": ["all molecules caus reaction"]}
{"id": "gram_C14_p13", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "observation", "predicate": "evidences", "obj": "claim", "quantifier": "many"}], "edges": []}, "accept_surfaces": ["many observations evidence claim"], "constraints": {"must_contain": ["many", "observations", "evidence", "claim"], "word_order": ["many", "observations", "evidence", "claim"], "max_words": 6}}

View file

@ -338,3 +338,27 @@ MASS_NOUNS: Final[frozenset[str]] = frozenset(
}
)
#: Predicates whose object is a **predicate nominal** — a second name for the
#: subject's category rather than an independent noun phrase. English makes the
#: object agree in number with the subject in exactly this construction:
#:
#: all dogs are mammals (not "are a mammal")
#: all molecules are defined as compounds
#:
#: and does **not** elsewhere, because a prepositional object carries its own
#: number, chosen by the speaker and not by the subject:
#:
#: all claims are grounded in evidence (mass — never "evidences")
#: some theories are caused by observation (generic singular is fine)
#:
#: So this is a CLOSED SET, not a productive rule — the same discipline
#: ``VES_PLURAL_SINGULARS`` enforces for f/fe → ves. A rule that pluralized
#: every object under a plural subject would produce "grounded in evidences".
#:
#: Deliberately excluded: ``is_distinguished_from``, ``is_caused_by``,
#: ``is_grounded_in``. Their objects read as generic and the reader refuses the
#: whole construction anyway (``reserved_word_in_np``), so widening the set buys
#: no round-trip and commits to a number English leaves open.
PREDICATIVE_NOMINAL: Final[frozenset[str]] = frozenset({"is_a", "is_defined_as"})

View file

@ -15,6 +15,8 @@ mammal``.
"""
from __future__ import annotations
from collections.abc import Callable
from generate.lexicon import (
INVARIANT_NUMBER,
VES_PLURAL_SINGULARS,
@ -152,13 +154,40 @@ _IRREGULAR_FORMS: dict[str, tuple[str, str]] = {
_IRREGULAR_PAST: dict[str, str] = {v: forms[0] for v, forms in _IRREGULAR_FORMS.items()}
_IRREGULAR_PARTICIPLE: dict[str, str] = {
# Present-participle (-ing) is almost always regular. Only handle
# the truly weird cases (lie→lying handled by the suffix rule;
# be→being is the one English present-participle that needs a
# special entry, but `is` doesn't normally surface as a content
# predicate in our realizer pipeline).
# Present-participle (-ing) is almost always regular (lie→lying is handled
# by the suffix rule). The auxiliaries are the exception, and they DO
# surface as predicate heads: every copular predicate in
# ``PREDICATE_DISPLAY`` begins with "is"/"has" ("is defined as", "has the
# following steps"), so the imperfective branch inflects them constantly.
# Without these entries ``present_participle("is")`` fell through
# ``_base_form("is") == "i"`` and produced **"iing"**.
"is": "being",
"are": "being",
"was": "being",
"were": "being",
"has": "having",
"have": "having",
"does": "doing",
"do": "doing",
}
#: 3sg present → bare infinitive, for the verbs whose base is not the stem left
#: behind by stripping ``-s``. ``_base_form`` is a suffix stripper, so without
#: this table ``base_form("is")`` returned **"i"** and the future branch emitted
#: "will i defined as". Same closed-set discipline as the ``ves`` plurals: a
#: table, not a rule, because the rule has no way to know.
_IRREGULAR_BASE: dict[str, str] = {
"is": "be", "are": "be", "was": "be", "were": "be", "am": "be",
"has": "have", "have": "have", "had": "have",
"does": "do", "do": "do", "did": "do",
}
#: Heads that take a bare ``not`` rather than do-support. "was not defined as",
#: never "did not be defined as"; but "did not belong to", never "belonged not".
_BARE_NOT_HEADS: frozenset[str] = frozenset(
{"is", "are", "was", "were", "has", "have", "had", "does", "do", "did"}
)
_IRREGULAR_PAST_PARTICIPLE: dict[str, str] = {v: forms[1] for v, forms in _IRREGULAR_FORMS.items()}
@ -179,6 +208,8 @@ _ES_STEM_ENDINGS = ("ss", "sh", "ch", "x", "z", "o")
def _base_form(verb_3sg: str) -> str:
if verb_3sg in _IRREGULAR_BASE:
return _IRREGULAR_BASE[verb_3sg]
if verb_3sg in _IES_KEEP_IE:
return verb_3sg[:-1]
if verb_3sg.endswith("ies"):
@ -207,24 +238,48 @@ def plural_present(verb_3sg: str) -> str:
return _base_form(verb_3sg)
def inflect_phrase_head(phrase: str, inflect: Callable[[str], str]) -> str:
"""Apply a SINGLE-VERB inflection to a predicate phrase's finite verb.
English marks tense, number and aspect on the finite verb, which is the
first token of every humanized predicate ("is defined as", "has the
following steps", "belongs to", "contrasts with"). Only that token
inflects; tokens 2..n are carried through **byte-identical**.
That tail-preservation is a falsifiable invariant, and it is the one this
module kept violating. Every function in here ``base_form``,
``past_tense``, ``present_participle``, ``past_participle`` is written
for a single verb, and ``_inflect_predicate`` was handing them whole
phrases on nine of its ten branches. Phase 3 fixed the two plural branches
by hand; the other eight still produced "belongs toed", "has belongs toed",
"is belongs toing" and "will is defined a". Routing every branch through
this one function is what makes the invariant checkable in one place
instead of eight.
"""
if not phrase:
return phrase
head, sep, rest = phrase.partition(" ")
return inflect(head) + sep + rest
def agree_plural_phrase(phrase: str) -> str:
"""Put a whole predicate PHRASE into plural agreement.
Number is marked on the finite verb, which is the first token of every
humanized predicate ("is defined as", "has the following steps",
"belongs to", "contrasts with"). Only that token inflects; the rest is
carried through untouched.
This exists because :func:`base_form` is a SINGLE-VERB function and was
being applied to whole phrases, stripping the last character-class of the
final word: "is defined as" -> "is defined a", "has the following steps"
-> "has the following step". Nine of the 26 seed predicates were wrong
that way, and every multi-word one was.
"""
return inflect_phrase_head(phrase, plural_present)
def takes_bare_not(phrase: str) -> bool:
"""True when negation attaches directly to the head ("was not defined as")
rather than through do-support ("did not belong to")."""
if not phrase:
return phrase
head, sep, rest = phrase.partition(" ")
return plural_present(head) + sep + rest
return False
return phrase.partition(" ")[0] in _BARE_NOT_HEADS
def past_tense(verb_3sg: str) -> str:

View file

@ -16,6 +16,7 @@ from generate.lexicon import (
IRREGULAR_PLURALS,
PLURAL_QUANTIFIERS,
PREDICATE_DISPLAY,
PREDICATIVE_NOMINAL,
)
from generate.articulation_legality import (
ArticulationLegality,
@ -25,11 +26,13 @@ from generate.graph_planner import RhetoricalMove
from generate.morphology import (
agree_plural_phrase,
base_form,
inflect_phrase_head,
is_mass_noun,
past_participle,
past_tense,
pluralize,
present_participle,
takes_bare_not,
)
@ -54,6 +57,16 @@ def _humanize_predicate(predicate: str) -> str:
return _PREDICATE_DISPLAY.get(predicate, predicate.replace("_", " "))
def _drop_indefinite_article(predicate_h: str) -> str:
"""Strip a trailing ``a``/``an`` from an inflected predicate.
``is_a`` humanizes to "is a" and pluralizes to "are a"; the article cannot
survive a plural nominal ("all dogs are a mammals" is not English).
"""
head, sep, rest = predicate_h.rpartition(" ")
return head if sep and rest in ("a", "an") else predicate_h
_MOVE_TEMPLATES: dict[RhetoricalMove, str] = {
RhetoricalMove.ASSERT: "{subject} {predicate_h} {obj}",
RhetoricalMove.ELABORATE: "furthermore, {subject} {predicate_h} {obj}",
@ -83,25 +96,42 @@ def _inflect_predicate(
predicate_h.startswith(prefix)
for prefix in ("is ", "are ", "has ", "have ", "belongs ")
)
base = base_form(verb)
match (aspect, tense, negated, plural_subject):
# Every branch below inflects the phrase HEAD and carries tokens 2..n
# through untouched. Phase 3 fixed only the two plural branches, so
# these eight were still handing whole phrases to single-verb
# functions: "belongs to" came back "has belongs toed" (perfective),
# "is belongs toing" (imperfective), "belongs toed" (past), and
# "is defined as" came back "will is defined a" (future).
case ("perfective", _, _, True):
return f"have {past_participle(verb)}"
return f"have {inflect_phrase_head(verb, past_participle)}"
case ("perfective", _, _, False):
return f"has {past_participle(verb)}"
return f"has {inflect_phrase_head(verb, past_participle)}"
case ("imperfective", _, _, True):
return f"are {present_participle(verb)}"
return f"are {inflect_phrase_head(verb, present_participle)}"
case ("imperfective", _, _, False):
return f"is {present_participle(verb)}"
return f"is {inflect_phrase_head(verb, present_participle)}"
case (_, "past", True, _):
return f"did not {base}"
# A be/have head negates in place and carries its own past tense
# ("was not defined as"); anything else takes do-support in the
# past ("did not belong to"), where the head reverts to the base.
if takes_bare_not(verb):
past = inflect_phrase_head(verb, past_tense)
p_head, sep, rest = past.partition(" ")
if plural_subject:
p_head = {"was": "were", "has": "have", "did": "did"}.get(p_head, p_head)
return f"{p_head} not{sep}{rest}" if rest else f"{p_head} not"
return f"did not {inflect_phrase_head(verb, base_form)}"
case (_, "past", False, _):
return past_tense(verb)
past = inflect_phrase_head(verb, past_tense)
if plural_subject:
p_head, sep, rest = past.partition(" ")
return {"was": "were"}.get(p_head, p_head) + sep + rest
return past
case (_, "future", True, _):
return f"will not {base}"
return f"will not {inflect_phrase_head(verb, base_form)}"
case (_, "future", False, _):
return f"will {base}"
return f"will {inflect_phrase_head(verb, base_form)}"
case (_, _, True, True):
# Plural + negated. Agree the head first, then negate around it:
# a plural copula takes a bare "not" ("are not defined as"), while
@ -125,9 +155,13 @@ def _inflect_predicate(
return "have not " + predicate_h[5:]
if predicate_h.startswith("belongs "):
return "does not belong " + predicate_h[8:]
return f"is not {base}"
return f"is not {inflect_phrase_head(verb, base_form)}"
case (_, _, True, False):
return f"does not {base}"
# Do-support puts the head in the bare infinitive. This branch was
# the ninth instance of the same defect: ``base_form`` on the whole
# phrase left "contrasts with" untouched (no -s/-es/-ies suffix to
# strip from "with"), yielding "does not contrasts with".
return f"does not {inflect_phrase_head(verb, base_form)}"
case (_, _, False, True):
# Plural agreement on the whole phrase, not base_form() of it.
# This is the branch the 9-of-26 defect lived in: it returned the
@ -170,6 +204,16 @@ def render_step(
plural_subject=plural,
)
obj_display = obj if obj != "<pending>" else "..."
# A predicate nominal names the subject's category, so it agrees with the
# subject in number and sheds its indefinite article: "all dogs are
# mammals", never "all dogs are a mammal". Restricted to the closed
# PREDICATIVE_NOMINAL set — a prepositional object carries its own number
# ("all claims are grounded in evidence") and pluralizing it produces
# "evidences". The reader accepts the agreeing form and refuses the other,
# which is why this was the last writer-side blocker on G-round-trip.
if plural and predicate in PREDICATIVE_NOMINAL:
obj_display = pluralize(obj_display)
predicate_h = _drop_indefinite_article(predicate_h)
subject_form = pluralize(subject) if plural else subject
subject_display = f"{quantifier} {subject_form}" if quantifier else subject_form
return template.format(

View file

@ -134,12 +134,22 @@ def test_shuffles_are_byte_stable_across_calls():
# --------------------------------------------------------------------------- #
def test_g_roundtrip_baseline_is_zero(report):
"""BASELINE PIN (a defect, not a goal): CORE reads 0% of what it writes.
def test_g_roundtrip_baseline_is_one_case_of_293(report):
"""RATCHET PIN. Was 0/293 on main @ 9696443a — CORE read *nothing* it wrote.
Measured on main @ 9696443a. When the grammar is unified this must be
revised **upward**; it must never be revised downward to accommodate a
regression.
Phase 4 made it **1**. This must only ever be revised upward; never
downward to accommodate a regression.
The one case is ``gram_C14_p01``, and what unblocked it is worth recording
because it is the opposite of what §6 of the plan predicted. The writer was
emitting ``all molecules are defined as compound`` a predicate nominal
that does not agree with its subject. The reader accepts
``...as compounds`` and refuses ``...as compound``, so the blocker was a
one-line **writer** defect, not the MeaningGraph/PropositionGraph type
mismatch of §1.8.
The remaining 292 decompose cleanly, and none of them is a model mismatch
either see ``test_the_remaining_blockers_are_reader_construction_coverage``.
"""
# 293, not the original 280: Phase 3 added 13 quantified-copular cases
# (construction C14) to grammatical_coverage/public/v1, and this lane
@ -148,8 +158,42 @@ def test_g_roundtrip_baseline_is_zero(report):
# or shrinks should require a deliberate edit here, not pass silently.
assert report.metrics["graph_cases"] == 293
assert report.metrics["g_write_rate"] == 1.0
assert report.metrics["g_read_rate"] == 0.0
assert report.metrics["g_exact_rate"] == 0.0
assert report.metrics["g_read_rate"] >= 0.003413, "the ratchet may not go down"
assert report.metrics["g_read_rate"] == 0.003413
def test_the_remaining_blockers_are_reader_construction_coverage(report):
"""WHY g_read_rate is 1/293 and not 293/293 — the measurement §6 turns on.
The plan pre-committed to reading a near-zero rate as evidence for §1.8:
two incompatible graph models, next step an ADR. The refusal reasons say
otherwise. Every one of them is the reader declining a CONSTRUCTION it has
no template for not a projection disagreeing about a graph it parsed:
no_template_match 289 reader has no SUBJ-VERB-OBJ template at all
unknown_morphology 2 prepositional objects (reserved_word_in_np)
unsupported_negation 1 reader has no negated-categorical template
Where a construction IS in both inventories, the round trip closes exactly
(``s_surface_match_rate == s_renderable_rate``). So the barrier is the
*overlap* of the two construction inventories, which is currently one
construction wide and that is Phase 5's item 1, not an ADR.
"""
reasons: dict[str, int] = {}
for row in report.case_details:
if "wrote" not in row:
continue
reasons[row.get("refusal_reason") or "READ"] = (
reasons.get(row.get("refusal_reason") or "READ", 0) + 1
)
assert reasons == {
"no_template_match": 289,
"unknown_morphology": 2,
"unsupported_negation": 1,
"READ": 1,
}
# The load-bearing claim: not one refusal is a graph-model disagreement.
assert "projection_mismatch" not in reasons
def test_s_roundtrip_closes_for_every_renderable_surface(report):

View file

@ -62,6 +62,10 @@ RECORDED_CONSUMERS: dict[str, frozenset[str]] = {
"QUANTIFIER_LEAD": frozenset({"generate.proof_chain.english"}),
"QUANTIFIER_TOKENS": frozenset({"generate.proof_chain.member"}),
"PLURAL_QUANTIFIERS": frozenset({"generate.templates"}),
# Phase 4: the closed set of predicates whose object is a predicate nominal
# and therefore agrees in number with the subject. Closed, not productive —
# a rule would pluralize "grounded in evidence" into "evidences".
"PREDICATIVE_NOMINAL": frozenset({"generate.templates"}),
"PREDICATE_DISPLAY": frozenset({
"generate.templates",
"generate.semantic_templates",

View file

@ -217,3 +217,152 @@ def test_f_to_ves_is_a_closed_set_not_a_rule(singular: str, plural: str) -> None
from generate.morphology import pluralize as _pluralize
assert _pluralize(singular) == plural
# ---------------------------------------------------------------------------
# Phase 4: the phrase-vs-single-verb defect on the EIGHT branches Phase 3 left
# ---------------------------------------------------------------------------
#
# Phase 3 fixed the two plural branches of ``_inflect_predicate`` by hand and
# pinned them with a hand-written oracle. That left the other eight branches
# handing whole predicate phrases to single-verb functions, so the same root
# cause was still live:
#
# "belongs to" --perfective--> "has belongs toed"
# "belongs to" --imperfective-> "is belongs toing"
# "belongs to" --past--------> "belongs toed"
# "is defined as" --future----> "will is defined a"
#
# 49 of 80 (branch x multi-word predicate) pairs were wrong. A per-branch
# oracle would have to be extended by hand every time a branch is added, and a
# branch added without one is invisible — which is how eight of them survived
# Phase 3. So the pin here is a STRUCTURAL INVARIANT instead:
#
# English marks tense, number and aspect on the FINITE VERB. Inflecting a
# predicate phrase must leave tokens 2..n byte-identical.
#
# It is falsifiable, it needs no oracle, and it covers branches nobody has
# written yet.
_INFLECTION_BRANCHES: list[tuple[str, dict[str, object]]] = [
("plural", {"plural_subject": True}),
("plural+negated", {"plural_subject": True, "negated": True}),
("negated", {"negated": True}),
("past", {"tense": "past"}),
("past+plural", {"tense": "past", "plural_subject": True}),
("past+negated", {"tense": "past", "negated": True}),
("future", {"tense": "future"}),
("future+negated", {"tense": "future", "negated": True}),
("perfective", {"aspect": "perfective"}),
("perfective+plural", {"aspect": "perfective", "plural_subject": True}),
("imperfective", {"aspect": "imperfective"}),
("imperfective+plural", {"aspect": "imperfective", "plural_subject": True}),
]
def _multi_word_predicates() -> list[str]:
from generate.lexicon import PREDICATE_DISPLAY
return sorted({d for d in PREDICATE_DISPLAY.values() if " " in d})
@pytest.mark.parametrize(("branch", "kwargs"), _INFLECTION_BRANCHES, ids=[b for b, _ in _INFLECTION_BRANCHES])
def test_inflection_only_touches_the_head_verb(branch: str, kwargs: dict[str, object]) -> None:
"""Tokens 2..n of a predicate phrase survive inflection byte-identically."""
from generate.templates import _inflect_predicate
violations = []
for display in _multi_word_predicates():
tail = display.split(" ")[1:]
got = _inflect_predicate(display, **kwargs) # type: ignore[arg-type]
if got.split(" ")[-len(tail):] != tail:
violations.append((display, got))
assert not violations, f"{branch} mangled the phrase tail: {violations}"
@pytest.mark.parametrize(
("kwargs", "expected"),
[
({"tense": "past"}, "was defined as"),
({"tense": "past", "plural_subject": True}, "were defined as"),
({"tense": "past", "negated": True}, "was not defined as"),
({"tense": "future"}, "will be defined as"),
({"tense": "future", "negated": True}, "will not be defined as"),
({"aspect": "perfective"}, "has been defined as"),
({"aspect": "perfective", "plural_subject": True}, "have been defined as"),
({"aspect": "imperfective"}, "is being defined as"),
({"aspect": "imperfective", "plural_subject": True}, "are being defined as"),
],
)
def test_copular_head_inflects_as_be(kwargs: dict[str, object], expected: str) -> None:
"""The head of every copular predicate is a form of BE, and BE is irregular
in all four of these paradigms. ``_base_form`` is a suffix stripper, so
before the irregular tables ``base_form("is")`` was **"i"** and
``present_participle("is")`` was **"iing"**."""
from generate.templates import _inflect_predicate
assert _inflect_predicate("is defined as", **kwargs) == expected # type: ignore[arg-type]
def test_do_support_is_used_when_the_head_is_not_an_auxiliary() -> None:
""""did not belong to", never "did not belonged to" or "belonged not to"."""
from generate.templates import _inflect_predicate
assert _inflect_predicate("belongs to", tense="past", negated=True) == "did not belong to"
assert _inflect_predicate("belongs to", tense="past") == "belonged to"
assert _inflect_predicate("belongs to", tense="future") == "will belong to"
# ---------------------------------------------------------------------------
# Phase 4: predicate-nominal object agreement
# ---------------------------------------------------------------------------
@pytest.mark.parametrize(
("predicate", "obj", "expected"),
[
# Predicate nominal: the object names the subject's category, so it
# agrees in number and the indefinite article goes.
("is_a", "mammal", "all dogs are mammals"),
("is_defined_as", "compound", "all dogs are defined as compounds"),
# Prepositional object: number is the speaker's, not the subject's.
("is_grounded_in", "evidence", "all dogs are grounded in evidence"),
("is_caused_by", "observation", "all dogs are caused by observation"),
("belongs_to", "pack", "all dogs belong to pack"),
],
)
def test_only_predicate_nominals_agree_in_number(predicate: str, obj: str, expected: str) -> None:
"""``all dogs are a mammal`` was the last writer-side blocker on
G-round-trip: the reader accepts "all dogs are mammals" and refuses the
other. But the fix must NOT be "pluralize the object under a plural
subject" — that yields "grounded in evidences". Hence a closed set."""
assert render_step(RhetoricalMove.ASSERT, "dog", predicate, obj, quantifier="all") == expected
def test_singular_subjects_keep_the_article_and_the_singular_object() -> None:
assert render_step(RhetoricalMove.ASSERT, "dog", "is_a", "mammal") == "dog is a mammal"
def test_predicative_nominal_is_a_closed_set_not_every_copular_predicate() -> None:
"""If this set ever becomes "anything starting with is", the mass-noun
control above starts failing."""
from generate.lexicon import PREDICATIVE_NOMINAL
assert PREDICATIVE_NOMINAL == frozenset({"is_a", "is_defined_as"})
def test_do_support_puts_the_head_in_the_bare_infinitive() -> None:
"""The ninth instance of the same defect, and the one the tail invariant
CANNOT see: "does not contrasts with" preserves the tail perfectly and is
still wrong, because the error is on the head.
``base_form("contrasts with")`` returned the phrase unchanged "with" has
no -s/-es/-ies suffix to strip so the 3sg -s survived do-support. A tail
invariant is necessary, not sufficient; this is the sufficiency half.
"""
from generate.templates import _inflect_predicate
assert _inflect_predicate("contrasts with", negated=True) == "does not contrast with"
assert _inflect_predicate("belongs to", negated=True) == "does not belong to"
assert _inflect_predicate("causes", negated=True) == "does not cause"
# A copular head negates in place and keeps its finite form.
assert _inflect_predicate("is defined as", negated=True) == "is not defined as"