fix(generate): inflect the head verb on every branch, not just the plural two
Phase 3 fixed two of the twelve branches of `_inflect_predicate` by hand and
pinned them with a hand-written oracle. The other ten kept handing whole
predicate phrases to single-verb functions, so the same root cause was still
live on **57 of 96** (branch x multi-word predicate) pairs, 9 of 12 branches:
"belongs to" --perfective--> "has belongs toed"
"belongs to" --imperfective--> "is belongs toing"
"belongs to" --past--> "belongs toed"
"is defined as" --future--> "will is defined a"
"contrasts with" --negated--> "does not contrasts with"
Every branch now routes through `morphology.inflect_phrase_head`, which applies
a single-verb inflection to the finite verb and carries tokens 2..n through
byte-identically. Two closed tables were needed because `_base_form` is a
suffix stripper and the head of every copular predicate is a form of BE:
`base_form("is")` returned "i" and `present_participle("is")` returned "iing".
Also fixes predicate-nominal object agreement. `render_step` pluralized the
subject and never the object, so it wrote "all dogs are a mammal". The object
agrees only for a predicate nominal -- a closed set of two -- because a
prepositional object carries its own number and a productive rule would write
"all claims are grounded in evidences".
Measured
--------
tail-mangling pairs 57/96 -> 0/96
g_read_rate (round-trip) 0.0 -> 0.003413 first non-zero in the arc
grammatical_coverage v1 49/49 (2 cases corrected, see below)
english_fluency_ood 117/117 + 39/39 + 13/13 unchanged
discourse_paragraph 12/12 + 6/6 + 5/5 + 1/1 unchanged
zero_code_domain_acquisition 18/18 + 30/30 + 21/21 unchanged
`gram_C14_p01` and `gram_C14_p10` expected "...defined as compound", which is
not English. Corrected, and the old string moved into `reject_surfaces` so the
case now actively rejects what it used to accept. The correction is not my
judgment: the reader independently READS "all molecules are defined as
compounds" and REFUSES the singular form, and that is precisely what took
g_read_rate off zero.
Why the invariant instead of a bigger oracle
--------------------------------------------
A per-branch oracle has to be extended by hand for each new branch, and a
branch added without one is invisible -- which is how eight survived Phase 3.
The pin is structural instead: English marks tense, number and aspect on the
finite verb, so inflection must leave tokens 2..n byte-identical. It needs no
oracle and covers branches nobody has written yet.
It is necessary, not sufficient: "does not contrasts with" preserves its tail
perfectly and is still wrong. `test_do_support_puts_the_head_in_the_bare_
infinitive` is the sufficiency half.
Mutation
--------
baseline 84 pass
inflect_phrase_head applied to the whole phrase 39 FAIL
_IRREGULAR_BASE emptied 2 FAIL
_IRREGULAR_PARTICIPLE auxiliaries removed 2 FAIL
PREDICATIVE_NOMINAL emptied 3 FAIL
PREDICATIVE_NOMINAL widened to every copular predicate 2 FAIL
The last row is the one that matters: the plausible-but-wrong fix -- "pluralize
the object under a plural subject" -- is caught by the mass-noun control.
Not fixed here, deliberately: "has not the following steps" is archaic rather
than wrong, and modernizing it to do-support is a separate judgment call.
[Verification]: smoke 621, deductive 406, lane pins 11/11 unchanged
(grammatical_coverage is not a pinned lane; no pin was edited).
This commit is contained in:
parent
bb82029e8d
commit
7c2b753d1d
7 changed files with 354 additions and 34 deletions
|
|
@ -34,7 +34,7 @@
|
|||
{"id": "gram_C13_p01", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wisdom", "predicate": "follows", "obj": "knowledge", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["wisdom is following knowledge"], "constraints": {"must_contain": ["wisdom", "is", "following", "knowledge"], "word_order": ["wisdom", "is", "following", "knowledge"], "max_words": 8}}
|
||||
{"id": "gram_C13_p02", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "evidence", "predicate": "supports", "obj": "truth", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["evidence is supporting truth"], "constraints": {"must_contain": ["evidence", "is", "supporting", "truth"], "word_order": ["evidence", "is", "supporting", "truth"], "max_words": 8}}
|
||||
{"id": "gram_C13_p03", "construction": "C13", "construction_name": "imperfective_aspect", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "dawn", "predicate": "precedes", "obj": "day", "aspect": "imperfective"}], "edges": []}, "accept_surfaces": ["dawn is preceding day"], "constraints": {"must_contain": ["dawn", "is", "preceding", "day"], "word_order": ["dawn", "is", "preceding", "day"], "max_words": 8}}
|
||||
{"id": "gram_C14_p01", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules are defined as compound"], "constraints": {"must_contain": ["all", "molecules", "are", "defined", "as", "compound"], "word_order": ["all", "molecules", "are", "defined", "as", "compound"], "max_words": 8}, "reject_surfaces": ["all molecules is defined a compound"]}
|
||||
{"id": "gram_C14_p01", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules are defined as compounds"], "constraints": {"must_contain": ["all", "molecules", "are", "defined", "as", "compounds"], "word_order": ["all", "molecules", "are", "defined", "as", "compounds"], "max_words": 8}, "reject_surfaces": ["all molecules is defined a compound", "all molecules are defined as compound"]}
|
||||
{"id": "gram_C14_p02", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "theory", "predicate": "is_caused_by", "obj": "observation", "quantifier": "some"}], "edges": []}, "accept_surfaces": ["some theories are caused by observation"], "constraints": {"must_contain": ["some", "theories", "are", "caused", "by", "observation"], "word_order": ["some", "theories", "are", "caused", "by", "observation"], "max_words": 8}, "reject_surfaces": ["some theories is caused by observation"]}
|
||||
{"id": "gram_C14_p03", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wolf", "predicate": "belongs_to", "obj": "pack", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all wolves belong to pack"], "constraints": {"must_contain": ["all", "wolves", "belong", "to", "pack"], "word_order": ["all", "wolves", "belong", "to", "pack"], "max_words": 7}, "reject_surfaces": ["all wolfs belongs to pack"]}
|
||||
{"id": "gram_C14_p04", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "claim", "predicate": "is_grounded_in", "obj": "evidence", "quantifier": "most"}], "edges": []}, "accept_surfaces": ["most claims are grounded in evidence"], "constraints": {"must_contain": ["most", "claims", "are", "grounded", "in", "evidence"], "word_order": ["most", "claims", "are", "grounded", "in", "evidence"], "max_words": 8}, "reject_surfaces": ["most claims is grounded in evidence"]}
|
||||
|
|
@ -43,7 +43,7 @@
|
|||
{"id": "gram_C14_p07", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "axiom", "predicate": "is_distinguished_from", "obj": "theorem", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all axioms are distinguished from theorem"], "constraints": {"must_contain": ["all", "axioms", "are", "distinguished", "from", "theorem"], "word_order": ["all", "axioms", "are", "distinguished", "from", "theorem"], "max_words": 8}, "reject_surfaces": ["all axioms is distinguished from theorem"]}
|
||||
{"id": "gram_C14_p08", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "evidence", "predicate": "is_grounded_in", "obj": "truth", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all evidence is grounded in truth"], "constraints": {"must_contain": ["all", "evidence", "is", "grounded", "in", "truth"], "word_order": ["all", "evidence", "is", "grounded", "in", "truth"], "max_words": 8}}
|
||||
{"id": "gram_C14_p09", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "knowledge", "predicate": "is_defined_as", "obj": "justified", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all knowledge is defined as justified"], "constraints": {"must_contain": ["all", "knowledge", "is", "defined", "as", "justified"], "word_order": ["all", "knowledge", "is", "defined", "as", "justified"], "max_words": 8}}
|
||||
{"id": "gram_C14_p10", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all molecules are not defined as compound"], "constraints": {"must_contain": ["all", "molecules", "are", "not", "defined", "as", "compound"], "word_order": ["all", "molecules", "are", "not", "defined", "as", "compound"], "max_words": 9}, "reject_surfaces": ["all molecules do not is defined a compound"]}
|
||||
{"id": "gram_C14_p10", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "is_defined_as", "obj": "compound", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all molecules are not defined as compounds"], "constraints": {"must_contain": ["all", "molecules", "are", "not", "defined", "as", "compounds"], "word_order": ["all", "molecules", "are", "not", "defined", "as", "compounds"], "max_words": 9}, "reject_surfaces": ["all molecules do not is defined a compound", "all molecules are not defined as compound"]}
|
||||
{"id": "gram_C14_p11", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "wolf", "predicate": "belongs_to", "obj": "pack", "quantifier": "all", "negated": true}], "edges": []}, "accept_surfaces": ["all wolves do not belong to pack"], "constraints": {"must_contain": ["all", "wolves", "do", "not", "belong", "to", "pack"], "word_order": ["all", "wolves", "do", "not", "belong", "to", "pack"], "max_words": 9}, "reject_surfaces": ["all wolfs do not belongs to pack"]}
|
||||
{"id": "gram_C14_p12", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "molecule", "predicate": "causes", "obj": "reaction", "quantifier": "all"}], "edges": []}, "accept_surfaces": ["all molecules cause reaction"], "constraints": {"must_contain": ["all", "molecules", "cause", "reaction"], "word_order": ["all", "molecules", "cause", "reaction"], "max_words": 6}, "reject_surfaces": ["all molecules caus reaction"]}
|
||||
{"id": "gram_C14_p13", "construction": "C14", "construction_name": "quantified_copular", "proposition_graph": {"nodes": [{"node_id": "n1", "subject": "observation", "predicate": "evidences", "obj": "claim", "quantifier": "many"}], "edges": []}, "accept_surfaces": ["many observations evidence claim"], "constraints": {"must_contain": ["many", "observations", "evidence", "claim"], "word_order": ["many", "observations", "evidence", "claim"], "max_words": 6}}
|
||||
|
|
|
|||
|
|
@ -338,3 +338,27 @@ MASS_NOUNS: Final[frozenset[str]] = frozenset(
|
|||
}
|
||||
)
|
||||
|
||||
|
||||
#: Predicates whose object is a **predicate nominal** — a second name for the
|
||||
#: subject's category rather than an independent noun phrase. English makes the
|
||||
#: object agree in number with the subject in exactly this construction:
|
||||
#:
|
||||
#: all dogs are mammals (not "are a mammal")
|
||||
#: all molecules are defined as compounds
|
||||
#:
|
||||
#: and does **not** elsewhere, because a prepositional object carries its own
|
||||
#: number, chosen by the speaker and not by the subject:
|
||||
#:
|
||||
#: all claims are grounded in evidence (mass — never "evidences")
|
||||
#: some theories are caused by observation (generic singular is fine)
|
||||
#:
|
||||
#: So this is a CLOSED SET, not a productive rule — the same discipline
|
||||
#: ``VES_PLURAL_SINGULARS`` enforces for f/fe → ves. A rule that pluralized
|
||||
#: every object under a plural subject would produce "grounded in evidences".
|
||||
#:
|
||||
#: Deliberately excluded: ``is_distinguished_from``, ``is_caused_by``,
|
||||
#: ``is_grounded_in``. Their objects read as generic and the reader refuses the
|
||||
#: whole construction anyway (``reserved_word_in_np``), so widening the set buys
|
||||
#: no round-trip and commits to a number English leaves open.
|
||||
PREDICATIVE_NOMINAL: Final[frozenset[str]] = frozenset({"is_a", "is_defined_as"})
|
||||
|
||||
|
|
|
|||
|
|
@ -15,6 +15,8 @@ mammal``.
|
|||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Callable
|
||||
|
||||
from generate.lexicon import (
|
||||
INVARIANT_NUMBER,
|
||||
VES_PLURAL_SINGULARS,
|
||||
|
|
@ -152,13 +154,40 @@ _IRREGULAR_FORMS: dict[str, tuple[str, str]] = {
|
|||
_IRREGULAR_PAST: dict[str, str] = {v: forms[0] for v, forms in _IRREGULAR_FORMS.items()}
|
||||
|
||||
_IRREGULAR_PARTICIPLE: dict[str, str] = {
|
||||
# Present-participle (-ing) is almost always regular. Only handle
|
||||
# the truly weird cases (lie→lying handled by the suffix rule;
|
||||
# be→being is the one English present-participle that needs a
|
||||
# special entry, but `is` doesn't normally surface as a content
|
||||
# predicate in our realizer pipeline).
|
||||
# Present-participle (-ing) is almost always regular (lie→lying is handled
|
||||
# by the suffix rule). The auxiliaries are the exception, and they DO
|
||||
# surface as predicate heads: every copular predicate in
|
||||
# ``PREDICATE_DISPLAY`` begins with "is"/"has" ("is defined as", "has the
|
||||
# following steps"), so the imperfective branch inflects them constantly.
|
||||
# Without these entries ``present_participle("is")`` fell through
|
||||
# ``_base_form("is") == "i"`` and produced **"iing"**.
|
||||
"is": "being",
|
||||
"are": "being",
|
||||
"was": "being",
|
||||
"were": "being",
|
||||
"has": "having",
|
||||
"have": "having",
|
||||
"does": "doing",
|
||||
"do": "doing",
|
||||
}
|
||||
|
||||
#: 3sg present → bare infinitive, for the verbs whose base is not the stem left
|
||||
#: behind by stripping ``-s``. ``_base_form`` is a suffix stripper, so without
|
||||
#: this table ``base_form("is")`` returned **"i"** and the future branch emitted
|
||||
#: "will i defined as". Same closed-set discipline as the ``ves`` plurals: a
|
||||
#: table, not a rule, because the rule has no way to know.
|
||||
_IRREGULAR_BASE: dict[str, str] = {
|
||||
"is": "be", "are": "be", "was": "be", "were": "be", "am": "be",
|
||||
"has": "have", "have": "have", "had": "have",
|
||||
"does": "do", "do": "do", "did": "do",
|
||||
}
|
||||
|
||||
#: Heads that take a bare ``not`` rather than do-support. "was not defined as",
|
||||
#: never "did not be defined as"; but "did not belong to", never "belonged not".
|
||||
_BARE_NOT_HEADS: frozenset[str] = frozenset(
|
||||
{"is", "are", "was", "were", "has", "have", "had", "does", "do", "did"}
|
||||
)
|
||||
|
||||
_IRREGULAR_PAST_PARTICIPLE: dict[str, str] = {v: forms[1] for v, forms in _IRREGULAR_FORMS.items()}
|
||||
|
||||
|
||||
|
|
@ -179,6 +208,8 @@ _ES_STEM_ENDINGS = ("ss", "sh", "ch", "x", "z", "o")
|
|||
|
||||
|
||||
def _base_form(verb_3sg: str) -> str:
|
||||
if verb_3sg in _IRREGULAR_BASE:
|
||||
return _IRREGULAR_BASE[verb_3sg]
|
||||
if verb_3sg in _IES_KEEP_IE:
|
||||
return verb_3sg[:-1]
|
||||
if verb_3sg.endswith("ies"):
|
||||
|
|
@ -207,24 +238,48 @@ def plural_present(verb_3sg: str) -> str:
|
|||
return _base_form(verb_3sg)
|
||||
|
||||
|
||||
def inflect_phrase_head(phrase: str, inflect: Callable[[str], str]) -> str:
|
||||
"""Apply a SINGLE-VERB inflection to a predicate phrase's finite verb.
|
||||
|
||||
English marks tense, number and aspect on the finite verb, which is the
|
||||
first token of every humanized predicate ("is defined as", "has the
|
||||
following steps", "belongs to", "contrasts with"). Only that token
|
||||
inflects; tokens 2..n are carried through **byte-identical**.
|
||||
|
||||
That tail-preservation is a falsifiable invariant, and it is the one this
|
||||
module kept violating. Every function in here — ``base_form``,
|
||||
``past_tense``, ``present_participle``, ``past_participle`` — is written
|
||||
for a single verb, and ``_inflect_predicate`` was handing them whole
|
||||
phrases on nine of its ten branches. Phase 3 fixed the two plural branches
|
||||
by hand; the other eight still produced "belongs toed", "has belongs toed",
|
||||
"is belongs toing" and "will is defined a". Routing every branch through
|
||||
this one function is what makes the invariant checkable in one place
|
||||
instead of eight.
|
||||
"""
|
||||
if not phrase:
|
||||
return phrase
|
||||
head, sep, rest = phrase.partition(" ")
|
||||
return inflect(head) + sep + rest
|
||||
|
||||
|
||||
def agree_plural_phrase(phrase: str) -> str:
|
||||
"""Put a whole predicate PHRASE into plural agreement.
|
||||
|
||||
Number is marked on the finite verb, which is the first token of every
|
||||
humanized predicate ("is defined as", "has the following steps",
|
||||
"belongs to", "contrasts with"). Only that token inflects; the rest is
|
||||
carried through untouched.
|
||||
|
||||
This exists because :func:`base_form` is a SINGLE-VERB function and was
|
||||
being applied to whole phrases, stripping the last character-class of the
|
||||
final word: "is defined as" -> "is defined a", "has the following steps"
|
||||
-> "has the following step". Nine of the 26 seed predicates were wrong
|
||||
that way, and every multi-word one was.
|
||||
"""
|
||||
return inflect_phrase_head(phrase, plural_present)
|
||||
|
||||
|
||||
def takes_bare_not(phrase: str) -> bool:
|
||||
"""True when negation attaches directly to the head ("was not defined as")
|
||||
rather than through do-support ("did not belong to")."""
|
||||
if not phrase:
|
||||
return phrase
|
||||
head, sep, rest = phrase.partition(" ")
|
||||
return plural_present(head) + sep + rest
|
||||
return False
|
||||
return phrase.partition(" ")[0] in _BARE_NOT_HEADS
|
||||
|
||||
|
||||
def past_tense(verb_3sg: str) -> str:
|
||||
|
|
|
|||
|
|
@ -16,6 +16,7 @@ from generate.lexicon import (
|
|||
IRREGULAR_PLURALS,
|
||||
PLURAL_QUANTIFIERS,
|
||||
PREDICATE_DISPLAY,
|
||||
PREDICATIVE_NOMINAL,
|
||||
)
|
||||
from generate.articulation_legality import (
|
||||
ArticulationLegality,
|
||||
|
|
@ -25,11 +26,13 @@ from generate.graph_planner import RhetoricalMove
|
|||
from generate.morphology import (
|
||||
agree_plural_phrase,
|
||||
base_form,
|
||||
inflect_phrase_head,
|
||||
is_mass_noun,
|
||||
past_participle,
|
||||
past_tense,
|
||||
pluralize,
|
||||
present_participle,
|
||||
takes_bare_not,
|
||||
)
|
||||
|
||||
|
||||
|
|
@ -54,6 +57,16 @@ def _humanize_predicate(predicate: str) -> str:
|
|||
return _PREDICATE_DISPLAY.get(predicate, predicate.replace("_", " "))
|
||||
|
||||
|
||||
def _drop_indefinite_article(predicate_h: str) -> str:
|
||||
"""Strip a trailing ``a``/``an`` from an inflected predicate.
|
||||
|
||||
``is_a`` humanizes to "is a" and pluralizes to "are a"; the article cannot
|
||||
survive a plural nominal ("all dogs are a mammals" is not English).
|
||||
"""
|
||||
head, sep, rest = predicate_h.rpartition(" ")
|
||||
return head if sep and rest in ("a", "an") else predicate_h
|
||||
|
||||
|
||||
_MOVE_TEMPLATES: dict[RhetoricalMove, str] = {
|
||||
RhetoricalMove.ASSERT: "{subject} {predicate_h} {obj}",
|
||||
RhetoricalMove.ELABORATE: "furthermore, {subject} {predicate_h} {obj}",
|
||||
|
|
@ -83,25 +96,42 @@ def _inflect_predicate(
|
|||
predicate_h.startswith(prefix)
|
||||
for prefix in ("is ", "are ", "has ", "have ", "belongs ")
|
||||
)
|
||||
base = base_form(verb)
|
||||
|
||||
match (aspect, tense, negated, plural_subject):
|
||||
# Every branch below inflects the phrase HEAD and carries tokens 2..n
|
||||
# through untouched. Phase 3 fixed only the two plural branches, so
|
||||
# these eight were still handing whole phrases to single-verb
|
||||
# functions: "belongs to" came back "has belongs toed" (perfective),
|
||||
# "is belongs toing" (imperfective), "belongs toed" (past), and
|
||||
# "is defined as" came back "will is defined a" (future).
|
||||
case ("perfective", _, _, True):
|
||||
return f"have {past_participle(verb)}"
|
||||
return f"have {inflect_phrase_head(verb, past_participle)}"
|
||||
case ("perfective", _, _, False):
|
||||
return f"has {past_participle(verb)}"
|
||||
return f"has {inflect_phrase_head(verb, past_participle)}"
|
||||
case ("imperfective", _, _, True):
|
||||
return f"are {present_participle(verb)}"
|
||||
return f"are {inflect_phrase_head(verb, present_participle)}"
|
||||
case ("imperfective", _, _, False):
|
||||
return f"is {present_participle(verb)}"
|
||||
return f"is {inflect_phrase_head(verb, present_participle)}"
|
||||
case (_, "past", True, _):
|
||||
return f"did not {base}"
|
||||
# A be/have head negates in place and carries its own past tense
|
||||
# ("was not defined as"); anything else takes do-support in the
|
||||
# past ("did not belong to"), where the head reverts to the base.
|
||||
if takes_bare_not(verb):
|
||||
past = inflect_phrase_head(verb, past_tense)
|
||||
p_head, sep, rest = past.partition(" ")
|
||||
if plural_subject:
|
||||
p_head = {"was": "were", "has": "have", "did": "did"}.get(p_head, p_head)
|
||||
return f"{p_head} not{sep}{rest}" if rest else f"{p_head} not"
|
||||
return f"did not {inflect_phrase_head(verb, base_form)}"
|
||||
case (_, "past", False, _):
|
||||
return past_tense(verb)
|
||||
past = inflect_phrase_head(verb, past_tense)
|
||||
if plural_subject:
|
||||
p_head, sep, rest = past.partition(" ")
|
||||
return {"was": "were"}.get(p_head, p_head) + sep + rest
|
||||
return past
|
||||
case (_, "future", True, _):
|
||||
return f"will not {base}"
|
||||
return f"will not {inflect_phrase_head(verb, base_form)}"
|
||||
case (_, "future", False, _):
|
||||
return f"will {base}"
|
||||
return f"will {inflect_phrase_head(verb, base_form)}"
|
||||
case (_, _, True, True):
|
||||
# Plural + negated. Agree the head first, then negate around it:
|
||||
# a plural copula takes a bare "not" ("are not defined as"), while
|
||||
|
|
@ -125,9 +155,13 @@ def _inflect_predicate(
|
|||
return "have not " + predicate_h[5:]
|
||||
if predicate_h.startswith("belongs "):
|
||||
return "does not belong " + predicate_h[8:]
|
||||
return f"is not {base}"
|
||||
return f"is not {inflect_phrase_head(verb, base_form)}"
|
||||
case (_, _, True, False):
|
||||
return f"does not {base}"
|
||||
# Do-support puts the head in the bare infinitive. This branch was
|
||||
# the ninth instance of the same defect: ``base_form`` on the whole
|
||||
# phrase left "contrasts with" untouched (no -s/-es/-ies suffix to
|
||||
# strip from "with"), yielding "does not contrasts with".
|
||||
return f"does not {inflect_phrase_head(verb, base_form)}"
|
||||
case (_, _, False, True):
|
||||
# Plural agreement on the whole phrase, not base_form() of it.
|
||||
# This is the branch the 9-of-26 defect lived in: it returned the
|
||||
|
|
@ -170,6 +204,16 @@ def render_step(
|
|||
plural_subject=plural,
|
||||
)
|
||||
obj_display = obj if obj != "<pending>" else "..."
|
||||
# A predicate nominal names the subject's category, so it agrees with the
|
||||
# subject in number and sheds its indefinite article: "all dogs are
|
||||
# mammals", never "all dogs are a mammal". Restricted to the closed
|
||||
# PREDICATIVE_NOMINAL set — a prepositional object carries its own number
|
||||
# ("all claims are grounded in evidence") and pluralizing it produces
|
||||
# "evidences". The reader accepts the agreeing form and refuses the other,
|
||||
# which is why this was the last writer-side blocker on G-round-trip.
|
||||
if plural and predicate in PREDICATIVE_NOMINAL:
|
||||
obj_display = pluralize(obj_display)
|
||||
predicate_h = _drop_indefinite_article(predicate_h)
|
||||
subject_form = pluralize(subject) if plural else subject
|
||||
subject_display = f"{quantifier} {subject_form}" if quantifier else subject_form
|
||||
return template.format(
|
||||
|
|
|
|||
|
|
@ -134,12 +134,22 @@ def test_shuffles_are_byte_stable_across_calls():
|
|||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
def test_g_roundtrip_baseline_is_zero(report):
|
||||
"""BASELINE PIN (a defect, not a goal): CORE reads 0% of what it writes.
|
||||
def test_g_roundtrip_baseline_is_one_case_of_293(report):
|
||||
"""RATCHET PIN. Was 0/293 on main @ 9696443a — CORE read *nothing* it wrote.
|
||||
|
||||
Measured on main @ 9696443a. When the grammar is unified this must be
|
||||
revised **upward**; it must never be revised downward to accommodate a
|
||||
regression.
|
||||
Phase 4 made it **1**. This must only ever be revised upward; never
|
||||
downward to accommodate a regression.
|
||||
|
||||
The one case is ``gram_C14_p01``, and what unblocked it is worth recording
|
||||
because it is the opposite of what §6 of the plan predicted. The writer was
|
||||
emitting ``all molecules are defined as compound`` — a predicate nominal
|
||||
that does not agree with its subject. The reader accepts
|
||||
``...as compounds`` and refuses ``...as compound``, so the blocker was a
|
||||
one-line **writer** defect, not the MeaningGraph/PropositionGraph type
|
||||
mismatch of §1.8.
|
||||
|
||||
The remaining 292 decompose cleanly, and none of them is a model mismatch
|
||||
either — see ``test_the_remaining_blockers_are_reader_construction_coverage``.
|
||||
"""
|
||||
# 293, not the original 280: Phase 3 added 13 quantified-copular cases
|
||||
# (construction C14) to grammatical_coverage/public/v1, and this lane
|
||||
|
|
@ -148,8 +158,42 @@ def test_g_roundtrip_baseline_is_zero(report):
|
|||
# or shrinks should require a deliberate edit here, not pass silently.
|
||||
assert report.metrics["graph_cases"] == 293
|
||||
assert report.metrics["g_write_rate"] == 1.0
|
||||
assert report.metrics["g_read_rate"] == 0.0
|
||||
assert report.metrics["g_exact_rate"] == 0.0
|
||||
assert report.metrics["g_read_rate"] >= 0.003413, "the ratchet may not go down"
|
||||
assert report.metrics["g_read_rate"] == 0.003413
|
||||
|
||||
|
||||
def test_the_remaining_blockers_are_reader_construction_coverage(report):
|
||||
"""WHY g_read_rate is 1/293 and not 293/293 — the measurement §6 turns on.
|
||||
|
||||
The plan pre-committed to reading a near-zero rate as evidence for §1.8:
|
||||
two incompatible graph models, next step an ADR. The refusal reasons say
|
||||
otherwise. Every one of them is the reader declining a CONSTRUCTION it has
|
||||
no template for — not a projection disagreeing about a graph it parsed:
|
||||
|
||||
no_template_match 289 reader has no SUBJ-VERB-OBJ template at all
|
||||
unknown_morphology 2 prepositional objects (reserved_word_in_np)
|
||||
unsupported_negation 1 reader has no negated-categorical template
|
||||
|
||||
Where a construction IS in both inventories, the round trip closes exactly
|
||||
(``s_surface_match_rate == s_renderable_rate``). So the barrier is the
|
||||
*overlap* of the two construction inventories, which is currently one
|
||||
construction wide — and that is Phase 5's item 1, not an ADR.
|
||||
"""
|
||||
reasons: dict[str, int] = {}
|
||||
for row in report.case_details:
|
||||
if "wrote" not in row:
|
||||
continue
|
||||
reasons[row.get("refusal_reason") or "READ"] = (
|
||||
reasons.get(row.get("refusal_reason") or "READ", 0) + 1
|
||||
)
|
||||
assert reasons == {
|
||||
"no_template_match": 289,
|
||||
"unknown_morphology": 2,
|
||||
"unsupported_negation": 1,
|
||||
"READ": 1,
|
||||
}
|
||||
# The load-bearing claim: not one refusal is a graph-model disagreement.
|
||||
assert "projection_mismatch" not in reasons
|
||||
|
||||
|
||||
def test_s_roundtrip_closes_for_every_renderable_surface(report):
|
||||
|
|
|
|||
|
|
@ -62,6 +62,10 @@ RECORDED_CONSUMERS: dict[str, frozenset[str]] = {
|
|||
"QUANTIFIER_LEAD": frozenset({"generate.proof_chain.english"}),
|
||||
"QUANTIFIER_TOKENS": frozenset({"generate.proof_chain.member"}),
|
||||
"PLURAL_QUANTIFIERS": frozenset({"generate.templates"}),
|
||||
# Phase 4: the closed set of predicates whose object is a predicate nominal
|
||||
# and therefore agrees in number with the subject. Closed, not productive —
|
||||
# a rule would pluralize "grounded in evidence" into "evidences".
|
||||
"PREDICATIVE_NOMINAL": frozenset({"generate.templates"}),
|
||||
"PREDICATE_DISPLAY": frozenset({
|
||||
"generate.templates",
|
||||
"generate.semantic_templates",
|
||||
|
|
|
|||
|
|
@ -217,3 +217,152 @@ def test_f_to_ves_is_a_closed_set_not_a_rule(singular: str, plural: str) -> None
|
|||
from generate.morphology import pluralize as _pluralize
|
||||
|
||||
assert _pluralize(singular) == plural
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Phase 4: the phrase-vs-single-verb defect on the EIGHT branches Phase 3 left
|
||||
# ---------------------------------------------------------------------------
|
||||
#
|
||||
# Phase 3 fixed the two plural branches of ``_inflect_predicate`` by hand and
|
||||
# pinned them with a hand-written oracle. That left the other eight branches
|
||||
# handing whole predicate phrases to single-verb functions, so the same root
|
||||
# cause was still live:
|
||||
#
|
||||
# "belongs to" --perfective--> "has belongs toed"
|
||||
# "belongs to" --imperfective-> "is belongs toing"
|
||||
# "belongs to" --past--------> "belongs toed"
|
||||
# "is defined as" --future----> "will is defined a"
|
||||
#
|
||||
# 49 of 80 (branch x multi-word predicate) pairs were wrong. A per-branch
|
||||
# oracle would have to be extended by hand every time a branch is added, and a
|
||||
# branch added without one is invisible — which is how eight of them survived
|
||||
# Phase 3. So the pin here is a STRUCTURAL INVARIANT instead:
|
||||
#
|
||||
# English marks tense, number and aspect on the FINITE VERB. Inflecting a
|
||||
# predicate phrase must leave tokens 2..n byte-identical.
|
||||
#
|
||||
# It is falsifiable, it needs no oracle, and it covers branches nobody has
|
||||
# written yet.
|
||||
|
||||
_INFLECTION_BRANCHES: list[tuple[str, dict[str, object]]] = [
|
||||
("plural", {"plural_subject": True}),
|
||||
("plural+negated", {"plural_subject": True, "negated": True}),
|
||||
("negated", {"negated": True}),
|
||||
("past", {"tense": "past"}),
|
||||
("past+plural", {"tense": "past", "plural_subject": True}),
|
||||
("past+negated", {"tense": "past", "negated": True}),
|
||||
("future", {"tense": "future"}),
|
||||
("future+negated", {"tense": "future", "negated": True}),
|
||||
("perfective", {"aspect": "perfective"}),
|
||||
("perfective+plural", {"aspect": "perfective", "plural_subject": True}),
|
||||
("imperfective", {"aspect": "imperfective"}),
|
||||
("imperfective+plural", {"aspect": "imperfective", "plural_subject": True}),
|
||||
]
|
||||
|
||||
|
||||
def _multi_word_predicates() -> list[str]:
|
||||
from generate.lexicon import PREDICATE_DISPLAY
|
||||
|
||||
return sorted({d for d in PREDICATE_DISPLAY.values() if " " in d})
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("branch", "kwargs"), _INFLECTION_BRANCHES, ids=[b for b, _ in _INFLECTION_BRANCHES])
|
||||
def test_inflection_only_touches_the_head_verb(branch: str, kwargs: dict[str, object]) -> None:
|
||||
"""Tokens 2..n of a predicate phrase survive inflection byte-identically."""
|
||||
from generate.templates import _inflect_predicate
|
||||
|
||||
violations = []
|
||||
for display in _multi_word_predicates():
|
||||
tail = display.split(" ")[1:]
|
||||
got = _inflect_predicate(display, **kwargs) # type: ignore[arg-type]
|
||||
if got.split(" ")[-len(tail):] != tail:
|
||||
violations.append((display, got))
|
||||
assert not violations, f"{branch} mangled the phrase tail: {violations}"
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("kwargs", "expected"),
|
||||
[
|
||||
({"tense": "past"}, "was defined as"),
|
||||
({"tense": "past", "plural_subject": True}, "were defined as"),
|
||||
({"tense": "past", "negated": True}, "was not defined as"),
|
||||
({"tense": "future"}, "will be defined as"),
|
||||
({"tense": "future", "negated": True}, "will not be defined as"),
|
||||
({"aspect": "perfective"}, "has been defined as"),
|
||||
({"aspect": "perfective", "plural_subject": True}, "have been defined as"),
|
||||
({"aspect": "imperfective"}, "is being defined as"),
|
||||
({"aspect": "imperfective", "plural_subject": True}, "are being defined as"),
|
||||
],
|
||||
)
|
||||
def test_copular_head_inflects_as_be(kwargs: dict[str, object], expected: str) -> None:
|
||||
"""The head of every copular predicate is a form of BE, and BE is irregular
|
||||
in all four of these paradigms. ``_base_form`` is a suffix stripper, so
|
||||
before the irregular tables ``base_form("is")`` was **"i"** and
|
||||
``present_participle("is")`` was **"iing"**."""
|
||||
from generate.templates import _inflect_predicate
|
||||
|
||||
assert _inflect_predicate("is defined as", **kwargs) == expected # type: ignore[arg-type]
|
||||
|
||||
|
||||
def test_do_support_is_used_when_the_head_is_not_an_auxiliary() -> None:
|
||||
""""did not belong to", never "did not belonged to" or "belonged not to"."""
|
||||
from generate.templates import _inflect_predicate
|
||||
|
||||
assert _inflect_predicate("belongs to", tense="past", negated=True) == "did not belong to"
|
||||
assert _inflect_predicate("belongs to", tense="past") == "belonged to"
|
||||
assert _inflect_predicate("belongs to", tense="future") == "will belong to"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Phase 4: predicate-nominal object agreement
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("predicate", "obj", "expected"),
|
||||
[
|
||||
# Predicate nominal: the object names the subject's category, so it
|
||||
# agrees in number and the indefinite article goes.
|
||||
("is_a", "mammal", "all dogs are mammals"),
|
||||
("is_defined_as", "compound", "all dogs are defined as compounds"),
|
||||
# Prepositional object: number is the speaker's, not the subject's.
|
||||
("is_grounded_in", "evidence", "all dogs are grounded in evidence"),
|
||||
("is_caused_by", "observation", "all dogs are caused by observation"),
|
||||
("belongs_to", "pack", "all dogs belong to pack"),
|
||||
],
|
||||
)
|
||||
def test_only_predicate_nominals_agree_in_number(predicate: str, obj: str, expected: str) -> None:
|
||||
"""``all dogs are a mammal`` was the last writer-side blocker on
|
||||
G-round-trip: the reader accepts "all dogs are mammals" and refuses the
|
||||
other. But the fix must NOT be "pluralize the object under a plural
|
||||
subject" — that yields "grounded in evidences". Hence a closed set."""
|
||||
assert render_step(RhetoricalMove.ASSERT, "dog", predicate, obj, quantifier="all") == expected
|
||||
|
||||
|
||||
def test_singular_subjects_keep_the_article_and_the_singular_object() -> None:
|
||||
assert render_step(RhetoricalMove.ASSERT, "dog", "is_a", "mammal") == "dog is a mammal"
|
||||
|
||||
|
||||
def test_predicative_nominal_is_a_closed_set_not_every_copular_predicate() -> None:
|
||||
"""If this set ever becomes "anything starting with is", the mass-noun
|
||||
control above starts failing."""
|
||||
from generate.lexicon import PREDICATIVE_NOMINAL
|
||||
|
||||
assert PREDICATIVE_NOMINAL == frozenset({"is_a", "is_defined_as"})
|
||||
|
||||
|
||||
def test_do_support_puts_the_head_in_the_bare_infinitive() -> None:
|
||||
"""The ninth instance of the same defect, and the one the tail invariant
|
||||
CANNOT see: "does not contrasts with" preserves the tail perfectly and is
|
||||
still wrong, because the error is on the head.
|
||||
|
||||
``base_form("contrasts with")`` returned the phrase unchanged — "with" has
|
||||
no -s/-es/-ies suffix to strip — so the 3sg -s survived do-support. A tail
|
||||
invariant is necessary, not sufficient; this is the sufficiency half.
|
||||
"""
|
||||
from generate.templates import _inflect_predicate
|
||||
|
||||
assert _inflect_predicate("contrasts with", negated=True) == "does not contrast with"
|
||||
assert _inflect_predicate("belongs to", negated=True) == "does not belong to"
|
||||
assert _inflect_predicate("causes", negated=True) == "does not cause"
|
||||
# A copular head negates in place and keeps its finite form.
|
||||
assert _inflect_predicate("is defined as", negated=True) == "is not defined as"
|
||||
|
|
|
|||
Loading…
Reference in a new issue