Merge branch 'feat/tier-s-evidence-and-vocab-instrument' into main

This commit is contained in:
Shay 2026-07-24 17:37:44 -07:00
commit a1f3947228
6 changed files with 936 additions and 0 deletions

View file

@ -185,6 +185,7 @@ TEST_SUITES: dict[str, tuple[str, ...]] = {
"tests/test_deduction_serve_e2e.py",
"tests/test_curriculum_serve.py",
"tests/test_ratified_ledger_bridge.py",
"tests/test_vocab_trigger_instrument.py",
),
"full": ("tests/",),
}

View file

@ -0,0 +1,106 @@
# Curriculum-Volume Quantification — ADR-0262 §5.1
**Tier S6** (generalization-arc-2026-07-24 §6). Evidence only. This
document does not propose ratifying curriculum content or flipping
`curriculum_serving_enabled` — both are Shay's decisions
([[project-curriculum-serving-ratified]] doctrine: a flag flips on earned
evidence, not proximity to it). It answers one question precisely: *how
short of a SERVE license is each served `(subject × relation family)` band
today, measured against the actual ratified corpora, not an estimate.*
Method: `teaching.curriculum_premises.load_curriculum(domain)` — the same
loader the curriculum composer uses in serving — run against each of the
four subjects `chat.curriculum_surface.SERVED_DOMAINS` currently routes to.
Counts below are chain counts (ratified, vocabulary-resident, `review_status
== "reviewed"`), i.e. exactly the population `compile_premises` would draw
premises from; nothing here touches unmounted or unratified content.
## 1. The floor, restated precisely (§1.2 of the ratification packet)
`Action.SERVE` needs a one-sided Wilson lower bound ≥ θ=0.99
(`core/reliability_gate/ceilings.py`, `core/reliability_gate/floor.py`). For
a perfect record this reduces to `committed / (committed + z²)` with
`z = 2.576`, which crosses 0.99 at **committed ≥ 657**.
657 committed cases alone is not sufficient — the plan (§6) and the
ratification packet both require a **genuine outcome mix**, not 657
restatements of the same fact padded with paraphrase (repetition-padding
was already considered and rejected as dishonest earlier in this arc:
`conservative_floor` cannot distinguish "657 independent facts" from "16
facts asked 41 times each," but a ledger that gamed the floor that way
would not have *earned* anything). Treating "genuine mix" as roughly a
three-way split across `entailed` / `refuted` / `unknown` (the plan's own
working assumption, not a formally ratified policy — flagged here as an
assumption, not fact) puts the **entailed bucket's floor at ≈219 distinct
taught relations** — `⌈657 / 3⌉`. That is where the plan's "≈219 per
subject × family" figure comes from; this document supplies the measured
`n` for every currently-served band to compare against it.
## 2. Measured ratified-chain counts, per subject × family (2026-07-24)
| domain | family | ratified chains (n) | 657÷n | 219÷n |
|---|---|---:|---:|---:|
| physics | causal | 7 | 94× | 31× |
| physics | modal | 9 | 73× | 24× |
| mathematics_logic | modal | 4 | 164× | 55× |
| mathematics_logic | sequence | 4 | 164× | 55× |
| mathematics_logic | contrast | 8 | 82× | 27× |
| mathematics_logic | evidential | 8 | 82× | 27× |
| systems_software | causal | 7 | 94× | 31× |
| systems_software | modal | 6 | 110× | 37× |
| systems_software | sequence | 3 | 219× | 73× |
| philosophy_theology | modal | 8 | 82× | 27× |
| philosophy_theology | contrast | 8 | 82× | 27× |
11 bands are live (a subject × family pair with ≥1 ratified chain routes
questions today, per `chat.curriculum_surface.band_for`); every one is
between 24× and 73× short of the entailed-bucket floor alone, before even
authoring the matching refuted/unknown volume. **No served subject has a
family close enough to flag as "next."** Systems-software `sequence` (n=3)
is the thinnest; mathematics_logic `modal`/`sequence` (n=4 each) are tied
for second-thinnest; physics `modal` (n=9) is the closest of the eleven,
still 24× short.
Total ratified chains per subject (all families, informational context —
not a band-level floor since bands are scored per family, not per subject):
physics 16, mathematics_logic 24, systems_software 16,
philosophy_theology 16.
## 3. Vocabulary is not the constraint
`Curriculum.vocabulary` (every lemma the subject's mounted packs teach) is
16 lemmas for physics/mathematics_logic/systems_software and **151** for
philosophy_theology — an order of magnitude more vocabulary than chains in
every subject. The gate `untaught_vocabulary` (§4.6) is not what is holding
any band back; the terms are taught, the *relations between them* are not
authored in ratified volume. This confirms ADR-0262 §5.1's framing: the
blocker is curriculum content authorship, not reader/vocabulary coverage —
see also S3's mechanism-vs-coverage instrument
(`docs/research/vocab-trigger-instrument-2026-07-24.md`), which finds the
curriculum path's own refusals are overwhelmingly coverage-class for a
different reason (relation, not term, absent) while carrying zero
`untaught_vocabulary` refusals in the committed lane corpus.
## 4. What one subject fully reaching the floor would look like
Taking physics `modal` (the closest band, n=9) as an illustration only: it
would need roughly 210 more ratified `requires`/`enables` relations,
authored and reviewed through the existing curriculum ratification
pipeline, before an entailed bucket alone could support a 657-committed
band — and that is before authoring the matching refuted/unknown volume
this document is not attempting to size (refuted questions need explicit
negative curriculum content the corpora do not currently carry at all;
unknown questions are comparatively cheap to author from
in-vocabulary-but-untaught pairings, so they are not expected to be the
binding constraint).
## 5. Non-goals of this document
- Does not recommend which subject or family to grow first — that is
content strategy, Shay's call.
- Does not estimate authoring effort/time (no clock-time estimates per
standing convention).
- Does not touch `curriculum_serving_enabled` or any corpus file.
Relates to [[project-curriculum-grounded-serving]],
[[project-generalization-arc]], `docs/adr/ADR-0262-curriculum-grounded-serving.md` §5.1.

View file

@ -0,0 +1,196 @@
# Ratification Packet — `deduction_serving_enabled` and `curriculum_serving_enabled`
**Tier S1** (generalization-arc-2026-07-24 §6). Evidence collation only — the
flag flip is Shay's decision alone, never inferred from this document or from
approval of adjacent work. This packet has two independent sections: one
recording a decision already made, one recording why a decision has not been
made.
---
## 1. `deduction_serving_enabled` — RATIFIED ON (2026-07-24, main @ `5fd9a861`)
This section is a **record**, not an input to the decision. Shay ratified the
flag flip on 2026-07-24; the evidence below is what stood behind that
ratification, preserved for audit and for anyone rolling back or re-deriving
the decision later.
### 1.1 Sealed-practice ledger — 25 bands, 720/720 wrong=0 each
`chat/data/deduction_serve_ledger.json` (schema `deduction_serve_ledger_v1`,
provenance `evals.deduction_serve.practice.runner.seal_ledger`, content
`a056f899…`). The engine only *reads* this artifact — it is committed,
SHA-sealed, and never written by serving.
| # | classes | committed | correct | wrong | refused |
|---|---|---|---|---|---|
| 25 | see below | 720 each | 720 each | **0** | 0 |
Classes (25): `atomic`, `categorical`, `conditional_chain`,
`conditional_single`, `disjunctive` (Band v1/v1b) · `en_atomic`,
`en_conditional_chain`, `en_conditional_single`, `en_disjunctive` (Band
v2-EN, ADR-0257) · `en_member_atomic`, `en_member_chain`,
`en_member_negative`, `en_member_single` (Band v3-MEM, ADR-0258) ·
`en_condmem_chain`, `en_condmem_conditional`, `en_condmem_disjunctive`,
`en_condmem_fused` (Band v4-CM, ADR-0259) · `en_verb_chain`, `en_verb_fact`,
`en_verb_negative`, `en_verb_universal` (Band v5-VP, ADR-0260) ·
`en_exist_chain`, `en_exist_negative`, `en_exist_universal`,
`en_exist_witness` (Band v6-EX, ADR-0261).
### 1.2 Earned-license gate ceilings (ADR-0175 §4a, `core/reliability_gate/`)
`Action.SERVE` ceiling is θ=0.99 (`core/reliability_gate/ceilings.py`). The
floor is a pinned one-sided Wilson lower bound
(`core/reliability_gate/floor.py::conservative_floor`), deterministic and
replay-stable (no learned weights, no sampling). For a perfect record
(`successes == committed`) the bound reduces to
`committed / (committed + z²)` with `z = WILSON_Z = 2.576`; solving
`committed / (committed + z²) ≥ 0.99` gives `committed ≥ 99·z² ≈ 656.9`,
i.e. **n ≥ 657** — where the standing "n≥657 committed" figure comes from,
not a chosen round number. All 25 classes commit 720 ≥ 657 with a perfect
record, clearing the floor with margin.
### 1.3 Serve-lane results — production pipeline, independent gold
`evals/deduction_serve/report.json` runs the *exact* pipeline
`chat/deduction_surface.py` calls in serving (shape-gate → reader →
projector → ROBDD engine), scored against independently-authored gold —
**166/166 correct, 0 wrong, 0 declined**, across six splits:
| split | n | band |
|---|---|---|
| v1 | 28 | propositional / categorical |
| v2_en | 26 | English opaque-clause arguments |
| v2_member | 26 | singular-membership arguments |
| v2_condmem | 26 | conditional-membership fusion |
| v2_verb | 28 | verb-predicate arguments |
| v2_exist | 32 | existential arguments |
`wrong` staying 0 is the lane's non-negotiable gate (a decline is a coverage
miss, tracked but never conflated with a wrong answer); `all_correct=true`
confirms every committed case, not just the wrong-count, matches gold.
### 1.4 Commit gate — narrow by construction
`chat/deduction_surface.py::looks_like_deductive_argument` is a single
regex: a sentence-initial `"therefore"` conclusion clause
(`_ARGUMENT_CONCLUSION_RE`). This is a genuine signal of intent — text
shaped that way IS an argument — so the composer can commit on shape alone
without risking hijacking ordinary turns. (Contrast with curriculum's gate,
§2.4 below, which cannot use shape alone.) Once committed, every branch is
fail-closed (INV-34): a decided verdict, a typed reader refusal, or the
generic out-of-band surface — never a silent fall-through to a different
composer.
### 1.5 Blast radius (verified before the flip, not assumed)
- Deductive suite (`core test --suite deductive`) — 268 passed with the flag
on.
- `chat/runtime.py::_maybe_pack_grounded_surface` — the deduction branch is
checked **before** the empty-vault gate (unlike pack/teaching/partial/oov):
a pure function of input text, independent of vault/field warm-or-cold
state, so it cannot perturb any other dispatch path's preconditions. A
claimed turn sets `grounding_source="deduction"`; every other path is
reached exactly as before when the composer declines
(`DispatchAttempt(source="deduction", outcome="skipped", ...)`).
- Lane pins (`scripts/verify_lane_shas.py`) stayed 10/11 across the flip —
unchanged, i.e. the flip touched nothing any other pinned lane measures.
### 1.6 Rollback
One line: `deduction_serving_enabled: bool = True``False` in
`core/config.py`. Flag-off is byte-identical to pre-arc dispatch — pinned by
`tests/test_deduction_surface.py::test_flag_off_preserves_pack_token_gloss_byte_identity`.
No residue: the composer is a pure function gated entirely by the one flag
read at `chat/runtime.py:1748`; nothing it touches persists state.
### 1.7 Unearned shapes stay disclosed, not withheld
Every band above is licensed. A future band, or a deployment whose ledger is
stripped, still gets the sound ROBDD answer — served DISCLOSED via the
`_UNVERIFIED_SHAPE_DISCLOSURE` hedge (`chat/deduction_surface.py`), never
presented as verified. This is what makes serving an *earned* capability
rather than a merely flagged one, and it is why disclosure is not itself
evidence against ratifying a *licensed* band.
---
## 2. `curriculum_serving_enabled` — stays OFF (ADR-0262)
This section exists to make the non-decision as auditable as the decision
above: **every curriculum band is unearned**, so flipping this flag would
enable a capability that answers *disclosed*, never authoritatively — sound,
but with no demonstrated track record on the pipeline serving it.
### 2.1 Same shape of evidence, opposite conclusion
`evals/curriculum_serve/report.json`: physics lane **32/32 correct, 0 wrong,
0 declined**, `all_correct=true`, 5 anti-recall probes (≥ the §4.7 floor of
3). This is a clean lane pass — the mechanism is proven — but it says
nothing about *volume*, which is the actual gate.
### 2.2 Why no band can earn SERVE from present data (ADR-0262 §5.1)
The reliability floor (§1.2 above) needs **n ≥ 657 committed** per band, and
a band here is `(subject × relation family)` — e.g.
`curriculum_physics_causal`. Physics teaches 7 causal + 9 modal relations
total (§6 below has the full count), so **at most 16 questions in the
entire subject can ever come back `entailed`** — roughly 25× short of a
single band's floor, before even asking for a *mix* of entailed/refuted/
unknown outcomes (a real reliability claim needs more than one outcome
class repeated). This is a volume ceiling on the ratified curriculum, not a
defect in the reader, the bands, or the engine — all of which are the same
verified components §1 already licensed.
### 2.3 The ledger says it plainly
Every one of the curriculum bands present in a ledger would show `committed`
in the low tens, nowhere near 657. (No curriculum-specific sealed ledger
exists yet, precisely because sealing one at this volume would just
document the shortfall — see [[project-curriculum-grounded-serving]] and
S6, `docs/research/curriculum-volume-quantification-2026-07-24.md`.)
### 2.4 Commit gate — routability, not shape (ADR-0262 §5.3, already fixed)
Unlike deduction's `"therefore"` gate, curriculum's `"Does …?"` shape is not
itself a signal of intent — it is one of the most common ways to open *any*
English question. `chat/curriculum_surface.py::is_curriculum_question`
commits only when the question parses to three tokens AND both terms are
vocabulary some served subject actually teaches (`resolve_domain`); an
`untaught_vocabulary` refusal is explicitly **not** a claim (the composer
declines the turn entirely, byte-identical to before this composer
existed). This was a real defect found and fixed this arc (commit-gate
firing on any `Does …?` text, e.g. "Does the build pass?"); the fix is
already on `main` and is not part of what remains gated by this flag.
### 2.5 What flipping ON today would mean
Every served answer would carry the `_UNVERIFIED_SHAPE_DISCLOSURE` hedge —
"(answered from my curriculum, but I haven't yet earned a verified track
record on questions of this shape)" — because `_license_gate`
(`chat/curriculum_surface.py`) finds no `Action.SERVE` license for any
`curriculum_<domain>_<family>` band. The capability would be real and
sound (same argument bands, same ROBDD engine as §1) but never
authoritative. This is not unsafe to ship; it is simply not yet *earned*,
and the standing rule (Shay, 2026-07-24) is that a flag flips on earned
evidence, not on proximity to it — see
[[project-deduction-serving-ratified]].
### 2.6 Rollback / blast radius, for parity with §1.6
Same shape as deduction: one line in `core/config.py`
(`curriculum_serving_enabled: bool = False`, already the default — nothing
to roll back), checked immediately after the deduction branch in
`chat/runtime.py:1764`, pure function of text + ratified curriculum, no
vault/field dependency. Flag-off is byte-identical to pre-arc dispatch by
construction (the branch is dead code when the flag is False).
---
## 3. What would close §2 (evidence only — not a proposal to act)
Nothing in Tier S closes this gap; it closes only when ratified curriculum
content grows roughly 25× per served `(subject × family)` — a content
question for Shay, not an engineering one. See S6's quantification,
`docs/research/curriculum-volume-quantification-2026-07-24.md`, for the
per-subject arithmetic.

View file

@ -0,0 +1,108 @@
# Vocab-Trigger Instrument — Baseline Measurement
**Tier S3** (generalization-arc-2026-07-24 §6). Implements the measurable
test named in `docs/handoffs/COMPREHENSION-READER-AUDIT.md` (§Q5) — a
refusal histogram split mechanism-vs-coverage plus an admissions-per-batch
counter — generalized from the math reader (the audit's original subject)
to the two composers this arc actually ships: `chat/deduction_surface.py`
and `chat/curriculum_surface.py`. Implementation:
`evals/vocab_trigger_instrument.py`; tests:
`tests/test_vocab_trigger_instrument.py` (11, registered in the
`deductive` suite).
This document does not propose any pack/curriculum expansion. It reports
where the two composers stand today and hands over a rerunnable
instrument, per the standing doctrine
([[project-base-knowledge-pack-expansion-trigger]]): pack/base-knowledge
expansion becomes load-bearing exactly when refusals shift from mechanism
scope to vocabulary scope — surface that shift, don't act on it here.
## 1. The rule (closed, not invented)
A decline is one of three classes:
- **mechanism** — the reader/gate cannot parse the SHAPE (a closed-grammar
gap: `sentence_shape_out_of_band`, `quantifier_out_of_band`, ...).
- **coverage** — a term or relation the ratified content never taught
(`untaught_vocabulary`, `out_of_curriculum` — the curriculum composer's
own naming, taken verbatim, not reinterpreted).
- **engine_refused** — a fully-read, well-formed argument the ROBDD engine
itself declines (`inconsistent_premises`, `out_of_regime_or_malformed`
from `generate/proof_chain/entail.py`). Neither a shape gap nor a
vocabulary gap; no lexicon or curriculum batch moves this bucket.
This third bucket was not in the brief's two-way framing and was found
while building the instrument: an early version misattributed
engine-declined arguments (read successfully, but self-contradictory) to
whichever reader band happened to be tried last in the fallback cascade —
a real bug, not a design choice, caught by `test_deduction_refusal_reason_recovers_engine_level_decline`
pinning `"p. Not p. Therefore q."` to `(band="v1", reason="inconsistent_premises")`
rather than a v6-EX shape reason it never actually hit.
## 2. Baseline measurement (2026-07-24, committed lane corpora)
Run: `uv run python -m evals.vocab_trigger_instrument`. Deterministic —
pure over `evals/deduction_serve/*/cases.jsonl`,
`evals/curriculum_serve/*/cases.jsonl`, and the same production decision
paths the lanes score against (`test_build_report_is_deterministic` pins
byte-identical output across repeat runs).
| composer | n | admitted | declined | mechanism | coverage | engine_refused |
|---|---:|---:|---:|---:|---:|---:|
| deduction | 166 | 126 | 40 | 35 | **0** | 5 |
| curriculum | 32 | 26 | 6 | 2 | 4 | 0 |
| **combined** | 198 | 152 | 46 | 37 | 4 | 5 |
### 2.1 Deduction: zero coverage-class refusals, and why that's correct
Every one of the 40 declined deduction cases is mechanism (35, closed-shape
gates — the largest single bucket is `v6_exist:mixed_structure_out_of_band`
at 11) or engine_refused (5, all `inconsistent_premises` — 4 pinned
Band-v1 contradiction probes plus one natural-English contradiction admitted
by Band v2-EN). **Zero coverage refusals is the honest count, not a gap in
this instrument**: none of the seven deduction bands read open vocabulary —
they read a small closed set of connective/quantifier words
(`generate/proof_chain/{english,member,cond_member,verb,exist}.py`), so
there is no vocabulary axis for them to refuse on. A future band that reads
open English prose (rather than a closed relation grammar) would be the
first one this instrument could ever show a nonzero coverage count for.
### 2.2 Curriculum: the coverage axis is real and already exercised
4 of the physics split's 6 declines are coverage (3 `untaught_vocabulary`,
1 `out_of_curriculum`); the other 2 are mechanism
(`question_shape_out_of_band`). This is the axis
[[project-curriculum-grounded-serving]] and S6's volume quantification
already describe from the ratified-chain-count side; this instrument
confirms it from the refusal side using the SAME production decision path,
independently.
## 3. How to use this for a future batch
```
uv run python -m evals.vocab_trigger_instrument --report before.json
# ... ratify new curriculum content or widen a band's grammar ...
uv run python -m evals.vocab_trigger_instrument --report after.json
uv run python -m evals.vocab_trigger_instrument --compare before.json --report after.json
```
The last command prints the admitted-count delta per composer — the
falsifiable measurement Option A of the audit's Q5 calls for, generalized:
a real batch should move `admitted` up and the matching declined-class
count down. No claim is made here about what that delta *should* be for
any specific batch; that is content strategy, not this instrument's job.
## 4. Non-goals
- Does not run against live chat traffic — only the committed, sealed lane
corpora (adding a live-traffic feed would be a different, larger
instrument and is not requested here).
- Does not add a new SHA-pinned CI lane; the report is deterministic and
could be pinned later if a gate on this split is wanted, but nothing
requested that for Tier S.
- Does not reclassify `ambiguous_reading` / `empty_curriculum` as coverage
even though they are vocabulary-adjacent — only the two reasons the
brief named explicitly are coverage-class (§1).
Relates to [[project-base-knowledge-pack-expansion-trigger]],
[[project-curriculum-grounded-serving]], [[project-generalization-arc]].

View file

@ -0,0 +1,407 @@
"""evals/vocab_trigger_instrument.py — mechanism-vs-coverage refusal instrument.
The measurable test from `docs/handoffs/COMPREHENSION-READER-AUDIT.md`
(Q5 "measurable tests") generalized to the two composers this arc actually
ships: split declined turns by WHY they declined a closed-grammar SHAPE
the reader cannot yet parse ("mechanism"), or a TERM/RELATION the ratified
content never taught ("coverage") and count admissions, so a future
lexicon/curriculum batch's effect can be measured as a before/after delta
instead of asserted.
This is the same axis `project-base-knowledge-pack-expansion-trigger`
(memory) already names for the math reader: pack/base-knowledge expansion
becomes load-bearing exactly when refusals shift from mechanism scope to
vocabulary scope. This module makes that shift countable for the deduction
and curriculum composers, using their own closed, typed refusal reasons
no heuristic text-matching, no invented policy.
**Three buckets, not two.** A decline is "mechanism" (an unreadable
closed-grammar SHAPE), "coverage" (an absent TERM/RELATION), or
"engine_refused" (a fully-read, well-formed argument the ROBDD engine
itself declines inconsistent premises, or out of the decidable regime;
`generate/proof_chain/entail.py`'s own REFUSED reasons). The third bucket
is neither a shape gap nor a vocabulary gap; no lexicon or curriculum batch
can move it, so folding it into "mechanism" would misdirect exactly the
signal this instrument exists to make legible.
**Deduction composer every READER/gate refusal is mechanism-class.**
All ~20 typed shape reasons across the six bands
(`generate/proof_chain/{english,member,cond_member,verb,exist}.py`, plus
the shared reader `generate/meaning_graph/reader.py`) name a closed-grammar
SHAPE the reader will not attempt (``sentence_shape_out_of_band``,
``quantifier_out_of_band``, ``mixed_structure_out_of_band``, ...). None of
them names a missing vocabulary term these bands read a small closed set
of connective/quantifier words, not open vocabulary, so there is no
vocabulary axis for them to refuse on yet. A future band that reads open
English prose would change this; today this instrument would correctly
report 0 coverage-class deduction refusals, and that is the honest count,
not a bug in the classifier. A deduction argument CAN still decline
engine-side (inconsistent premises) after being read just fine by any band,
including Band v1 itself see ``deduction_refusal_reason``.
**Curriculum composer has a real coverage axis.**
`chat/curriculum_surface.py::resolve_domain` refuses ``untaught_vocabulary``
when a query's subject/object term is not in ANY served subject's mounted
vocabulary, and `resolve_family` refuses ``out_of_curriculum`` when the
relation itself was never taught. Both name an absent TERM/RELATION, not an
unreadable shape coverage-class by the brief's own naming.
``ambiguous_reading`` and ``empty_curriculum`` are arguably vocabulary-
adjacent (the terms ARE taught) but were not named as coverage-class by the
brief, so this instrument does not invent that reclassification they
stay mechanism-class; see ``_CURRICULUM_COVERAGE_REASONS`` for the exact
closed set of what IS treated as coverage.
Usage before/after a lexicon or curriculum batch::
uv run python -m evals.vocab_trigger_instrument --report before.json
# ... add ratified vocabulary/curriculum content ...
uv run python -m evals.vocab_trigger_instrument --report after.json
uv run python -m evals.vocab_trigger_instrument --compare before.json --report after.json
The last form is the admissions-per-batch counter: it re-runs the instrument
against the CURRENT corpora and prints the admitted-count delta against a
prior snapshot, per composer.
"""
from __future__ import annotations
import argparse
import json
from collections import Counter
from collections.abc import Callable
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Literal, Protocol, runtime_checkable
from chat.curriculum_surface import decide_curriculum_question
from chat.deduction_surface import looks_like_deductive_argument
from evals.deduction_serve.runner import decide as decide_deduction
from generate.meaning_graph.projectors import to_deductive_logic, to_syllogism
from generate.meaning_graph.reader import Comprehension, comprehend
from generate.proof_chain.categorical import CategoricalError, decide_syllogism
from generate.proof_chain.cond_member import CondMemberArgument, read_cond_member_argument
from generate.proof_chain.english import EnglishArgument, read_english_argument
from generate.proof_chain.entail import (
INCONSISTENT_PREMISES,
OUT_OF_REGIME_OR_MALFORMED,
Entailment,
evaluate_entailment_with_trace,
)
from generate.proof_chain.exist import ExistArgument, read_exist_argument
from generate.proof_chain.member import MemberArgument, read_member_argument
from generate.proof_chain.verb import VerbArgument, read_verb_argument
#: The engine's REFUSED-only reasons (`generate/proof_chain/entail.py`'s
#: closed vocabulary also contains entailed/refuted/unknown reasons under the
#: same ``ENTAILMENT_REASONS`` set; only these two are ever attached to a
#: REFUSED verdict, which is the only outcome that reaches this classifier).
_ENGINE_REFUSED_REASONS: frozenset[str] = frozenset(
{INCONSISTENT_PREMISES, OUT_OF_REGIME_OR_MALFORMED}
)
#: Three buckets, not two: a decline is either a READER/gate refusal on an
#: unreadable SHAPE ("mechanism"), a composer refusal on an absent TERM or
#: RELATION ("coverage"), or a fully-read, well-formed argument the ROBDD
#: engine itself declines because its premises are inconsistent or the
#: structure is outside the decidable regime ("engine_refused" —
#: `generate/proof_chain/entail.py`'s own REFUSED reasons). The third bucket
#: is neither a shape gap nor a vocabulary gap — expanding a lexicon or a
#: curriculum cannot move it, so lumping it into "mechanism" would misdirect
#: exactly the trigger this instrument exists to make legible.
RefusalClass = Literal["mechanism", "coverage", "engine_refused"]
#: The curriculum composer's own closed reason vocabulary that names an
#: absent TERM or RELATION rather than an unreadable shape (module docstring).
_CURRICULUM_COVERAGE_REASONS: frozenset[str] = frozenset(
{"untaught_vocabulary", "out_of_curriculum"}
)
@runtime_checkable
class _DecidableArgument(Protocol):
"""The shape every band's ``*Argument`` success type shares
(`generate/proof_chain/{english,member,cond_member,verb,exist}.py`)
named here so the fallback loop below can call
``evaluate_entailment_with_trace`` without a per-band branch. Declared
as read-only properties, not plain attributes: every concrete
``*Argument`` is a frozen dataclass, and a mutable Protocol attribute is
incompatible with a read-only frozen field."""
@property
def premise_formulas(self) -> tuple[str, ...]: ...
@property
def query_formula(self) -> str: ...
#: The deduction fallback cascade, in the exact order
#: ``chat/deduction_surface.py`` tries them, paired with the ``Argument``
#: type each reader returns on success (vs. a typed ``*Refusal`` on decline).
_DEDUCTION_BAND_READERS: tuple[
tuple[str, Callable[[str], Any], type[_DecidableArgument]], ...
] = (
("v2_en", read_english_argument, EnglishArgument),
("v3_mem", read_member_argument, MemberArgument),
("v4_condmem", read_cond_member_argument, CondMemberArgument),
("v5_verb", read_verb_argument, VerbArgument),
("v6_exist", read_exist_argument, ExistArgument),
)
def _engine_reason(premise_formulas: tuple[str, ...], query_formula: str) -> str | None:
"""The engine's REFUSED reason for an already-admitted argument, or
``None`` when the engine reached a definite verdict (entailed/refuted/
unknown) i.e. this band did NOT decline it; something downstream in
the cascade order already claimed the case before this point is reached
in practice, but the caller still checks."""
trace = evaluate_entailment_with_trace(premise_formulas, query_formula)
return trace.reason if trace.outcome is Entailment.REFUSED else None
def refusal_class_deduction(reason: str) -> RefusalClass:
"""A reason from the engine's closed REFUSED vocabulary is engine-level
(inconsistent premises / out of regime) regardless of WHICH band parsed
the argument before handing it to the engine any of the seven bands
can admit a well-formed but inconsistent argument. Everything else is a
reader/gate SHAPE refusal; no deduction band refuses on vocabulary
(module docstring)."""
if reason in _ENGINE_REFUSED_REASONS:
return "engine_refused"
return "mechanism"
def refusal_class_curriculum(reason: str) -> RefusalClass:
"""``untaught_vocabulary`` / ``out_of_curriculum`` are coverage-class;
the engine's own REFUSED reasons (reachable here too — the curriculum
composer runs the same ROBDD engine over compiled premises) are
engine-level; every other typed curriculum reason is mechanism-class."""
if reason in _CURRICULUM_COVERAGE_REASONS:
return "coverage"
if reason in _ENGINE_REFUSED_REASONS:
return "engine_refused"
return "mechanism"
def deduction_refusal_reason(text: str) -> tuple[str, str]:
"""``(band, reason)`` this text declined at — the band whose engine
verdict or reader refusal produced the decline
:func:`evals.deduction_serve.runner.decide` already reported. Mirrors
``chat/deduction_surface.py``'s FULL cascade, including Band v1/v1b
(which the production composer's surface never exposes a reason for
either an inconsistent-premises decline there renders the same generic
wording as a reader refusal) recovering the typed reason precisely is
the reason this function exists at all."""
if not looks_like_deductive_argument(text):
return "gate", "not_argument_shaped"
comp = comprehend(text)
if isinstance(comp, Comprehension):
projected = to_deductive_logic(comp)
if projected is not None:
premises, query = projected
reason = _engine_reason(premises, query)
return "v1", reason or "" # pragma: no cover - decide() said declined
syllogism = to_syllogism(comp)
if syllogism is not None:
structure, s_query = syllogism
try:
trace = decide_syllogism(structure, s_query)
except CategoricalError as exc:
return "v1b", str(exc) or "categorical_error"
reason = trace.reason if trace.outcome is Entailment.REFUSED else ""
return "v1b", reason # pragma: no cover - decide() said declined
band, reason = "reader", ""
else:
band, reason = "reader", comp.reason
for band_name, reader, argument_type in _DEDUCTION_BAND_READERS:
arg = reader(text)
if isinstance(arg, argument_type):
engine_reason = _engine_reason(arg.premise_formulas, arg.query_formula)
return band_name, engine_reason or "" # pragma: no cover
band, reason = band_name, arg.reason
return band, reason
def curriculum_refusal_reason(text: str) -> str:
"""The typed reason ``decide_curriculum_question`` declined *text* for.
``decide_curriculum_question`` already carries ``.reason`` on its
returned ``CurriculumDecision`` for every ``declined`` verdict (gate
reasons AND the engine's own REFUSED reasons alike) — no second cascade
walk needed here, unlike the deduction side."""
return decide_curriculum_question(text).reason
@dataclass(frozen=True, slots=True)
class RefusalHistogram:
"""One composer's admitted/declined split, with declines broken out by
typed reason and by mechanism-vs-coverage class."""
n: int
admitted: int
declined: int
by_reason: dict[str, int]
by_class: dict[str, int]
def to_json_dict(self) -> dict[str, Any]:
return {
"n": self.n,
"admitted": self.admitted,
"declined": self.declined,
"by_reason": dict(sorted(self.by_reason.items())),
"by_class": dict(sorted(self.by_class.items())),
}
#: Split directories, discovered rather than duplicated from the runners'
#: own (module-private) split tuples — every ``<split>/cases.jsonl`` under
#: each lane root, sorted for determinism. ``deduction_serve/practice/`` and
#: ``curriculum_serve/oracle.py``'s directory carry no ``cases.jsonl`` and
#: are silently skipped by the glob, not filtered explicitly.
_DEDUCTION_ROOT = Path(__file__).resolve().parent / "deduction_serve"
_CURRICULUM_ROOT = Path(__file__).resolve().parent / "curriculum_serve"
def _load_jsonl(path: Path) -> list[dict]:
with path.open(encoding="utf-8") as fh:
return [json.loads(line) for line in fh if line.strip()]
def _all_deduction_texts() -> list[str]:
texts: list[str] = []
for path in sorted(_DEDUCTION_ROOT.glob("*/cases.jsonl")):
texts.extend(case["text"] for case in _load_jsonl(path))
return texts
def _all_curriculum_texts() -> list[str]:
texts: list[str] = []
for path in sorted(_CURRICULUM_ROOT.glob("*/cases.jsonl")):
texts.extend(case["text"] for case in _load_jsonl(path))
return texts
def build_deduction_histogram(texts: list[str]) -> RefusalHistogram:
by_reason: Counter[str] = Counter()
by_class: Counter[str] = Counter()
admitted = 0
for text in texts:
if decide_deduction(text) != "declined":
admitted += 1
continue
band, reason = deduction_refusal_reason(text)
by_reason[f"{band}:{reason}"] += 1
by_class[refusal_class_deduction(reason)] += 1
return RefusalHistogram(
n=len(texts),
admitted=admitted,
declined=len(texts) - admitted,
by_reason=dict(by_reason),
by_class=dict(by_class),
)
def build_curriculum_histogram(texts: list[str]) -> RefusalHistogram:
by_reason: Counter[str] = Counter()
by_class: Counter[str] = Counter()
admitted = 0
for text in texts:
if decide_curriculum_question(text).verdict != "declined":
admitted += 1
continue
reason = curriculum_refusal_reason(text)
by_reason[reason] += 1
by_class[refusal_class_curriculum(reason)] += 1
return RefusalHistogram(
n=len(texts),
admitted=admitted,
declined=len(texts) - admitted,
by_reason=dict(by_reason),
by_class=dict(by_class),
)
def build_report() -> dict[str, Any]:
"""The deterministic, pinnable snapshot over the CURRENT committed lane
corpora (`evals/deduction_serve/*/cases.jsonl`,
`evals/curriculum_serve/*/cases.jsonl`) pure over those files and the
production decision paths, so two runs against the same commit produce
byte-identical JSON."""
deduction = build_deduction_histogram(_all_deduction_texts())
curriculum = build_curriculum_histogram(_all_curriculum_texts())
combined_class: Counter[str] = Counter()
combined_class.update(deduction.by_class)
combined_class.update(curriculum.by_class)
return {
"schema_version": 1,
"instrument": "vocab_trigger",
"deduction": deduction.to_json_dict(),
"curriculum": curriculum.to_json_dict(),
"combined": {
"admitted": deduction.admitted + curriculum.admitted,
"declined": deduction.declined + curriculum.declined,
"by_class": dict(sorted(combined_class.items())),
},
}
def write_report(path: Path) -> dict[str, Any]:
report = build_report()
path.write_text(json.dumps(report, indent=2, sort_keys=True) + "\n", encoding="utf-8")
return report
def admissions_delta(before: dict[str, Any], after: dict[str, Any]) -> dict[str, Any]:
"""The admissions-per-batch counter: per-composer admitted-count delta
between two snapshots (``build_report``/``write_report`` output). Run
once before a lexicon/curriculum batch, once after, and diff this is
the falsifiable measurement Option A of the audit's Q5 asks for, applied
to whichever composer the batch targeted."""
return {
"deduction_admitted_delta": (
after["deduction"]["admitted"] - before["deduction"]["admitted"]
),
"curriculum_admitted_delta": (
after["curriculum"]["admitted"] - before["curriculum"]["admitted"]
),
"combined_admitted_delta": (
after["combined"]["admitted"] - before["combined"]["admitted"]
),
}
def _print_histogram(name: str, hist: RefusalHistogram) -> None:
print(f"[{name}] n={hist.n} admitted={hist.admitted} declined={hist.declined}")
print(f" by_class: {dict(sorted(hist.by_class.items()))}")
for reason, count in sorted(hist.by_reason.items()):
print(f" {reason}: {count}")
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--report", type=Path, default=None,
help="write the deterministic JSON snapshot to this path",
)
parser.add_argument(
"--compare", type=Path, default=None,
help="a prior snapshot (from --report) to diff the current run against",
)
args = parser.parse_args(argv)
report = build_report()
if args.report is not None:
write_report(args.report)
if args.compare is not None:
before = json.loads(args.compare.read_text(encoding="utf-8"))
delta = admissions_delta(before, report)
print(f"admissions since {args.compare}: {delta}")
return 0
_print_histogram("deduction", RefusalHistogram(**report["deduction"]))
_print_histogram("curriculum", RefusalHistogram(**report["curriculum"]))
print(f"[combined] {report['combined']}")
return 0
if __name__ == "__main__":
raise SystemExit(main())

View file

@ -0,0 +1,118 @@
"""tests/test_vocab_trigger_instrument.py — mechanism-vs-coverage refusal instrument (Tier S3).
Pins the classifier's closed rule and the instrument's arithmetic against the
CURRENT committed lane corpora, so a future band/curriculum change that shifts
the split is a visible diff here rather than a silent behavior change.
"""
from __future__ import annotations
from evals.vocab_trigger_instrument import (
_all_curriculum_texts,
_all_deduction_texts,
admissions_delta,
build_curriculum_histogram,
build_deduction_histogram,
build_report,
curriculum_refusal_reason,
deduction_refusal_reason,
refusal_class_curriculum,
refusal_class_deduction,
)
def test_deduction_refusals_are_never_coverage_class() -> None:
"""The deduction bands read a closed connective/quantifier grammar, not
open vocabulary every declined deduction case classifies as mechanism
or engine_refused, never coverage (module docstring's central claim)."""
hist = build_deduction_histogram(_all_deduction_texts())
assert hist.declined > 0
assert "coverage" not in hist.by_class
def test_curriculum_has_real_coverage_refusals() -> None:
"""The curriculum composer's untaught_vocabulary/out_of_curriculum gate
is a genuine coverage axis the committed physics split exercises it."""
hist = build_curriculum_histogram(_all_curriculum_texts())
assert hist.by_class.get("coverage", 0) > 0
def test_refusal_class_curriculum_closed_set() -> None:
assert refusal_class_curriculum("untaught_vocabulary") == "coverage"
assert refusal_class_curriculum("out_of_curriculum") == "coverage"
assert refusal_class_curriculum("question_shape_out_of_band") == "mechanism"
assert refusal_class_curriculum("ambiguous_reading") == "mechanism"
assert refusal_class_curriculum("empty_curriculum") == "mechanism"
assert refusal_class_curriculum("inconsistent_premises") == "engine_refused"
assert refusal_class_curriculum("out_of_regime_or_malformed") == "engine_refused"
def test_refusal_class_deduction_only_engine_reasons_are_engine_refused() -> None:
assert refusal_class_deduction("inconsistent_premises") == "engine_refused"
assert refusal_class_deduction("out_of_regime_or_malformed") == "engine_refused"
assert refusal_class_deduction("sentence_shape_out_of_band") == "mechanism"
assert refusal_class_deduction("quantifier_out_of_band") == "mechanism"
assert refusal_class_deduction("not_argument_shaped") == "mechanism"
def test_deduction_refusal_reason_recovers_engine_level_decline() -> None:
"""A well-formed but self-contradictory Band-v1 argument declines at the
ENGINE, not at any reader a real bug this test pins (an earlier version
of this instrument misattributed it to whichever fallback band happened
to run last)."""
band, reason = deduction_refusal_reason("p. Not p. Therefore q.")
assert band == "v1"
assert reason == "inconsistent_premises"
def test_deduction_refusal_reason_recovers_reader_level_decline() -> None:
band, reason = deduction_refusal_reason("colorless green ideas therefore sleep furiously")
assert band in {"gate", "reader", "v2_en", "v3_mem", "v4_condmem", "v5_verb", "v6_exist"}
assert reason != ""
assert refusal_class_deduction(reason) == "mechanism"
def test_curriculum_refusal_reason_matches_decide() -> None:
from chat.curriculum_surface import decide_curriculum_question
text = "Does the build pass?"
assert curriculum_refusal_reason(text) == decide_curriculum_question(text).reason
def test_histogram_totals_partition_the_corpus() -> None:
hist = build_deduction_histogram(_all_deduction_texts())
assert hist.admitted + hist.declined == hist.n
assert sum(hist.by_reason.values()) == hist.declined
assert sum(hist.by_class.values()) == hist.declined
def test_build_report_is_deterministic() -> None:
a = build_report()
b = build_report()
assert a == b
def test_admissions_delta_is_zero_against_self() -> None:
report = build_report()
delta = admissions_delta(report, report)
assert delta == {
"deduction_admitted_delta": 0,
"curriculum_admitted_delta": 0,
"combined_admitted_delta": 0,
}
def test_admissions_delta_reflects_a_new_admission() -> None:
before = build_report()
after = {
"deduction": {**before["deduction"], "admitted": before["deduction"]["admitted"] + 3},
"curriculum": before["curriculum"],
"combined": {
**before["combined"],
"admitted": before["combined"]["admitted"] + 3,
},
}
delta = admissions_delta(before, after)
assert delta["deduction_admitted_delta"] == 3
assert delta["combined_admitted_delta"] == 3
assert delta["curriculum_admitted_delta"] == 0