The curriculum composer committed on question SHAPE — any `Does …?` text — and then, being fail-closed, always returned a surface. With the flag on, "Does the build pass?" would have been taken away from the rest of dispatch and answered "I haven't been taught the or pass"; "Does anyone know the time?" likewise. Found by checking the flag's readiness before recommending ratification; never live, since the flag is default-off. The asymmetry is the lesson: the deduction composer may commit on shape because a sentence-initial "therefore" IS a signal of intent — text shaped that way is an argument. `Does …?` is one of the most common ways to open any English question and signals nothing. A fail-closed composer is only as safe as its commit gate is narrow. The gate is now `is_curriculum_question`: claim the turn only when the question parses to three tokens AND its terms are vocabulary a served subject actually teaches. Everything past the gate stays fail-closed — including `ambiguous_reading` (both terms taught, two subjects claim them) and `out_of_curriculum` (terms taught, relation unknown), which are real curriculum questions with honest answers. The DECIDER is unchanged, so the lane still records untaught-vocabulary probes as declined: the curriculum path declines them AND does not speak for them. [Verification]: tests/test_curriculum_serve.py 28 passed (+8: five ordinary `Does …?` questions pass through untouched, three routable ones still answered); core test --suite deductive 268 passed; curriculum lane 32/32 wrong=0 with its pinned report SHA unchanged.
4.6 KiB
Curriculum-serve lane contract
What this lane scores
The production curriculum-grounded answering path — the exact pipeline
chat/curriculum_surface.py::decide_curriculum_question runs on a core chat
turn when curriculum_serving_enabled is on:
Does <term> <relation> <term>? (closed exam-question grammar)
→ subject routing (the domain whose ratified vocabulary holds BOTH terms)
→ relation family (the connective, normalized by ADR-0260 agreement)
→ premise compilation (that family's RATIFIED chains, and nothing else)
→ the argument bands (ADR-0260 verb reading → ADR-0261 existential)
→ the ROBDD engine (ADR-0201/0218)
There is no subject-specific decision code anywhere in the path. Physics differs from philosophy only in which curriculum rows load — which is the property that makes "add a subject" a data operation.
Gold vocabulary
Three classes are reachable: entailed, unknown, declined.
refuted is unreachable from a purely positive curriculum and that is not
an omission. The curriculum is read OPEN-world: a relation it does not state
is UNKNOWN, never "no". Refuting would require the curriculum to teach a
negative ("no X causes Y"), which no ratified corpus does yet. See ADR-0262 §5.
The three lane guards (all run before any case is scored)
Failing a guard is a lane failure, not a case miss — the lane is unsound, not merely under-covered.
- Provenance (plan §4.3) — every chain id a case pins must resolve in the subject's ratified curriculum. A case whose curriculum moved under it breaks loudly rather than quietly answering from what remains.
- Corpus soundness (§4.4) — the INDEPENDENT oracle
(
evals/curriculum_serve/oracle.py, sharing no code with the serving path: own loader, own ratification predicate, own family table, own agreement normalization, own verdict rule) must re-derive every committed gold. - Anti-recall coverage (§4.7) — the split must carry ≥3 probes whose answer is true in the world and absent from the curriculum, and no probe may carry a committed gold. Without them the lane cannot show the system decodes rather than recalls, and it must not ship.
Splits
physics/— 32 hand-authored cases overphysics_chains_v1: 13 taught edges (both connectives of each family, incl. the row whose declaredoperator_familydisagrees with its connective), 5 untaught compositions at depth 2–3, 4 reverse-direction and cross-relation near-misses, 5 anti-recall probes (3 with untaught vocabulary — gravity, current, pressure — and 2 with taught vocabulary but untaught relations), and 5 typed refusals.
wrong=0 discipline
Identical to the deduction-serve lane: wrong (a committed verdict that
disagrees with gold) MUST stay 0; a decline where gold expected a verdict is a
coverage miss, tracked in counts.declined, never conflated with a
confabulation. The runner requires correct == n.
Lane scope vs composer scope
The lane scores the DECIDER (decide_curriculum_question), which records a
verdict for every case including untaught_vocabulary ones. The COMPOSER
(curriculum_grounded_surface) is narrower: it claims a turn only when the
question is routable (ADR-0262 §5.3), so the anti-recall probes reach the rest
of dispatch rather than being answered with a curriculum remark. Both
statements hold at once — the curriculum path declines them, and it does not
speak for them.
What the lane deliberately does NOT do
- It does not compose chains.
force causes accelerationandacceleration causes motionare both taught;force causes motionis UNKNOWN. Causal transitivity is a substantive claim about the world, and no ratified corpus teaches it. The oracle reports the shortest path length alongside its verdict precisely so the lane can assert that a reachable pair is still answered UNKNOWN — composition is provably not happening. - It does not read the corpus's
operator_familyfield. The family comes from the CONNECTIVE, because a question carries a relation word and nothing else; deriving the family from a field the question cannot carry would let premise compilation and question routing disagree, and a taught edge could go missing from the premises compiled to decide it.
Reproduce
uv run python -m evals.curriculum_serve.runner # human-facing
uv run python -m evals.curriculum_serve.runner --report evals/curriculum_serve/report.json
Pinned in scripts/verify_lane_shas.py as lane id curriculum_serve_v1.
core test --suite deductive runs tests/test_curriculum_serve.py.