Phase 2 of the generalization arc, implementing the plan's §4 curriculum-entailment gold contract. "Does force cause acceleration?" is answered from the ratified physics chain corpus and nothing else; "Does gravity cause acceleration?" is declined because gravity is in no pack CORE has been taught; "Does force cause motion?" is unsettled even though the curriculum contains both links, because nothing ratified says causation composes. Path (flag-gated, default off): closed question grammar -> subject routing by ratified vocabulary -> family-scoped premise compilation from reviewed, pack-resident chains -> the SAME argument bands (ADR-0260/0261) -> the ROBDD engine. Zero subject-specific decision code: physics differs from philosophy only in which rows load. Epistemology enforced mechanically, not by convention: - gold is a function of (curriculum, question); cases pin chain ids and the runner FAILS a case whose pinned chain is absent or unratified; - untaught => UNKNOWN, never "no" (open-world; silence is not denial); - an independent oracle (own loader, ratification predicate, family table, agreement rules, verdict rule) re-derives every gold — it disagreed once, on "entropy reveals energy" vs "entropy causes energy", and the ORACLE was the side that was wrong; - anti-recall probes are a lane GUARD: a split without >=3 true-but-untaught probes refuses to run. Findings recorded rather than worked around (ADR-0262 §5): - every curriculum band is UNEARNED and every answer is hedged. A band needs n>=657 with a real outcome mix; physics teaches 7 causal + 9 modal relations, so at most 16 questions in the subject can ever be ENTAILED. A balanced band needs ~219 taught relations per subject×family — a ~25x gap that only ratified curriculum content closes. Phase 2's blocker is curriculum volume, not machinery. - REFUTED is unreachable from present corpora (every chain is positive). - there is NO biology domain-chain corpus; the biology OOD lane measures fluency, not truth. The four subjects with ratified chains and mounted packs are physics, mathematics_logic, systems_software, philosophy_theology — the composer serves all four. [Verification]: curriculum lane 32/32 wrong=0 with 5 anti-recall probes and all three contract guards passing; tests/test_curriculum_serve.py 20 passed; core test --suite deductive 252 passed; lane pinned as curriculum_serve_v1.
40 lines
768 B
JSON
40 lines
768 B
JSON
{
|
|
"aggregate": {
|
|
"correct": 32,
|
|
"declined": 0,
|
|
"n": 32,
|
|
"wrong": 0
|
|
},
|
|
"all_correct": true,
|
|
"arc": "generalization",
|
|
"lane": "curriculum_serve",
|
|
"schema_version": 1,
|
|
"splits": {
|
|
"physics": {
|
|
"all_cases_correct": true,
|
|
"anti_recall_probes": 5,
|
|
"by_band": {
|
|
"curriculum_physics_causal": 14,
|
|
"curriculum_physics_modal": 12
|
|
},
|
|
"by_gold": {
|
|
"declined": 6,
|
|
"entailed": 14,
|
|
"unknown": 12
|
|
},
|
|
"correct_by_gold": {
|
|
"declined": 6,
|
|
"entailed": 14,
|
|
"unknown": 12
|
|
},
|
|
"counts": {
|
|
"correct": 32,
|
|
"declined": 0,
|
|
"wrong": 0
|
|
},
|
|
"domain": "physics",
|
|
"n": 32
|
|
}
|
|
},
|
|
"wrong_is_zero": true
|
|
}
|