core/tests/test_determination_estimation_lane.py
Claude 71bf04fb44
feat(provenance,teaching): close the lateral gaps the assessment actually found
Squashes the arc's work into one commit; the workflow-file edit it originally
carried is excluded (see the end of this message).

## Lane 1 — Workbench recorded a proved answer as ungrounded

With deduction_serving_enabled ratified ON (ADR-0256), workbench/api.py's live
chat route builds a bare ChatRuntime(), so the deduction composer decides
Workbench turns and stamps grounding_source="deduction" — but
_coerce_grounding_source carried a hand-copied whitelist of the six pre-arc
labels and silently rewrote anything else to "none". The runtime comment
reasoned this was inert because "REPL turns do not flow through Workbench's
CognitivePipelineRecord path". True, and irrelevant: the traffic flows the
other way. Stale since 2026-07-24.

Scope is one field. workbench/api.py:818 prefers TurnEvent.epistemic_state,
which read epistemic_state_needed — honest. So the UNregistered path degraded
honestly while the hand-copied whitelist asserted a falsehood; a second copy of
a closed enum was worse than no copy. Hence registration AND derivation:
GROUNDING_SOURCES exposes the Literal's members, and the coercion reads it.
workbench-ui badges/tokens/snapshot follow; enumCoverage.test.ts forces atomicity.

## Lane 2 — the ratification ceremony

The discovery loop was instrumented but not closed. teaching/ratification.py
turns a reviewed decision into a chain record, a corpus commit, and a receipt.

Its design turns on one observation: _ratified_rows DROPS unadmissible rows
silently — correct when serving, a trap when ratifying, because the file grows,
the commit lands, and the band count does not move. So the ceremony refuses to
call an append a ratification until it has re-read the curriculum through the
real loader and seen the chain arrive; a non-admitted append is rolled back.
Validation is a pre-flight courtesy, admission is the proof.

Arena queue entry and ledger reseal are deliberately NOT performed (bridge rule
1); the receipt names them. Front door: `core proposal-queue ratify`, a sibling
of `review` rather than a flag on it.

## Lane 3 — structural closures

- ADR-0263 gains rule 5: absence policy is DECLARED in CAPABILITY_LEDGERS, not
  passed at the call site. An AST-matched test fails if a serving path passes
  missing_ok again.
- Deductive suite added WHOLE to the pre-push gate: 285 tests in 29s against
  smoke's 216 in 62s, so no coverage trade was needed.
- Smoke/CI parity assertion made bidirectional. It was one-directional, and had
  drifted.
- test_prior_surface_deduction_binding.py pins correction binding on the
  deduction path. The review's diagnosis did NOT reproduce — hash_surface moves
  in lockstep — so it pins what is there. Mutation-checked.
- Domain-keyed ADR index over 312 flat-numbered files, explicitly partial.
- Arc-close brief template, plus this arc's own brief filled in against it.

## Lanes 4 and 5 — two premises falsified by measurement, one of them mine

Math 4.2: baseline reproduced (correct=5 wrong=0 refused=495); all four named
cases traced to one seam with each gap isolated by one-variable probes. Then the
number that changes the recommendation: the gap blocking case 0000 affects 1
case in 500, the 'than' gap blocking 0001 affects 2. ADR-0251's prohibition on
per-case growth now rests on a count. No reader change made.

CGA: versor_condition is 0.22% of a turn, not the "~10x proof latency" I claimed
— that multiplied an isolated microbenchmark by a call count and compared it to
a single verdict's latency. The real cost is geometric_product at 33,986
calls/turn (~73%) via cga_inner in search paths. The obvious closed form is NOT
bit-exact (954/4000 in f32); backend.vault_recall's serial fold IS (3000/3000,
worst-rel 0) and is the correct target. cargo test could not run —
static.crates.io is denied by the sandbox network policy — so the Rust parity
question stays open and the typestate lane is carried forward, not shipped
uncompiled.

## Not landed: three lines owed to .github/workflows/smoke.yml

The CI smoke gate is narrower than the local one —
test_pack_draft_serve_boundary.py (ADR-0253 INV-33) has been local-only, unseen
because the parity pin checked one direction. The edit was authored and rejected
at push for lacking the `workflow` OAuth scope, so it is recorded as a named,
dated PENDING_IN_CI exception rather than dropped: the assertion still fires on
any new divergence, and a second guard fires once the three land.

[Verification]: pre-push gates all green — smoke 236 passed, warmed_session 10
passed, deductive 285 passed. Ratification 14, ADR index 5, CLI suites 10.
Grounding/epistemic sweep 741 passed 1 skipped. workbench-ui 598 passed across
73 files, tsc -b clean. capability index 11 passed, digest unchanged. Math
holdout correct=5 wrong=0 refused=495. Committed chain corpora byte-unchanged
after the tests that write to them.
Environment caveat: the repo pins requires-python ==3.12.13, which uv cannot
fetch for linux-x86_64, so `uv sync --locked` fails. All Python runs used a
scratch venv on 3.12.11 with declared deps — not the locked universe, not the
full ~12k suite. The pin was left untouched. Re-run on a 3.12.13 host before
treating this as merge evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FduW6Krm3PPQv3P5iwBYtx
2026-07-25 04:51:15 +00:00

216 lines
8.7 KiB
Python

"""Step E — the converse-guess estimator + its calibration gold lane.
The estimator is BLIND (never reads symmetry metadata); the reliability gate decides
licensing from MEASURED commitment precision. The load-bearing properties: the gate
DISCRIMINATES (a symmetric predicate's converse-guess earns SERVE; a directed one does
not), the SERVE license is earned by VOLUME (the Wilson floor binds at 657), and the
serving-side estimator fires only on a told converse.
"""
from __future__ import annotations
from dataclasses import replace as _dc_replace
from pathlib import Path
import pytest
from chat.runtime import ChatRuntime
from core.config import DEFAULT_CONFIG
from core.reliability_gate import N_MIN
from evals.determination_estimation import (
LICENSED_PREDICATE,
REFUSED_PREDICATE,
build_ledger,
load_symmetric_predicates,
reliability_at,
run,
)
from generate.determine.estimate import ConverseEstimate, converse_class_name, estimate_converse
from generate.meaning_graph.relational import comprehend_relational, load_relational_pack_lemmas
from generate.realize import realize_comprehension
from session.context import SessionContext
_HIGH = 10**9
@pytest.fixture(scope="module")
def vocab_persona():
rt = ChatRuntime(no_load_state=True)
return rt._context.vocab, rt._context.persona
@pytest.fixture(scope="module")
def rel_lemmas():
return load_relational_pack_lemmas()
def _ctx(vocab_persona) -> SessionContext:
vocab, persona = vocab_persona
return SessionContext(vocab=vocab, persona=persona, vault_reproject_interval=_HIGH)
def _tell_relational(text: str, ctx: SessionContext, lemmas) -> None:
realize_comprehension(comprehend_relational(text, lemmas), ctx)
# --------------------------------------------------------------------------- #
# The gate discriminates (the whole point of E)
# --------------------------------------------------------------------------- #
def test_gate_discriminates_symmetric_from_directed() -> None:
report = run()
assert report["gate_discriminates"] is True
licensed = report["classes"][converse_class_name(LICENSED_PREDICATE)]
refused = report["classes"][converse_class_name(REFUSED_PREDICATE)]
assert licensed["serve_licensed"] is True
assert refused["serve_licensed"] is False
# The symmetric class is right every time; the directed class is wrong every time.
assert licensed["tally"]["wrong"] == 0 and licensed["tally"]["correct"] > 0
assert refused["tally"]["correct"] == 0 and refused["tally"]["wrong"] > 0
def test_serve_license_is_earned_by_volume() -> None:
# Below the Wilson volume floor a PERFECT symmetric record is still NOT SERVE-licensed.
assert reliability_at(LICENSED_PREDICATE, 656) < 0.99
assert reliability_at(LICENSED_PREDICATE, 657) >= 0.99
# And below N_MIN no reliability is claimed at all.
assert reliability_at(LICENSED_PREDICATE, N_MIN - 1) == 0.0
def test_run_is_deterministic() -> None:
a, b = run(), run()
assert a == b
def test_gold_symmetry_matches_pack() -> None:
sym = load_symmetric_predicates()
assert LICENSED_PREDICATE in sym # sibling_of — graph.edge.symmetric
assert REFUSED_PREDICATE not in sym # parent_of — graph.edge.directed
# --------------------------------------------------------------------------- #
# The serving-side estimator
# --------------------------------------------------------------------------- #
def test_estimate_fires_only_on_a_told_converse(vocab_persona, rel_lemmas) -> None:
ctx = _ctx(vocab_persona)
_tell_relational("Alice is the sibling of Bob.", ctx, rel_lemmas)
# Told sibling_of(alice, bob); the converse query sibling_of(bob, alice) gets a guess.
est = estimate_converse(ctx, "sibling_of", "bob", "alice")
assert isinstance(est, ConverseEstimate)
assert est.answer is True
assert est.basis == "estimate_converse"
assert est.subject == "bob" and est.object == "alice"
assert est.told_structure_key # ties the guess to the realized fact
# No told converse → no guess (the estimator never invents evidence).
assert estimate_converse(ctx, "sibling_of", "carol", "dave") is None
def test_estimate_is_blind_to_symmetry(vocab_persona, rel_lemmas) -> None:
# The estimator commits the converse for a DIRECTED predicate too — being wrong
# there is exactly what the gate measures and refuses to license.
ctx = _ctx(vocab_persona)
_tell_relational("Alice is the parent of Bob.", ctx, rel_lemmas)
est = estimate_converse(ctx, "parent_of", "bob", "alice")
assert isinstance(est, ConverseEstimate) and est.answer is True
# --------------------------------------------------------------------------- #
# E-2 — the ratified ledger artifact + serving-side license
# --------------------------------------------------------------------------- #
def test_ratified_ledger_matches_sealed_practice() -> None:
# Provenance (the GSM8K-style --check): the committed artifact IS the deterministic
# sealed-practice output, not a hand-edited ledger.
from generate.determine.estimation_license import load_ratified_ledger
committed = load_ratified_ledger()
fresh = build_ledger()
assert {k: (t.correct, t.wrong, t.refused) for k, t in committed.items()} == {
k: (t.correct, t.wrong, t.refused) for k, t in fresh.items()
}
def test_serve_license_from_ratified_ledger() -> None:
from generate.determine.estimation_license import serve_license
assert serve_license(LICENSED_PREDICATE).licensed is True # sibling_of — earned SERVE
assert serve_license(REFUSED_PREDICATE).licensed is False # parent_of — never
assert serve_license("nonexistent_predicate") is None # no committed evidence → refuse
def test_tampered_ledger_is_rejected(tmp_path, monkeypatch) -> None:
# A hand-edited ledger (counts changed without the matching hash) must be REJECTED,
# never silently trusted — only the sealed-practice output is admissible.
import json as _json
import generate.determine.estimation_license as mod
from generate.determine.estimation_license import RatifiedLedgerError, load_ratified_ledger
from dataclasses import replace
from core.ratified_ledger import CAPABILITY_LEDGERS
good = _json.loads(mod._LEDGER_PATH.read_text(encoding="utf-8"))
good["classes"][converse_class_name(REFUSED_PREDICATE)]["correct"] = 999 # forge a SERVE
forged_path = tmp_path / "estimation_ledger.json"
forged_path.write_text(_json.dumps(good), encoding="utf-8")
load_ratified_ledger.cache_clear()
# Redirect at the manifest, which owns the path since ADR-0263 rule 5 —
# patching the module constant would no longer reach the load.
monkeypatch.setitem(
CAPABILITY_LEDGERS,
mod._LEDGER_CAPABILITY,
replace(CAPABILITY_LEDGERS[mod._LEDGER_CAPABILITY], path=forged_path),
)
try:
with pytest.raises(RatifiedLedgerError, match="content_sha256 mismatch"):
load_ratified_ledger()
finally:
load_ratified_ledger.cache_clear() # don't poison other tests' cached load
# --------------------------------------------------------------------------- #
# E-3 — the runtime wire (chat turn → disclosed estimate, license-gated)
# --------------------------------------------------------------------------- #
def _estimation_runtime(tmp_path: Path, *, enabled: bool = True) -> ChatRuntime:
cfg = _dc_replace(
DEFAULT_CONFIG,
estimation_enabled=enabled,
accrue_realized_knowledge=True,
persist_session_state=True,
)
return ChatRuntime(config=cfg, engine_state_path=tmp_path)
def test_licensed_converse_is_served_disclosed_approximate(tmp_path) -> None:
rt = _estimation_runtime(tmp_path)
rt.chat("Alice is the sibling of Bob.") # told sibling_of(alice, bob)
resp = rt.chat("Is Bob the sibling of Alice?") # converse — DETERMINE refuses
assert resp.reach_level == "strict"
assert not resp.surface.startswith("[approximate]")
assert "bob" in resp.surface and "alice" in resp.surface
def test_unlicensed_converse_stays_strict_refusal(tmp_path) -> None:
rt = _estimation_runtime(tmp_path)
rt.chat("Alice is the parent of Bob.") # told parent_of(alice, bob) — DIRECTED
resp = rt.chat("Is Bob the parent of Alice?")
# parent_of's converse-guess is not SERVE-licensed → no widening, honest refusal.
assert resp.reach_level == "strict"
assert not resp.surface.startswith("[approximate]")
def test_estimation_flag_off_is_strict(tmp_path) -> None:
rt = _estimation_runtime(tmp_path, enabled=False)
rt.chat("Alice is the sibling of Bob.")
resp = rt.chat("Is Bob the sibling of Alice?")
assert resp.reach_level == "strict"
assert not resp.surface.startswith("[approximate]")