Squashes the arc's work into one commit; the workflow-file edit it originally carried is excluded (see the end of this message). ## Lane 1 — Workbench recorded a proved answer as ungrounded With deduction_serving_enabled ratified ON (ADR-0256), workbench/api.py's live chat route builds a bare ChatRuntime(), so the deduction composer decides Workbench turns and stamps grounding_source="deduction" — but _coerce_grounding_source carried a hand-copied whitelist of the six pre-arc labels and silently rewrote anything else to "none". The runtime comment reasoned this was inert because "REPL turns do not flow through Workbench's CognitivePipelineRecord path". True, and irrelevant: the traffic flows the other way. Stale since 2026-07-24. Scope is one field. workbench/api.py:818 prefers TurnEvent.epistemic_state, which read epistemic_state_needed — honest. So the UNregistered path degraded honestly while the hand-copied whitelist asserted a falsehood; a second copy of a closed enum was worse than no copy. Hence registration AND derivation: GROUNDING_SOURCES exposes the Literal's members, and the coercion reads it. workbench-ui badges/tokens/snapshot follow; enumCoverage.test.ts forces atomicity. ## Lane 2 — the ratification ceremony The discovery loop was instrumented but not closed. teaching/ratification.py turns a reviewed decision into a chain record, a corpus commit, and a receipt. Its design turns on one observation: _ratified_rows DROPS unadmissible rows silently — correct when serving, a trap when ratifying, because the file grows, the commit lands, and the band count does not move. So the ceremony refuses to call an append a ratification until it has re-read the curriculum through the real loader and seen the chain arrive; a non-admitted append is rolled back. Validation is a pre-flight courtesy, admission is the proof. Arena queue entry and ledger reseal are deliberately NOT performed (bridge rule 1); the receipt names them. Front door: `core proposal-queue ratify`, a sibling of `review` rather than a flag on it. ## Lane 3 — structural closures - ADR-0263 gains rule 5: absence policy is DECLARED in CAPABILITY_LEDGERS, not passed at the call site. An AST-matched test fails if a serving path passes missing_ok again. - Deductive suite added WHOLE to the pre-push gate: 285 tests in 29s against smoke's 216 in 62s, so no coverage trade was needed. - Smoke/CI parity assertion made bidirectional. It was one-directional, and had drifted. - test_prior_surface_deduction_binding.py pins correction binding on the deduction path. The review's diagnosis did NOT reproduce — hash_surface moves in lockstep — so it pins what is there. Mutation-checked. - Domain-keyed ADR index over 312 flat-numbered files, explicitly partial. - Arc-close brief template, plus this arc's own brief filled in against it. ## Lanes 4 and 5 — two premises falsified by measurement, one of them mine Math 4.2: baseline reproduced (correct=5 wrong=0 refused=495); all four named cases traced to one seam with each gap isolated by one-variable probes. Then the number that changes the recommendation: the gap blocking case 0000 affects 1 case in 500, the 'than' gap blocking 0001 affects 2. ADR-0251's prohibition on per-case growth now rests on a count. No reader change made. CGA: versor_condition is 0.22% of a turn, not the "~10x proof latency" I claimed — that multiplied an isolated microbenchmark by a call count and compared it to a single verdict's latency. The real cost is geometric_product at 33,986 calls/turn (~73%) via cga_inner in search paths. The obvious closed form is NOT bit-exact (954/4000 in f32); backend.vault_recall's serial fold IS (3000/3000, worst-rel 0) and is the correct target. cargo test could not run — static.crates.io is denied by the sandbox network policy — so the Rust parity question stays open and the typestate lane is carried forward, not shipped uncompiled. ## Not landed: three lines owed to .github/workflows/smoke.yml The CI smoke gate is narrower than the local one — test_pack_draft_serve_boundary.py (ADR-0253 INV-33) has been local-only, unseen because the parity pin checked one direction. The edit was authored and rejected at push for lacking the `workflow` OAuth scope, so it is recorded as a named, dated PENDING_IN_CI exception rather than dropped: the assertion still fires on any new divergence, and a second guard fires once the three land. [Verification]: pre-push gates all green — smoke 236 passed, warmed_session 10 passed, deductive 285 passed. Ratification 14, ADR index 5, CLI suites 10. Grounding/epistemic sweep 741 passed 1 skipped. workbench-ui 598 passed across 73 files, tsc -b clean. capability index 11 passed, digest unchanged. Math holdout correct=5 wrong=0 refused=495. Committed chain corpora byte-unchanged after the tests that write to them. Environment caveat: the repo pins requires-python ==3.12.13, which uv cannot fetch for linux-x86_64, so `uv sync --locked` fails. All Python runs used a scratch venv on 3.12.11 with declared deps — not the locked universe, not the full ~12k suite. The pin was left untouched. Re-run on a 3.12.13 host before treating this as merge evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FduW6Krm3PPQv3P5iwBYtx
216 lines
8.7 KiB
Python
216 lines
8.7 KiB
Python
"""Step E — the converse-guess estimator + its calibration gold lane.
|
|
|
|
The estimator is BLIND (never reads symmetry metadata); the reliability gate decides
|
|
licensing from MEASURED commitment precision. The load-bearing properties: the gate
|
|
DISCRIMINATES (a symmetric predicate's converse-guess earns SERVE; a directed one does
|
|
not), the SERVE license is earned by VOLUME (the Wilson floor binds at 657), and the
|
|
serving-side estimator fires only on a told converse.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from dataclasses import replace as _dc_replace
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
from chat.runtime import ChatRuntime
|
|
from core.config import DEFAULT_CONFIG
|
|
from core.reliability_gate import N_MIN
|
|
from evals.determination_estimation import (
|
|
LICENSED_PREDICATE,
|
|
REFUSED_PREDICATE,
|
|
build_ledger,
|
|
load_symmetric_predicates,
|
|
reliability_at,
|
|
run,
|
|
)
|
|
from generate.determine.estimate import ConverseEstimate, converse_class_name, estimate_converse
|
|
from generate.meaning_graph.relational import comprehend_relational, load_relational_pack_lemmas
|
|
from generate.realize import realize_comprehension
|
|
from session.context import SessionContext
|
|
|
|
_HIGH = 10**9
|
|
|
|
|
|
@pytest.fixture(scope="module")
|
|
def vocab_persona():
|
|
rt = ChatRuntime(no_load_state=True)
|
|
return rt._context.vocab, rt._context.persona
|
|
|
|
|
|
@pytest.fixture(scope="module")
|
|
def rel_lemmas():
|
|
return load_relational_pack_lemmas()
|
|
|
|
|
|
def _ctx(vocab_persona) -> SessionContext:
|
|
vocab, persona = vocab_persona
|
|
return SessionContext(vocab=vocab, persona=persona, vault_reproject_interval=_HIGH)
|
|
|
|
|
|
def _tell_relational(text: str, ctx: SessionContext, lemmas) -> None:
|
|
realize_comprehension(comprehend_relational(text, lemmas), ctx)
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# The gate discriminates (the whole point of E)
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_gate_discriminates_symmetric_from_directed() -> None:
|
|
report = run()
|
|
assert report["gate_discriminates"] is True
|
|
licensed = report["classes"][converse_class_name(LICENSED_PREDICATE)]
|
|
refused = report["classes"][converse_class_name(REFUSED_PREDICATE)]
|
|
assert licensed["serve_licensed"] is True
|
|
assert refused["serve_licensed"] is False
|
|
# The symmetric class is right every time; the directed class is wrong every time.
|
|
assert licensed["tally"]["wrong"] == 0 and licensed["tally"]["correct"] > 0
|
|
assert refused["tally"]["correct"] == 0 and refused["tally"]["wrong"] > 0
|
|
|
|
|
|
def test_serve_license_is_earned_by_volume() -> None:
|
|
# Below the Wilson volume floor a PERFECT symmetric record is still NOT SERVE-licensed.
|
|
assert reliability_at(LICENSED_PREDICATE, 656) < 0.99
|
|
assert reliability_at(LICENSED_PREDICATE, 657) >= 0.99
|
|
# And below N_MIN no reliability is claimed at all.
|
|
assert reliability_at(LICENSED_PREDICATE, N_MIN - 1) == 0.0
|
|
|
|
|
|
def test_run_is_deterministic() -> None:
|
|
a, b = run(), run()
|
|
assert a == b
|
|
|
|
|
|
def test_gold_symmetry_matches_pack() -> None:
|
|
sym = load_symmetric_predicates()
|
|
assert LICENSED_PREDICATE in sym # sibling_of — graph.edge.symmetric
|
|
assert REFUSED_PREDICATE not in sym # parent_of — graph.edge.directed
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# The serving-side estimator
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_estimate_fires_only_on_a_told_converse(vocab_persona, rel_lemmas) -> None:
|
|
ctx = _ctx(vocab_persona)
|
|
_tell_relational("Alice is the sibling of Bob.", ctx, rel_lemmas)
|
|
# Told sibling_of(alice, bob); the converse query sibling_of(bob, alice) gets a guess.
|
|
est = estimate_converse(ctx, "sibling_of", "bob", "alice")
|
|
assert isinstance(est, ConverseEstimate)
|
|
assert est.answer is True
|
|
assert est.basis == "estimate_converse"
|
|
assert est.subject == "bob" and est.object == "alice"
|
|
assert est.told_structure_key # ties the guess to the realized fact
|
|
|
|
# No told converse → no guess (the estimator never invents evidence).
|
|
assert estimate_converse(ctx, "sibling_of", "carol", "dave") is None
|
|
|
|
|
|
def test_estimate_is_blind_to_symmetry(vocab_persona, rel_lemmas) -> None:
|
|
# The estimator commits the converse for a DIRECTED predicate too — being wrong
|
|
# there is exactly what the gate measures and refuses to license.
|
|
ctx = _ctx(vocab_persona)
|
|
_tell_relational("Alice is the parent of Bob.", ctx, rel_lemmas)
|
|
est = estimate_converse(ctx, "parent_of", "bob", "alice")
|
|
assert isinstance(est, ConverseEstimate) and est.answer is True
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# E-2 — the ratified ledger artifact + serving-side license
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_ratified_ledger_matches_sealed_practice() -> None:
|
|
# Provenance (the GSM8K-style --check): the committed artifact IS the deterministic
|
|
# sealed-practice output, not a hand-edited ledger.
|
|
from generate.determine.estimation_license import load_ratified_ledger
|
|
|
|
committed = load_ratified_ledger()
|
|
fresh = build_ledger()
|
|
assert {k: (t.correct, t.wrong, t.refused) for k, t in committed.items()} == {
|
|
k: (t.correct, t.wrong, t.refused) for k, t in fresh.items()
|
|
}
|
|
|
|
|
|
def test_serve_license_from_ratified_ledger() -> None:
|
|
from generate.determine.estimation_license import serve_license
|
|
|
|
assert serve_license(LICENSED_PREDICATE).licensed is True # sibling_of — earned SERVE
|
|
assert serve_license(REFUSED_PREDICATE).licensed is False # parent_of — never
|
|
assert serve_license("nonexistent_predicate") is None # no committed evidence → refuse
|
|
|
|
|
|
def test_tampered_ledger_is_rejected(tmp_path, monkeypatch) -> None:
|
|
# A hand-edited ledger (counts changed without the matching hash) must be REJECTED,
|
|
# never silently trusted — only the sealed-practice output is admissible.
|
|
import json as _json
|
|
|
|
import generate.determine.estimation_license as mod
|
|
from generate.determine.estimation_license import RatifiedLedgerError, load_ratified_ledger
|
|
|
|
from dataclasses import replace
|
|
|
|
from core.ratified_ledger import CAPABILITY_LEDGERS
|
|
|
|
good = _json.loads(mod._LEDGER_PATH.read_text(encoding="utf-8"))
|
|
good["classes"][converse_class_name(REFUSED_PREDICATE)]["correct"] = 999 # forge a SERVE
|
|
forged_path = tmp_path / "estimation_ledger.json"
|
|
forged_path.write_text(_json.dumps(good), encoding="utf-8")
|
|
|
|
load_ratified_ledger.cache_clear()
|
|
# Redirect at the manifest, which owns the path since ADR-0263 rule 5 —
|
|
# patching the module constant would no longer reach the load.
|
|
monkeypatch.setitem(
|
|
CAPABILITY_LEDGERS,
|
|
mod._LEDGER_CAPABILITY,
|
|
replace(CAPABILITY_LEDGERS[mod._LEDGER_CAPABILITY], path=forged_path),
|
|
)
|
|
try:
|
|
with pytest.raises(RatifiedLedgerError, match="content_sha256 mismatch"):
|
|
load_ratified_ledger()
|
|
finally:
|
|
load_ratified_ledger.cache_clear() # don't poison other tests' cached load
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# E-3 — the runtime wire (chat turn → disclosed estimate, license-gated)
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def _estimation_runtime(tmp_path: Path, *, enabled: bool = True) -> ChatRuntime:
|
|
cfg = _dc_replace(
|
|
DEFAULT_CONFIG,
|
|
estimation_enabled=enabled,
|
|
accrue_realized_knowledge=True,
|
|
persist_session_state=True,
|
|
)
|
|
return ChatRuntime(config=cfg, engine_state_path=tmp_path)
|
|
|
|
|
|
def test_licensed_converse_is_served_disclosed_approximate(tmp_path) -> None:
|
|
rt = _estimation_runtime(tmp_path)
|
|
rt.chat("Alice is the sibling of Bob.") # told sibling_of(alice, bob)
|
|
resp = rt.chat("Is Bob the sibling of Alice?") # converse — DETERMINE refuses
|
|
assert resp.reach_level == "strict"
|
|
assert not resp.surface.startswith("[approximate]")
|
|
assert "bob" in resp.surface and "alice" in resp.surface
|
|
|
|
|
|
def test_unlicensed_converse_stays_strict_refusal(tmp_path) -> None:
|
|
rt = _estimation_runtime(tmp_path)
|
|
rt.chat("Alice is the parent of Bob.") # told parent_of(alice, bob) — DIRECTED
|
|
resp = rt.chat("Is Bob the parent of Alice?")
|
|
# parent_of's converse-guess is not SERVE-licensed → no widening, honest refusal.
|
|
assert resp.reach_level == "strict"
|
|
assert not resp.surface.startswith("[approximate]")
|
|
|
|
|
|
def test_estimation_flag_off_is_strict(tmp_path) -> None:
|
|
rt = _estimation_runtime(tmp_path, enabled=False)
|
|
rt.chat("Alice is the sibling of Bob.")
|
|
resp = rt.chat("Is Bob the sibling of Alice?")
|
|
assert resp.reach_level == "strict"
|
|
assert not resp.surface.startswith("[approximate]")
|