diff --git a/docs/handoff/ADR-0246-Acceptance-Evidence.md b/docs/handoff/ADR-0246-Acceptance-Evidence.md new file mode 100644 index 00000000..19af8f15 --- /dev/null +++ b/docs/handoff/ADR-0246-Acceptance-Evidence.md @@ -0,0 +1,89 @@ +# ADR-0246 / 0247 / 0248 — Acceptance Evidence (intelligence-loop arc) + +**Status**: Evidence for Shay's §8 rulings — NOT an acceptance; no flags flipped, no ADR status changed. +**Date**: 2026-07-18 (branch `feat/intelligence-loop-arc`) +**Companion**: `docs/audit/adr-0246-acceptance-packet-2026-07-17.md` (the packet holding the +pending rulings), `docs/research/intelligence-loop-homestretch-plan-2026-07-18.md` (the arc plan), +`docs/research/spark-audit-adjudication-2026-07-18.md` (the evidence firewall). +**Reproduce**: `uv run python -m pytest tests/test_generalized_lift_instrument.py -q` and the +two entry points below — everything here is deterministic recompute, nothing is stored state. + +--- + +## 1. What this arc added (seams S1–S5, all merged to the arc branch) + +| Seam | Mechanism | Where | +| :--- | :--- | :--- | +| S1 | Egress `readback_eligible` → geometric token readback (⟨ψ ψ̃_T⟩₀ spectrum) + "hearing ourselves think" round-trip via `WaveManifold.phase_correlation` | `core/physics/linguistic_readback.py` | +| S2 | Chiral sgn(Q_top) precondition composed into the PASS-gated biography write-path (provenance v2) + first real harness-driven caller + ` 0.99 criterion). +- Reading the lift honestly (disclosed in the instrument's own notes): the recognition delta + isolates the **relax+readback stages'** contribution against a constraint-blind baseline — an + eigensolver baseline WITH access to H would reach parity. It proves the loop is closed and + articulate, not that the corridor out-reasons the symbolic engine. +- Multimodal parity detail: the vision token already resonates ≥ 0.4 with the audio-only partial + (compiler versors overlap), so the baseline also names both percepts. Measured, disclosed. + +**Scope limitation (recorded, not silently dropped):** GSM8K / natural-language arithmetic is NOT +ingestible by corridor v1 — no reader→Hamiltonian compiler exists beyond the ≤5-atom propositional +and quadratic-well domains. That compiler is the composition frontier; this instrument is the +harness already waiting to measure it. + +## 3. Ports + handoff evidence (ADR-0247 / ADR-0248) + +Entry point: `evals.lift_evidence_handoff.run_lift_evidence_handoff()` — two REAL certified +lifecycle turns (relax → egress → governed f64→f32 `serving_cast`), each run through +`IdentityPort` + `PrecisionPort` (Ring-2 seven-stage grammar) and fused by `coordinate_handoff` +(Ring-3). The acceptance evidence is the **contrast**: + +| Turn | IdentityPort | PrecisionPort | Replay chain | Handoff | +| :--- | :--- | :--- | :--- | :--- | +| identity-action (canonical lawful turn under H_id={I}) | proceed | proceed | verified | **proceed** (digest `580e686463b06af6…`) | +| frame-rotating (e1∧e2 rotor, θ=0.5) | **abstain: `d_stab>epsilon_turn`** | proceed | verified | **abstain: `port:identity:d_stab>epsilon_turn`** (digest `6ce9949fccdd224d…`) | + +The machinery discriminates: lawful turns route through, frame-rotating turns abstain with typed, +port-attributed reasons, and both replay chains verify. Flags stayed off; policy is +`AdmissionPolicy.placeholder_default()` (uncalibrated — activation would still be refused by the +serve gate, exactly as ADR-0246 §3.7 requires). + +## 4. Ring-2 Smith-chart conformity note (master-prompt §3) + +Assessed, nothing built (no ADR authorizes new machinery): the Ring-2 grammar already frames +ports as the conformal interconnect between non-identical native geometries (`core/ports/adapters.py` +docstring pins the future Atlas/Evidence/Articulation adapter slots). No Smith-chart math library +was written — the directive's own constraint. When a hyperbolic-atlas port lands, its Z→Γ mapping +belongs in the adapter as Spin(4,1) conformal rotor transport; that is a future ADR's work. + +## 5. Master-prompt adjudication deltas (applied vs corrected) + +- Applied to the REAL tree: `multimodal_lifecycle.py` → `cognitive_lifecycle.py`; + `ingest_sensorium` → `ingest_context`; `GoldTetherMonitor.calculate_coherence_residual` → + `coherence_residual` / `update`. +- "Wire Fibonacci to calibrate κ/θ" — already satisfied in the evals quarantine + (`evals/adr_0244_gamma_calibration`, `evals/analogical_transfer/kappa_calibration.py`); wiring it + into live streams stays rejected (A-04 / I-03 / R-04). +- §4.1 "fail closed, defaulting back to κ=1.0" — contradiction resolved in favor of the code's + actual contract: typed `OptimizationFailure`, consumers keep the last RATIFIED constant; a silent + 1.0 default is exactly what the same directive's Subsystem-B clause prohibits. +- §4.3 "remove `default=str` fallbacks in content-id serialization" — verified already clean: the + only `default=str` occurrences are human-readable CLI display printing (`core/cli.py`), which + ADR-0245 §2.3 permits; every content-id site refuses non-serializable payloads. +- §Phase-2 chiral composition — implemented, with the honesty theorem pinned: I₅ is central in odd + Cl(4,1), so closed versors have Q ≡ 0 exactly; the precondition is vacuous-by-theorem on + admissible trajectories and LIVE against raw non-versor mirror flips (both pinned in tests). diff --git a/docs/research/intelligence-loop-homestretch-plan-2026-07-18.md b/docs/research/intelligence-loop-homestretch-plan-2026-07-18.md index 11b90cc3..8aae930b 100644 --- a/docs/research/intelligence-loop-homestretch-plan-2026-07-18.md +++ b/docs/research/intelligence-loop-homestretch-plan-2026-07-18.md @@ -35,14 +35,14 @@ Close the comprehend → reason → articulate → contemplate → learn loop an ## 3. Phases -### Phase 0 — Worktree + record (small) — PARTIALLY DONE +### Phase 0 — Worktree + record (small) — DONE - [x] Spark PDFs copied to `docs/research/`. - [x] Adjudication doc written (`spark-audit-adjudication-2026-07-18.md`). - [x] This plan doc written. -- [ ] Create worktree off `forgejo/main`; run smoke + fast lanes → green baseline recorded. -- [ ] Commit the three docs + PDFs as the arc's opening record. +- [x] Worktree `feat/intelligence-loop-arc` off `forgejo/main`; baseline smoke 176 passed (133s). +- [x] Opening record committed (`fcea2d3a`). -### Phase 1 — Close the articulation seam, S1 (medium) +### Phase 1 — Close the articulation seam, S1 (medium) — DONE (`19d5731a`) Wire a `readback` stage into the lifecycle corridor (eval-tier, off-serving): - `egress` `route="readback_eligible"` → `VocabManifold.nearest()` token selection (`cognitive_lifecycle.py` gains a vocab consumer; today it imports no `vocab`, `:68-76`). @@ -53,7 +53,7 @@ Wire a `readback` stage into the lifecycle corridor (eval-tier, off-serving): - TDD anchors: `tests/test_adr_0243_cognitive_lifecycle.py`, `tests/test_vocab_manifold_invariants.py`. - Scope guard: within Accepted ADR-0243 §2.3 (readback rules); if design exceeds it → ADR amendment, not silent drift. -### Phase 2 — Activate the learning write-path, S2 (medium) +### Phase 2 — Activate the learning write-path, S2 (medium) — DONE (`f16b4a60`) - Compose the chiral Q_top latch (`chiral_gate.py`) into `integrate_validated_biography` (`biography_wiring.py:174`): charge conservation becomes a precondition of biography updates. - First real caller: harness-driven — after ADR-0240 validation PASS, integrate holonomy; prove @@ -62,7 +62,7 @@ Wire a `readback` stage into the lifecycle corridor (eval-tier, off-serving): - Cleanup-as-you-find: `biography.py:62` bare `.tobytes()` → explicit ` int: + return self.corridor_correct - self.baseline_correct + + @property + def verdict(self) -> str: + if self.delta_correct > 0: + return _LIFT + return _PARITY if self.delta_correct == 0 else _DEFICIT + + def as_dict(self) -> dict[str, Any]: + return { + "domain_id": self.domain_id, + "n_cases": self.n_cases, + "corridor": { + "correct": self.corridor_correct, + "wrong": self.corridor_wrong, + "refused": self.corridor_refused, + }, + "baseline": { + "correct": self.baseline_correct, + "wrong": self.baseline_wrong, + "refused": self.baseline_refused, + }, + "delta_correct": self.delta_correct, + "verdict": self.verdict, + "notes": list(self.notes), + "cases": list(self.cases), + } + + +@dataclass(frozen=True, slots=True) +class LiftInstrumentReport: + outcomes: tuple[DomainOutcome, ...] + wrong_zero_guard_held: bool + honest_null: bool + scope_limitations: tuple[str, ...] + + def as_dict(self) -> dict[str, Any]: + return { + "outcomes": [o.as_dict() for o in self.outcomes], + "wrong_zero_guard_held": self.wrong_zero_guard_held, + "honest_null": self.honest_null, + "scope_limitations": list(self.scope_limitations), + } + + +# --- Domain A: propositional entailment ------------------------------------------------ + +Literal = tuple[str, bool] +Clause = tuple[Literal, ...] + +# (case_id, atoms, premise clauses (CNF), conclusion clause (single literal)) +_PROP_CASES: tuple[tuple[str, tuple[str, ...], tuple[Clause, ...], Literal], ...] = ( + ("modus-ponens", ("a", "b"), ((("a", True),), (("a", False), ("b", True))), ("b", True)), + ( + "chain-3", + ("a", "b", "c"), + ((("a", True),), (("a", False), ("b", True)), (("b", False), ("c", True))), + ("c", True), + ), + ("disj-not-entailed", ("a", "b"), ((("a", True), ("b", True)),), ("a", True)), + ("unsat-ex-falso", ("a", "b"), ((("a", True),), (("a", False),)), ("b", True)), + ( + "neg-conclusion", + ("a", "b"), + ((("a", True),), (("a", False), ("b", False))), + ("b", False), + ), + ("no-information", ("a", "b"), ((("a", True),),), ("b", True)), + ( + "resolution", + ("a", "b", "c"), + ( + (("a", True), ("b", True)), + (("a", False), ("c", True)), + (("b", False), ("c", True)), + ), + ("c", True), + ), + ( + "contrapositive", + ("a", "b"), + ((("a", False), ("b", True)), (("b", False),)), + ("a", False), + ), + ("premise-restates", ("a", "b"), ((("a", True),),), ("a", True)), + ("wide-disj-not-entailed", ("a", "b", "c"), ((("a", True), ("b", True), ("c", True)),), ("c", True)), +) + + +def _lit_formula(lit: Literal) -> str: + atom, positive = lit + return atom if positive else f"(not {atom})" + + +def _clauses_formula(clauses: Sequence[Clause]) -> str: + return " and ".join("(" + " or ".join(_lit_formula(l) for l in clause) + ")" for clause in clauses) + + +def _truth_table_entailed( + atoms: Sequence[str], clauses: Sequence[Clause], conclusion: Literal +) -> bool: + """Independent gold: every model of the premises satisfies the conclusion.""" + for values in product((False, True), repeat=len(atoms)): + env = dict(zip(atoms, values)) + if all(any(env[a] == pos for a, pos in clause) for clause in clauses): + atom, positive = conclusion + if env[atom] != positive: + return False + return True + + +def run_propositional_domain() -> DomainOutcome: + corridor_correct = corridor_wrong = corridor_refused = 0 + baseline_correct = baseline_wrong = baseline_refused = 0 + rows: list[dict[str, Any]] = [] + for case_id, atoms, clauses, conclusion in _PROP_CASES: + gold = _truth_table_entailed(atoms, clauses, conclusion) + + corridor_entailed = propositional_entails( + PropositionalProblem(atoms=atoms, clauses=clauses), (conclusion,) + ).entailed + if corridor_entailed == gold: + corridor_correct += 1 + else: + corridor_wrong += 1 + + premises_f = _clauses_formula(clauses) + conjunction_f = f"({premises_f}) and ({_lit_formula(conclusion)})" + robdd = check_equivalence(conjunction_f, premises_f) + if robdd.verdict is Verdict.REFUSED: + baseline_refused += 1 + baseline_entailed: bool | None = None + else: + baseline_entailed = robdd.verdict is Verdict.EQUIVALENT + if baseline_entailed == gold: + baseline_correct += 1 + else: + baseline_wrong += 1 + + rows.append( + { + "case_id": case_id, + "gold_entailed": gold, + "corridor_entailed": corridor_entailed, + "baseline_entailed": baseline_entailed, + } + ) + return DomainOutcome( + domain_id="propositional", + n_cases=len(_PROP_CASES), + corridor_correct=corridor_correct, + corridor_wrong=corridor_wrong, + corridor_refused=corridor_refused, + baseline_correct=baseline_correct, + baseline_wrong=baseline_wrong, + baseline_refused=baseline_refused, + notes=( + "Baseline is the deductive flagship (ROBDD); PARITY here is the " + "expected honest outcome — both paths are exact on this regime.", + "wrong=0 guard binds the corridor on this domain.", + ), + cases=tuple(rows), + ) + + +# --- Domain B: constrained recognition + articulation ---------------------------------- + +_RECOGNITION_GRID: tuple[tuple[int, float], ...] = tuple( + (plane, angle) for plane in (6, 7, 8) for angle in (0.4, 0.8, 1.2) +) + + +def _grid_word(plane: int, angle: float) -> str: + return f"mode-p{plane}-a{int(round(angle * 10))}" + + +def _recognition_vocab() -> tuple[VocabManifold, tuple[np.ndarray, ...]]: + vocab = VocabManifold() + versors: list[np.ndarray] = [] + for plane, angle in _RECOGNITION_GRID: + v = np.asarray(make_rotor_from_angle(angle, bivector_idx=plane), dtype=np.float64) + vocab.add(_grid_word(plane, angle), v) + versors.append(v) + return vocab, tuple(versors) + + +def _argmax_word(psi: np.ndarray, vocab: VocabManifold, manifold: WaveManifold) -> str: + best_score, best_idx = -np.inf, -1 + for i in range(len(vocab)): + score = float(manifold.phase_correlation(psi, np.asarray(vocab.get_versor_at(i), dtype=np.float64))) / 2.0 + if score > best_score: + best_score, best_idx = score, i + return vocab.get_word_at(best_idx) + + +def run_recognition_domain() -> DomainOutcome: + vocab, versors = _recognition_vocab() + manifold = WaveManifold() + engine = CognitiveLifecycleEngine() + n = len(_RECOGNITION_GRID) + corridor_correct = corridor_wrong = corridor_refused = 0 + baseline_correct = baseline_wrong = 0 + rows: list[dict[str, Any]] = [] + for i, (plane, angle) in enumerate(_RECOGNITION_GRID): + target_word = _grid_word(plane, angle) + target = versors[i] + distractor = versors[(i + 1) % n] + packets = ( + ModalityPacket(modality_id="mix:target", coefficients=0.45 * target), + ModalityPacket(modality_id="mix:distractor", coefficients=0.55 * distractor), + ) + domain_id = f"recognition:{target_word}" + + baseline_word = _argmax_word( + ingest_context(packets, domain_id).psi, vocab, manifold + ) + if baseline_word == target_word: + baseline_correct += 1 + else: + baseline_wrong += 1 + + row: dict[str, Any] = { + "case_id": target_word, + "baseline_word": baseline_word, + } + try: + outcome = engine.solve( + packets, + domain_id, + compile_quadratic_well(target), + energy_inputs=_HOT_ENERGY, + ) + readback, roundtrip = articulate_outcome( + outcome, vocab, min_resonance=_MIN_RESONANCE, max_tokens=1 + ) + corridor_word = readback.tokens[0].word + row["corridor_word"] = corridor_word + row["roundtrip_agreement"] = roundtrip.agreement + if corridor_word == target_word: + corridor_correct += 1 + else: + corridor_wrong += 1 + except ReadbackRefusal as exc: + corridor_refused += 1 + row["corridor_word"] = None + row["corridor_refusal"] = exc.reason + rows.append(row) + return DomainOutcome( + domain_id="constrained-recognition", + n_cases=n, + corridor_correct=corridor_correct, + corridor_wrong=corridor_wrong, + corridor_refused=corridor_refused, + baseline_correct=baseline_correct, + baseline_wrong=baseline_wrong, + baseline_refused=0, + notes=( + "Baseline is the constraint-blind ingress argmax over the same " + "vocabulary; the delta isolates the relax+readback stages.", + "An eigensolver baseline WITH access to H would reach parity — " + "this domain measures loop integrity, not open-ended capability.", + ), + cases=tuple(rows), + ) + + +# --- Domain C: multimodal completion --------------------------------------------------- + + +def run_multimodal_domain() -> DomainOutcome: + from evals.adr_0243_cognitive_lifecycle import _fixed_audio_tone, _fixed_vision_tile + from core.physics.sensorium_wave_feed import packet_from_compilation_unit + from sensorium.audio.compiler import AudioCompiler + from sensorium.vision import VisionCompiler + + audio_unit = AudioCompiler().compile(_fixed_audio_tone(24_000, 0.25, 440.0), 24_000) + vision_unit = VisionCompiler().compile_tile(_fixed_vision_tile()) + audio_pkt = packet_from_compilation_unit("audio", audio_unit) + vision_pkt = packet_from_compilation_unit("vision", vision_unit) + + vocab = VocabManifold() + vocab.add("audio-tone", np.asarray(audio_pkt.coefficients, dtype=np.float64)) + vocab.add("vision-tile", np.asarray(vision_pkt.coefficients, dtype=np.float64)) + manifold = WaveManifold() + + full = ingest_context((audio_pkt, vision_pkt), "multimodal-completion") + partial = ingest_context((audio_pkt,), "multimodal-completion") + expected_words = {"audio-tone", "vision-tile"} + + def _resonant_words(psi: np.ndarray) -> set[str]: + found = set() + for i in range(len(vocab)): + score = ( + float( + manifold.phase_correlation( + psi, np.asarray(vocab.get_versor_at(i), dtype=np.float64) + ) + ) + / 2.0 + ) + if score >= _MIN_RESONANCE: + found.add(vocab.get_word_at(i)) + return found + + baseline_words = _resonant_words(partial.psi) + baseline_correct = int(baseline_words == expected_words) + + corridor_correct = corridor_wrong = corridor_refused = 0 + row: dict[str, Any] = { + "case_id": "audio-partial-to-full", + "baseline_words": sorted(baseline_words), + } + result = relax_to_ground(partial.psi, compile_quadratic_well(full.psi)) + verdict = egress_gate(result.psi_steady, result.certificate, **_HOT_ENERGY) + try: + readback = linguistic_readback( + result.psi_steady, + result.certificate, + verdict, + vocab, + min_resonance=_MIN_RESONANCE, + max_tokens=2, + ) + corridor_words = set(readback.words) + row["corridor_words"] = sorted(corridor_words) + if corridor_words == expected_words: + corridor_correct = 1 + else: + corridor_wrong = 1 + except ReadbackRefusal as exc: + corridor_refused = 1 + row["corridor_refusal"] = exc.reason + + return DomainOutcome( + domain_id="multimodal-completion", + n_cases=1, + corridor_correct=corridor_correct, + corridor_wrong=corridor_wrong, + corridor_refused=corridor_refused, + baseline_correct=baseline_correct, + baseline_wrong=1 - baseline_correct, + baseline_refused=0, + notes=( + "Correct = articulation names BOTH constituent percepts; baseline " + "articulates the raw audio-only partial ingress.", + ), + cases=(row,), + ) + + +# --- Composed instrument --------------------------------------------------------------- + + +def run_generalized_lift_instrument() -> LiftInstrumentReport: + outcomes = ( + run_propositional_domain(), + run_recognition_domain(), + run_multimodal_domain(), + ) + propositional = outcomes[0] + return LiftInstrumentReport( + outcomes=outcomes, + wrong_zero_guard_held=(propositional.corridor_wrong == 0), + honest_null=all(o.delta_correct <= 0 for o in outcomes), + scope_limitations=( + "GSM8K / natural-language arithmetic is NOT ingestible by corridor " + "v1: no reader-to-Hamiltonian compiler exists beyond the <=5-atom " + "propositional and quadratic-well domains. Recorded, not silently " + "dropped; that compiler is the composition frontier this " + "instrument is waiting to measure.", + ), + ) diff --git a/evals/lift_evidence_handoff.py b/evals/lift_evidence_handoff.py new file mode 100644 index 00000000..20fb0fae --- /dev/null +++ b/evals/lift_evidence_handoff.py @@ -0,0 +1,109 @@ +"""ADR-0247/0248 evidence run — lift-instrument turns through ports + handoff (seam S5). + +Routes two corridor turns through the Ring-2 residual protocol and the Ring-3 +integrity handoff, OFF-SERVING (no flags touched — this produces §8 ruling +evidence, not activation): + +* an **identity-action** turn (the canonical identity versor — under the + locked ADR-0246 stabilizer H_id={I} the identity action is the ONLY lawful + action, so this is the canonical lawful turn, not a toy) → expected: both + ports PROCEED, handoff PROCEED; +* a **frame-rotating** turn (e1∧e2 rotor — an in-span rotation the ADR-0246 + stabilizer refuses) → expected: identity port ABSTAIN with typed reasons, + handoff ABSTAIN. + +Each turn is a REAL certified lifecycle outcome (relax → egress → governed +serving cast), so the PrecisionPort witnesses a genuine f64→f32 transport. +The pair demonstrates the handoff DISCRIMINATES — the acceptance evidence is +the contrast, not a single green path. +""" + +from __future__ import annotations + +from typing import Any + +import numpy as np + +from algebra.cl41 import N_COMPONENTS +from core.epistemic_state import EpistemicState, NormativeClearance +from core.physics.cognitive_lifecycle import ( + CognitiveLifecycleEngine, + compile_quadratic_well, + serving_cast, +) +from core.physics.identity_action import AdmissionPolicy +from core.physics.identity_manifold import IdentityManifoldGeometry +from core.physics.sensorium_wave_feed import ModalityPacket +from core.ports.adapters import ( + IdentityPort, + PrecisionPort, + PrecisionSubject, +) +from core.ports.integrity_handoff import coordinate_handoff +from core.ports.residual_protocol import run_residual_protocol, verify_replay_chain + +__all__ = ["run_lift_evidence_handoff"] + +_E12 = 6 # grade-2 bivector block index of e1∧e2 + + +def _rotor_e12(theta: float) -> np.ndarray: + v = np.zeros(N_COMPONENTS, dtype=np.float64) + v[0] = np.cos(theta / 2.0) + v[_E12] = np.sin(theta / 2.0) + return v + + +def _identity_versor() -> np.ndarray: + v = np.zeros(N_COMPONENTS, dtype=np.float64) + v[0] = 1.0 + return v + + +def _certified_turn(target: np.ndarray, label: str) -> tuple[Any, Any]: + """One real corridor turn: relax onto the target versor, egress, cast.""" + engine = CognitiveLifecycleEngine() + packets = (ModalityPacket(modality_id=f"seed:{label}", coefficients=target),) + outcome = engine.solve(packets, f"handoff-evidence:{label}", compile_quadratic_well(target)) + serving = serving_cast( + outcome.relaxation.psi_steady, outcome.relaxation.certificate, outcome.verdict + ) + return outcome, serving + + +def run_lift_evidence_handoff() -> dict[str, Any]: + geometry = IdentityManifoldGeometry.from_directions( + ((1.0, 0.0, 0.0), (0.0, 1.0, 0.0), (0.0, 0.0, 1.0)) + ) + policy = AdmissionPolicy.placeholder_default() + identity_port = IdentityPort(geometry, policy) + precision_port = PrecisionPort(1e-6) + + artifact: dict[str, Any] = {"turns": {}, "adr_refs": ["ADR-0246", "ADR-0247", "ADR-0248"]} + for label, target in ( + ("identity-action", _identity_versor()), + ("frame-rotating", _rotor_e12(0.5)), + ): + outcome, serving = _certified_turn(target, label) + chain, identity_decision = run_residual_protocol( + identity_port, outcome.relaxation.psi_steady, () + ) + chain, precision_decision = run_residual_protocol( + precision_port, PrecisionSubject.from_serving_state(serving), chain + ) + handoff = coordinate_handoff( + chain, + epistemic_state=EpistemicState.DECODED, + normative_clearance=NormativeClearance.CLEARED, + ) + artifact["turns"][label] = { + "outcome_id": outcome.outcome_id, + "serving": serving.as_dict(), + "identity_action": identity_decision.as_dict(), + "precision_action": precision_decision.as_dict(), + "chain_verified": verify_replay_chain(chain), + "chain_records": [r.as_dict() for r in chain], + "handoff": handoff.as_dict(), + "handoff_digest": handoff.handoff_digest(), + } + return artifact diff --git a/tests/test_generalized_lift_instrument.py b/tests/test_generalized_lift_instrument.py new file mode 100644 index 00000000..a874d02d --- /dev/null +++ b/tests/test_generalized_lift_instrument.py @@ -0,0 +1,80 @@ +"""Seam S4/S5 pins — generalized-lift instrument + ports/handoff evidence. + +The instrument is the arbiter of "generalized lift, no overfitting": identical +compiled problems for corridor and baseline, independent truth-table gold, +wrong=0 guard on the propositional (flagship-regime) domain, and an +honest-NULL protocol. The handoff evidence pins that Ring-2/Ring-3 machinery +DISCRIMINATES (proceed on the lawful identity-action turn, typed abstain on a +frame-rotating turn) — off-serving, flags untouched. +""" + +from __future__ import annotations + +import json + +import pytest + +from evals.generalized_lift_instrument import ( + run_generalized_lift_instrument, +) +from evals.lift_evidence_handoff import run_lift_evidence_handoff + + +@pytest.fixture(scope="module") +def report(): + return run_generalized_lift_instrument() + + +def test_propositional_domain_wrong_zero_and_gold_agreement(report): + prop = report.outcomes[0] + assert prop.domain_id == "propositional" + assert prop.corridor_wrong == 0 # the flagship-regime guard + assert prop.baseline_wrong == 0 + for row in prop.cases: + assert row["corridor_entailed"] == row["gold_entailed"] + assert row["baseline_entailed"] == row["gold_entailed"] + assert prop.verdict == "PARITY" # honest expected outcome: both exact + + +def test_recognition_domain_shows_relax_readback_lift(report): + rec = report.outcomes[1] + assert rec.domain_id == "constrained-recognition" + assert rec.corridor_correct == rec.n_cases # relax+readback recovers every mode + assert rec.corridor_wrong == 0 and rec.corridor_refused == 0 + assert rec.baseline_correct < rec.n_cases # constraint-blind argmax fails + assert rec.delta_correct > 0 and rec.verdict == "LIFT" + for row in rec.cases: + assert row["roundtrip_agreement"] > 0.99 # hearing ourselves think + + +def test_multimodal_domain_recorded_honestly(report): + multi = report.outcomes[2] + assert multi.domain_id == "multimodal-completion" + assert multi.n_cases == 1 + assert multi.corridor_wrong == 0 + assert multi.verdict in ("LIFT", "PARITY") # measured, never assumed + + +def test_report_flags_and_scope_limitations(report): + assert report.wrong_zero_guard_held + assert isinstance(report.honest_null, bool) + assert any("GSM8K" in note for note in report.scope_limitations) + json.dumps(report.as_dict()) # JSON-safe artifact + + +def test_handoff_evidence_discriminates(): + artifact = run_lift_evidence_handoff() + lawful = artifact["turns"]["identity-action"] + rotated = artifact["turns"]["frame-rotating"] + + assert lawful["chain_verified"] and rotated["chain_verified"] + assert lawful["identity_action"]["action"] == "proceed" + assert lawful["precision_action"]["action"] == "proceed" + assert lawful["handoff"]["handoff"] == "proceed" + + assert rotated["identity_action"]["action"] == "abstain" + assert "d_stab>epsilon_turn" in rotated["identity_action"]["reasons"] + assert rotated["handoff"]["handoff"] == "abstain" + assert any(r.startswith("port:identity:") for r in rotated["handoff"]["reasons"]) + + json.dumps(artifact) # JSON-safe acceptance-evidence artifact