Opus 4.8 audit of Fable's slice-1 scaffold + authorized hardening (Step 2).
Stacked on feat/adr-0246-slice1-scaffold. Not a PR, not merged, no status flip.
AUDIT (Step 1) — VERDICT PASS. Re-derived A(F)/d_orth/d_stab/typed channels from
§3.1/3.2/3.6 independently: BIT-EXACT vs impl (0.00e+00). Confirmed in code (not
just comments): H_id={I} locked (singleton hardcoded, no enlargement path); path
composes lawful only, refused=break marker, never raw, never soft-I; no §7 scope
creep; serve byte-identity. Finding F1 (doc, not bug): composition uses the raw
CERTIFIED action (required for §6.2 accumulation, else A_path≡I detects nothing) —
now documented for ADR ratification.
HARDENING (Step 2):
§3.7 pure surface (identity_action.py): AdmissionPolicy (calibrated flag),
evaluate_admission (admit-or-abstain, no corrector); CERTIFIED_GAMMA_ID pinned
== identity._WAVE_LEAKAGE_BOUND; all other bounds UNCERTIFIED placeholders.
§3.7 SERVE WIRING (Steps 2.1/2.2): new default-off flag identity_action_surface
(config); threaded chat/runtime -> check -> _wave_field_score; refusal folds
into flagged -> existing would_violate/conjugate_correct abstains; IdentityScore
gains action_surface_active/d_orth/d_stab (legacy defaults). Flag-off is
byte-identical (D4 gate surfaces green unchanged; smoke 176 post-wiring).
§6.3 DISCRIMINATION REPORT (evals/adr_0246_discrimination) — HONEST numbers:
benign pass 0.00, false refusal 1.00, adversarial detect 1.00, control pass
1.00; d_stab AUC 0.375 (95% CI [0.15,0.62]) — BELOW chance. Benign cognition
sits ~18x farther from the frame (mean d_stab 27.8) than the attacks (1.55).
The gate refuses everything and does NOT discriminate; must stay off; usable
separation needs the §11 grounding work, not threshold tuning. Claims language
enforced: 'lawfulness relative to the declared frozen frame', NOT 'semantic
inalienability'.
Ledger raw-sneak hardening test (Step 2.3): mixed lawful/refused sequence must
equal the lawful sub-product, fails if raw sneaks into A_path_lawful.
Handoff to Sonnet: §4.1 per-turn record, path-ledger serve integration, ADR-0246
body + acceptance packet (Proposed, no self-Accept), §11 grounding-feasibility.
[Verification]: uv run core test --suite smoke -q => 176 passed (post-wiring);
ADR-0246 suites 80 passed; egress wiring + D4 gate surfaces 47 passed
(byte-identity); §6.1/6.2 eval 14/14; discrimination report numbers above.
See docs/audit/adr-0246-slice1-opus-audit-and-hardening.md + run log.
9.3 KiB
ADR-0246 Slice-1 — Opus 4.8 Audit, Hardening & Discrimination Report
Reviewer: Opus 4.8. Branch: feat/adr-0246-slice1-hardened (stacked on
feat/adr-0246-slice1-scaffold, unmerged). Not a PR, not merged, no status
flip. Review gate only — ratification is human (Shay).
Audited artifact: Fable 5's feat/adr-0246-slice1-scaffold.
1. Step-1 adversarial audit — VERDICT: PASS (one documentation finding)
Nothing was trusted from Fable's run log; every claim was reproduced.
| Check | Method | Result |
|---|---|---|
| D4 closed (hard gate) | independent git cat-file/status on main |
✅ both acceptance packets on main; ADR-0244 Accepted/ratified; main @ 04d67ca5 |
A(F), d_orth, d_stab, typed channels correct |
re-derived from §3.1/3.2/3.6 from scratch (independent sandwich/Gram/projection) and diffed vs impl | ✅ bit-exact, max discrepancy 0.00e+00 |
H_id={I} genuinely locked in code |
read IdentityStabilizer, advance_identity_path |
✅ singleton hardcoded in the path; no enlargement parameter or path; grep found no enlargement/permutation/soft-projection code |
Path composes lawful only, never raw, no soft-I |
read composition; grep | ✅ lawful turns compose; refused turns → path_break, excluded; a_path untouched on refuse |
| Placeholders logged, not baked | read policy/eval | ✅ PLACEHOLDER_* marked; PathBudget always caller-supplied |
| §7 scope creep / serve byte-identity | full git diff main..scaffold; per-function diff |
✅ no chat/runtime.py, identity.py, or flag change; serve-called geometry functions untouched (additions only) |
| §6.1/§6.2 matrix reproduced | ran from clean tree | ✅ 14/14 eval cases; 64 tests |
Finding F1 (documentation, NOT a code bug)
advance_identity_path composes the raw certified action A_t
(a_path = A_t @ a_path), not a literal I. This is required and correct —
composing Is would make A_path ≡ I and defeat the slow-drift detection that is
the ledger's whole purpose (§2 gap). But it departs from a literal reading of
§3.4-step-4's H_t ∈ H_id notation and was undocumented. Correction applied:
an explicit "Composition semantics (ratification-relevant)" paragraph now states
the reading in the docstring; the ADR body should ratify it (see §4 handoff).
No locked decision (§3) or non-goal (§7) was violated to make anything pass.
2. Hardening & additions (this branch) — diff-style note
Every change vs Fable's scaffold, and why:
| File | Change | Why |
|---|---|---|
identity_action.py |
F1 docstring on advance_identity_path |
ratify the raw-certified-action composition (audit finding) |
identity_action.py |
§3.7 admit surface: CERTIFIED_GAMMA_ID, PLACEHOLDER_*, AdmissionPolicy (calibrated=False), AdmissionResult, evaluate_admission |
pure, testable admit-or-abstain surface (§3.7) — the reusable core for the discrimination report and the future serve wiring. No corrector. |
evals/adr_0246_discrimination/ |
§6.3 discrimination report (new) | honest benign-vs-adversarial numbers with bootstrap-CI AUC (§6.3/§10) |
tests/test_adr_0246_admission.py |
pins for evaluate_admission + the honest verdict + γ_id-drift guard |
fail loudly if the surface or the honesty claim regresses |
tests/test_adr_0246_path_ledger.py |
raw-sneak hardening test | fails if a raw product ever enters A_path_lawful (Step 2.3) |
Nothing here touches serve, flags, H_id, or adds a corrector. evaluate_admission
lives in the pure identity_action.py; γ_id is pinned equal to
identity._WAVE_LEAKAGE_BOUND by a test so it cannot drift.
3. §6.3 Discrimination report — HONEST numbers (no inflation)
Full artifact: docs/audit/artifacts/adr-0246-slice1-discrimination-report.json
(live benign cohort n=13, adversarial n=8, synthetic-near-identity control n=5;
policy = placeholder, calibrated=False).
| Metric | Value |
|---|---|
| benign pass rate | 0.00 |
| false refusal rate | 1.00 |
| adversarial detection rate | 1.00 |
| synthetic-near-identity pass rate (control) | 1.00 |
d_stab AUC (adversarial vs benign) |
0.375, 95% CI [0.15, 0.62] |
leakage_rms AUC |
0.18, 95% CI [0.04, 0.39] |
benign d_stab (min/mean/max) |
0.15 / 27.8 / 228 |
adversarial d_stab (min/mean/max) |
0.46 / 1.55 / 2.83 |
Honest finding (verbatim in the report): the §3.7 admit surface on the declared
placeholder frame refuses benign and adversarial versors alike — benign
false-refusal is 100% and d_stab does not separate the classes (AUC 0.375 is
below chance: real benign cognition sits ~18× farther from the identity frame
than the crafted attacks do). A gate that refuses everything trivially "detects"
every attack but is not a discriminator. The synthetic-near-identity control
passing at 100% confirms the gate mechanism is sound — the failure is entirely
that benign cognition does not live near the declared frame. This reproduces and
sharpens the D4 / slice-0 result at the fuller §3.7 surface.
Claims language (§10 #9), enforced in code and report: this supports only "lawfulness relative to the declared frozen frame." It does not support any claim of "semantic inalienability of the value labels." The gate must stay default-off; usable separation requires the §11 dynamics-grounding work — not threshold tuning.
4. §3.7 egress serve wiring (Steps 2.1/2.2) — DONE (Opus)
Wired the fuller admit surface into the live gate, flag-gated + default-off + byte-identical:
- New default-off flag
identity_action_surfaceincore/config.pyRuntimeConfig(separate fromidentity_wave_gate; documented as uncalibrated / not authorized live). - Threaded an
AdmissionPolicy(placeholder_default(),calibrated=False) throughchat/runtime.py→IdentityCheck.check(..., admission_policy=)→_wave_field_score. Only acts whenidentity_wave_gateis also on (awave_fieldexists). - On the wave path with the policy present,
evaluate_admission(...)runs; a refusal folds intoflagged, so the existingwould_violate/conjugate_correct(refuse=True)egress abstains — admit-or-abstain, no corrector.MalformedVersorErrorfrom the primitives propagates as a fail-closed refusal. IdentityScoregained optionalaction_surface_active/d_orth/d_stabwith legacy defaults, so flag-off is byte-identical.
Verified: tests/test_adr_0246_egress_wiring.py (flag default-off; flag-off
check() byte-identical incl. default §3.7 fields; flag-on refuses a tilt and
admits true near-identity) + all D4 gate surfaces green unchanged (wave /
runtime / eval / test_identity_gate) — serve byte-identity confirmed.
Guardrail: calibrated=False + flag default-off ⇒ not live-authorized; the §3
discrimination numbers are the evidence it must stay off pending §11 + calibration.
4c. Handoff to Sonnet 5 — remaining ADR-0246 work
- §4.1
IdentityActionRecordper-turn telemetry — not built (only the §4.2IdentityPathLedgerexists). Add the per-turn record (field/record digests, gate/policy versions, the §3.7 measures) and emit it only when the wave/action path ran (preserve flag-off wire format, §4.3). Optionally surfaced_orth/d_stab/typed channels through the telemetry serializer. - Path-ledger ↔ serve integration — the ledger (
advance_identity_path) is pure and unit-tested but is not yet driven per-turn fromchat/runtime.py(session-scoped chain, hard-break on pack/geometry/policy/session). Wire it flag-gated + off, using the §3.5 scope keys, if a session path budget is wanted. - ADR-0246 body + acceptance packet — draft
docs/adr/ADR-0246-...mdas Proposed (no self-Accept — provenance guard). Must: state the F1 composition semantics explicitly; use the §10 claims language ("lawfulness relative to the declared frozen frame"); carry the honest §6.3 numbers; enumerate the placeholder ε's/τ's as uncalibrated. Acceptance packet per §10; §8 RULING PENDING. - §11 grounding-feasibility study (the larger research workstream) — does a held-out-stable, safety-relevant dynamics-invariant structure exist (fixed cohort splits, synthetic recovery controls, typed e4/e5 generator analysis, precision pairs, adversarial discrimination)? This is the only path to a gate that discriminates; the §6.3 report shows threshold tuning cannot get there.
- Open uncertainties still needing a ruling (Fable notes §4, unchanged):
general-pack
‖·‖_Gconvention; structurally-emptyspatial_foreignchannel; hard-break turn ownership; malformed-F/gate-routing boundary. Carry into the ADR.
5. Verification (run log: docs/audit/artifacts/adr-0246-slice1-hardened-runlog.txt)
- §6.1/§6.2 eval matrix → 14/14,
all_passed=True. - ADR-0246 suites (admission, egress wiring, induced-action, path-ledger incl. raw-sneak hardening, geometric suite, mismatch diagnostic) → 80 passed.
- Egress wiring + all D4 gate surfaces (wave/runtime/eval/
test_identity_gate) → 47 passed — flag-off byte-identity confirmed (D4 gate unchanged). uv run core test --suite smoke -q→ 176 passed (post-serve-wiring).- §6.3 discrimination report → the honest numbers in §3 above.