Rank 4 on the docket. The L10 always-on lane has passed a 5000-beat soak since
2026-07-19. That result lived as PROSE IN A CONTRACT FILE — no machine-readable
artifact, no run SHA, no rerun path, and a deterministic_digest that report.py
computed and nothing pinned. The contract's own closing line asked for the pin
("Pin it once the lane is trusted so a regression flips it") and it was never
added. A recorded prose result with no re-derivable artifact is testimony, not
evidence, and that file is exactly where the distinction was meant to be
enforced.
R-2 first, before the run — as the plan required, so the soak measured a
decided question rather than deciding one by accident.
R-2 was ruled B (beats govern cognition; wall-clock permitted in telemetry
only). Measured: this lane carries NO WALL-CLOCK AT ALL — runner.py,
predicates.py and report.py contain no time, datetime, or timestamp of any
kind. That is stronger than B requires, and it is KEPT rather than "completed",
for three reasons now recorded in the contract:
1. B's permissive half is a permission, not an obligation. Adding a field so
the code matches a ruling's letter, then adding a pin to guard the field,
is a mechanism whose only purpose is to be guarded.
2. "When was this last run?" is already the artifact's COMMIT DATE —
immutable, verifiable, stated once. A generated_at inside the report would
be a second copy of one fact, drifting the moment a report is regenerated
and not committed. That is G-23's defect class, registered this same week.
3. It keeps the digest honestly deterministic. A timestamp would have to be
excluded from it — a value in the artifact that the artifact's own
integrity check ignores.
The contract also records what WOULD change this: a predicate whose claim is
about duration rather than sequence (H5's resource-cost leg is the live
candidate). B is already ruled, so that needs no new decision — only the
predicate that requires it.
The re-run: 5000 beats, reboot at 2500, under the POST-R-3 profile. All four
predicates pass. Committed as
evals/l10_always_on/results/report-5000-6c67e88d.json, digest f814fd97….
The most useful result was not the pass. Every DETERMINISTIC fact reproduced
exactly — vault bounded at 6, convergence at beat 1 with a 4999-beat tail,
clean reboot with derived learning intact — while the machine-variant float
moved from the prose's 1.389e-07 to a worst of 6.177e-08. That is the first
real evidence that excluding versor_condition from the digest was correct
rather than convenient: a digest covering it would have flipped on a machine
change and taught every reader to ignore it. The prose figure is superseded,
not contradicted — it was never a reproducible quantity.
tests/test_l10_soak_evidence.py, on the gate, guarding two DIFFERENT failure
modes that are easy to conflate:
REGRESSION — the digest is pinned, plus the horizon (a 24-beat report pinned
as though it were the 5000-beat one would pass everything else), the four
predicates by name (all_gates_pass alone is satisfied by an empty predicate
list), and a vacuity guard on the attested set.
STALENESS — the mode nothing guarded, and the one that had ALREADY FIRED.
R-9 ruled the cadence change-triggered, not clock-triggered: local-first
doctrine makes "nightly" a ruling rather than a cron line, and a nightly that
does not run is green for the wrong reason (H-9). So instead of "when did we
last run it?", the pin asks the question that matters — DOES THIS EVIDENCE
STILL DESCRIBE THE CODE THAT SHIPS? — by holding the SHA-256 of every source
the soak attests: chat/always_on.py, chat/always_on_daemon.py, and the three
harness modules.
That is not hypothetical. It had already fired silently for six weeks: the
2026-07-19 soak ran with accrue_realized_knowledge absent from
CONTINUOUS_LIFE_CONFIG_FLAGS, so from R-3 (PR-5, this same day) onward the
recorded result attested a process that no longer existed — and nothing in the
repository could notice. HAD THIS PIN EXISTED, PR-5 WOULD HAVE FAILED IT. That
is precisely the intended behaviour, and it is verified by sabotage: touching
always_on_daemon.py turns it red.
The failure message says RE-RUN THE SOAK — DO NOT EDIT THE HASHES, because
updating a hash to match changed source without re-running converts the pin
into a record asserting evidence that does not exist. That is the failure this
whole arc is about, committed inside its own remedy.
Three sabotages observed red: a touched daemon (the staleness mode), a
regressed verdict in the artifact, and a shorter horizon pinned as if it were
5000 beats.
H1-H4 promoted onto the gate, and the cheap option was deliberately refused.
Measured 20.95s for the four real-soak "holds" and 0.46s for the eight mutation
"bites" — 98% of the cost is the holds. Shipping only the bites, or parking the
file in a curated-but-unreachable suite, would have satisfied R-9's letter and
bought nothing; a suite nobody runs is the non-guarantee this arc spent PR-4
closing. Those four holds are the ONLY tests on any gate that run
chat/always_on's run_continuous end-to-end, and that loop just shipped a
six-week silent divergence. 21s per push to actually run the process that broke
is the most defensible time on this gate. Removed from full_only_baseline.txt
in the same edit; the baseline shrank 746 -> 745.
Closes G-5. Proof-of-life has moved from prose to committed, pinned,
change-guarded evidence.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
Rank 3 on the docket. RuntimeConfig carries 32 boolean fields and nothing in
the repository stated that set, distinguished a capability flag from a posture
flag from a deployment knob, or recorded what evidence would flip any of them.
Two independent counts of the same set differed by eleven — the assessment said
"seventeen capability flags," N-7 measured 28 default-off plus 4 default-on —
and neither was wrong about what it looked at. They counted different things
because nothing declared what the set was.
R-3 (ruled A — incomplete flag set). accrue_realized_knowledge joins
CONTINUOUS_LIFE_CONFIG_FLAGS. Step D consolidates the realized facts Step B
writes, and the daemon forced the consumer without the producer: a continuous
life consolidating an empty set, which the mastery framework calls "garbage at
high speed." It shipped that way for six weeks.
The decisive evidence was in the code, not the documents. BOTH flags' comment
blocks in core/config.py already described the corrected profile —
accrue_realized_knowledge says "the production L10 process enables it alongside
persist_session_state", consolidate_determinations says "...alongside
accrue_realized_knowledge + persist_session_state". Three records (two
docstrings and the 07-25 verification doc) agreed with each other and disagreed
with four lines of code. Dormancy was a coherent ruling and would have cost
correcting three records that were RIGHT about the design; incompleteness cost
one line.
So the planned "N-6 docstring correction" deliverable DISSOLVED rather than
shipping: adding the flag made both docstrings true. Recorded explicitly,
because a PR that quietly ships less than it promised is the same divergence
class this arc exists to close.
R-4 (ruled A) — docs/specs/flag_register.md. All 32 booleans with class,
governing ADR or an honest "none", recorded rationale, and what evidence would
flip it. The classification is the load-bearing half and is stated as a
taxonomy rather than a label:
CAPABILITY — what CORE can do; engineering may flip it on named evidence
POSTURE — what CORE serves or refuses AS TRUE; the wrong=0 boundary,
ruling-only
DEPLOYMENT — cost, lifecycle, process shape; per-deployment
estimation_enabled and composed_surface are both "= False" and are not the same
kind of thing. Nothing said so before.
Profiles are declared as the UNIT of decision (R-4's mechanism): one-shot/eval
and continuous-life. The absence of a SERVING profile is stated deliberately —
the four POSTURE serving flags each ride their own license, and grouping them
before the ledger is sealed would license by association, the outcome R-8
exists to prevent.
tests/test_flag_register.py, on the gate. BIDIRECTIONAL, because either
direction alone would have passed through both failures it exists to prevent:
a flag added to config.py and not registered fails; a register row outliving
its flag fails. The counts the register's prose states (32 / 4 ON) are pinned,
since a stale count is how this register came to be needed. Four sabotages
observed red — new unregistered flag, deleted row, orphaned row, and a flag
silently flipped ON, the last caught by two independent pins.
THREE FINDINGS THE ITEM DID NOT ANTICIPATE, all from measuring rather than
reading.
1. Accumulated permissiveness, not just hesitancy. Of the four default-ON
flags, ONE has a governing ADR (deduction_serving_enabled / ADR-0256) and
TWO have no recorded reason of any kind — allow_cross_language_recall and
use_salience carry no comment block, no ADR, no criterion. In an
architecture built on earned licenses, two permanently-on capability flags
with no recorded decision is G-8's question inverted. Registered and
deliberately NOT fixed: writing a rationale after the fact would invent a
decision nobody made, which is worse than recording the absence.
2. A hollow gate inside the governance pin itself.
test_default_on_flag_is_not_governed_by_a_proposed_adr is parameterized over
(default-ON flag x cited ADR), so the three uncited ON flags contribute ZERO
cases — it covered exactly one flag. Its non-vacuity guard asserted only
that SOME flag cites an ADR, which default-off flags satisfy in abundance.
One reformatted comment away from zero coverage with everything green: the
failure state indistinguishable from the success state, inside the module
written to prevent exactly that. Guard tightened to require a non-empty
parametrization, and observed red by stripping ADR-0256's citation.
3. The daemon's own tests ran on no gate. R-3 changes what the always-on
process does every beat, and tests/test_l10_always_on_daemon.py was in no
curated suite — so this behaviour change would have shipped guarded solely
by tests no gate invokes, which is the hollow-gate pattern landing on the
exact file covering the change. Promoted to smoke WITH the change it guards.
Measured +9.4s, +4.3% of smoke — paid deliberately, not estimated. Removed
from full_only_baseline.txt in the same edit (both-directions ratchet; the
baseline shrank 747 -> 746).
Red before green throughout: the assertion cfg.accrue_realized_knowledge is
True was added to the daemon test and OBSERVED FAILING before the flag existed.
H-8(c) closed by AMENDING the correction note rather than deleting it, and the
amendment is the more useful artifact. PR-1's note said the 07-25 document's
claim was false; R-3 then made that claim TRUE, so the correction note itself
went stale — in the direction nobody guards, a correction made wrong by its own
subject being fixed. Both notes now carry the commit window they apply to. A
dated correction survives its subject changing; an undated "this is false" does
not.
Also delivered, from the 2026-07-28 external-assessment triage: §5, the
declared-table index. Eight single-source-of-truth tables and the pin that makes
each true. A reader's aid, explicitly NOT a central contracts.toml — that would
add a fifth copy needing agreement with four generators and put four different
authorities in one merge surface.
Handed forward to PR-11: the 5000-beat soak recorded in
evals/l10_always_on/contract.md (2026-07-19) predates this change and therefore
describes a configuration that NO LONGER SHIPS. PR-11's re-run is the first soak
of the corrected profile, and Step B feeding Step D is exactly what a long
horizon stresses.
Closes G-6, G-8, H-6, H-8(c), H-8(d) — and H-8 in full. Wave 4's F-6 half-gate
is lifted; the Track C half is untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
Ruled: fix, don't document. The argument that settled it — a red test on a
serving-path contract left with an explanatory note IS the H-8 failure mode, a
document acknowledging a gap instead of closing it, which becomes the next
assessor's finding.
Two distinct failures, and the second is the serious one.
1. THE CEREMONY PINS WERE STALE AGAINST CORRECT CODE (H-8e's class).
ChainRecord.polarity defaults to AFFIRMATIVE and is DELIBERATELY OMITTED
from affirmative rows (ADR-0264 R1; teaching/ratification.py:119). Both pins
built ChainRecord(**{k: row[k] for k in __slots__}), so they KeyError'd on
every correctly-serialized row. Now built from the row's own keys, letting
absent fields take their dataclass defaults — which RESTORES the guarantee
rather than weakening it. Consequence of the staleness: ratification-corpus
byte compatibility and unreviewed-status refusal were UNVERIFIED since
ADR-0264 R1 landed.
2. CLAIMS.md PUBLISHED A SUPERSEDED EVIDENCE DIGEST FOR A LICENSED CAPABILITY.
f9e9cc0c (2026-07-26, "the deduction lane hashes the prose it serves")
strengthened the lane, rewrote evals/deduction_serve/report.json, and moved
the authoritative pin in scripts/verify_lane_shas.py:60 to c855d55c —
recording the superseded 0b461a5a on line 56. It did NOT regenerate
CLAIMS.md. So for two days the published claim for deduction_serve_v1
pointed at evidence that no longer existed. The claim TEXT never changed;
only the pointer was wrong — which is exactly the failure a digest exists to
prevent. Regenerated; the diff is one line, the digest.
The drift had reached a THIRD artifact: this assessment's own
M3-comprehension-reasoning.md card cited the same stale 0b461a5a as its
would-fail-if-absent evidence. Corrected.
THE MECHANISM IS G-7's, AGAIN. One commit updated the lane, its verifier and
its lane test, and missed one downstream artifact — and the pin that guards
that artifact (tests/test_claims_md_is_current.py) EXISTED AND RAN ON NO GATE.
Nothing surfaced it until this arc ran the full tree once. Both it and
test_ratification_ceremony.py are now registered in TEST_SUITES["smoke"].
The membership ratchet caught the author a second time: promoting
test_claims_md_is_current.py onto the gate without removing its baseline line
failed test_baseline_has_no_stale_entries immediately. Baseline 748 -> 747.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
PR-6 promoted tests/test_safety_pack.py onto the gate and left it listed in
tests/full_only_baseline.txt. test_baseline_has_no_stale_entries failed the very
next smoke run, naming the file and the direction.
That is the both-directions half of the ratchet doing exactly the work it was
written for, roughly an hour after it landed, against the person who wrote it.
A baseline that silently keeps entries it no longer needs stops being a
measurement and becomes a story about one — so the failure is the feature.
Recorded in the pin's own docstring and in G-7, because a mechanism catching its
author is better evidence than a sabotage test: the sabotage was contrived, this
was not.
Baseline: 749 -> 748. The ratchet only turns down.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
N-9 is the finding, and it PREVENTS work rather than re-scoping it.
WHAT THE PLAN GOT WRONG. N-3 concluded the 23-vs-13 smoke delta was unintended
drift, on two in-repo comments asserting parity, and called that "the strongest
evidence available that the drift was never intended". H-12 repeated it; R-14
built three options on it recommending "raise CI, +46s measured"; PR-4
specified a bidirectional parity pin as deliverable one.
A third artifact answers all of it, and none of them read it.
tests/test_cli_test_suites.py::test_cli_smoke_suite_covers_ci_smoke_gate has
pinned smoke.yml SUBSET-OF TEST_SUITES["smoke"] since before this assessment —
which is why every measurement finds CI-only == 0; a pin, not luck — beneath a
comment headed "DELIBERATELY ONE-DIRECTIONAL — do not 'complete' this by also
asserting local_paths <= ci_paths". Its history is 50fa287d (2026-07-25):
"revert(tests): drop the local-to-CI smoke parity assertion — CI is not a
gate." An earlier agent read the same comment, drew the same inference, made
the assertion symmetric, and reverted it hours later: "That was wrong, and
AGENTS.md says so in a line I had already read."
AGENTS.md:280 — GitHub is "a mirror only; its Actions are billing-locked and
produce dead signals — never chase them." §277 — the workflows are "secondary
observability only." So the local suite is the SOURCE and is deliberately
broader; the delta is the design.
- PR-4 pin 1: WITHDRAWN. It exists, in the direction that protects something.
- R-14: premise corrected in the packet. Option A spends 46s on a workflow
that does not gate; option B would delete real protection for a fiction;
option C's live half is a two-line comment fix, done here.
- N-3's exposure claim was too GENEROUS to CI: for a push that skipped the
local gate there is no automatic check at all. Larger than stated, and not
closable by editing smoke.yml.
WHAT LANDED. Measured: 749 of 877 test files are in no curated suite. The
plan's mechanism ("assign every orphan, or 749 exclusion reasons") is hollow at
that scale — glob topic-suites ("adr": tests/test_adr_*.py absorbs 172) satisfy
the assertion while changing nothing about what executes, which is precisely the
Third-Door objection the plan levels at G-7's own first formulation. All four
demonstrated incidents (#113, #136, negation-survives-articulation,
speculative-subject-lifecycle) were caused by a NEWLY LANDED orphan; none by a
legacy one.
So: a ratchet. tests/test_suite_membership.py + tests/full_only_baseline.txt.
New orphans fail. The baseline is enforced in BOTH directions — a promoted or
deleted file must leave it, so the list can only shrink — and its count is
pinned so bulk movement is a reviewed decision, in the discipline
test_volume_honesty.py uses.
Also: the gate-guarding pin now runs on the gate. test_cli_test_suites.py lived
in `fast`; the pre-push gate runs smoke + deductive. The pin guarding the gate
did not run on the gate.
Three sabotages, each OBSERVED RED: a new unregistered test file, a stale
baseline entry, the parity pin demoted out of smoke.
Note on the delta: now 12 (local 27, CI 15) because this session added three
pins locally. Under the corrected reading that is the design, not widening
drift — the invariant is CI-only == 0, and that is pinned. The registers now
say so; "ten files" was a measurement, never a target.
Process note: `git checkout --` during sabotage cleanup destroyed uncommitted
edits to core/cli_test.py, which were redone. Backups, not checkout, on files
with unstaged work.
Registers: G-7 CLOSED, H-12 amended, R-14 dissolved in the packet, PR-4
re-specified, N-9 added to plan §0.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4