Merge the Wave 0 record amendments (PR-1) into the assessment docs branch

This commit is contained in:
Claude 2026-07-28 02:55:30 +00:00
commit 0c11b53083
No known key found for this signature in database
6 changed files with 76 additions and 29 deletions

View file

@ -2,7 +2,9 @@
**Version:** 0.1.0 **Version:** 0.1.0
**Date:** 2026-05-12 **Date:** 2026-05-12
**Status:** Draft — Under Active Development **Status:** Draft — **stale since 2026-05-12; not under active development**
> **Staleness banner — 2026-07-27 (holistic assessment Phase 6).** This blueprint has not moved since its authoring date and was one of CORE's five unreconciled macro articulations of its own cognitive cycle. Those five were dissolved into the two-axis taxonomy in `docs/assessment/02-layer-taxonomy.md` (decision D1) — read that first for the current structural account. What this document uniquely contributed, the allocation/attention physics layer (ADR-0008), was verified **live** in Phase 2 (`generate/stream.py:255-263`, `use_salience` defaults `True`) and is ungoverned rather than absent; it is registered as G-14 awaiting a one-page ADR. Nothing here is authoritative; it is kept for its intent.
--- ---

View file

@ -3,6 +3,7 @@
**Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27) **Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27)
**Standing:** This is CORE's first *live* gap register since `docs/gaps.md` closed its 26th entry. Proposal (for ruling): this register supersedes `docs/gaps.md`, which is marked historical; two dead registers plus a live one is worse than one live one. **Standing:** This is CORE's first *live* gap register since `docs/gaps.md` closed its 26th entry. Proposal (for ruling): this register supersedes `docs/gaps.md`, which is marked historical; two dead registers plus a live one is worse than one live one.
**Discipline:** A gap is an *absence the telos requires filled* with no explicit deferral ruling. Deferred-with-ruling is not a gap (scripture content is the model). Every entry carries evidence, its **deciding authority**, and a leverage rank. The register decides nothing. **Discipline:** A gap is an *absence the telos requires filled* with no explicit deferral ruling. Deferred-with-ruling is not a gap (scripture content is the model). Every entry carries evidence, its **deciding authority**, and a leverage rank. The register decides nothing.
**Amended:** 2026-07-27 at `ed06dd64` (Opus 5, Phase 6). Four entries carried claims that code, workflows, and an ADR's own supersession banner falsify. Each amendment is marked **[AMENDED N-x]** inline and derived in [`50-execution-plan.md`](50-execution-plan.md) §0: **G-5** (the soak ran and passed; the pins run twice a day), **G-7** (the orphan scan was an artifact; the mechanism is respecified), **G-8** (28 flags, not 17), **G-10** (the engineering blocker was discharged 2026-07-26). G-19 gains a note: its exposure is already pinned in-repo.
--- ---
@ -28,10 +29,19 @@ The comprehension frontier itself, measured. Standing ruling: close fabrications
Confirmed at component depth: drive objects exist (`DriveGradientMap` — constructed, never read; `ExertionMeter` — telemetry only); idle mechanisms exist (consolidation, proposal review, contemplation — each flag-gated, each doing one thing); **nothing ranks what matters next**. The daemon heartbeat advances `idle_tick` and nothing more ambitious. This is the AGI-grade conceptual absence: everything CORE does is chosen by the operator. Design work, not a flag flip. Confirmed at component depth: drive objects exist (`DriveGradientMap` — constructed, never read; `ExertionMeter` — telemetry only); idle mechanisms exist (consolidation, proposal review, contemplation — each flag-gated, each doing one thing); **nothing ranks what matters next**. The daemon heartbeat advances `idle_tick` and nothing more ambitious. This is the AGI-grade conceptual absence: everything CORE does is chosen by the operator. Design work, not a flag flip.
**Evidence:** `attention-allocation` + `always-on-process` cards; `02-layer-taxonomy.md` CR-2. · **Authority:** design + ruling (does the L10 process own an agenda, governed by what). **Evidence:** `attention-allocation` + `always-on-process` cards; `02-layer-taxonomy.md` CR-2. · **Authority:** design + ruling (does the L10 process own an agenda, governed by what).
### G-5 · L10 proof debt — the soak has never produced an artifact, and nothing runs its pins ### G-5 · L10 proof debt — the soak ran and passed; what is owed is the artifact, the pinned digest, and a cadence **[AMENDED N-1/N-2/N-4]**
**Layer:** M6 / MV · **Leverage: 5** **Layer:** M6 / MV · **Leverage: 5**
The always-on process is built; the falsifiable harness (H1H4, holds/bites pairs, vacuity-guarded) is built; **no recorded long-horizon artifact exists, no suite contains any `l10`/`always_on` test, no nightly cadence exists** — and the local-first/Mac-runner doctrine makes "nightly" itself need a ruling rather than a cron line. The still-owed ADR-0146 Phase-4 spike, in its modern form: run the soak, record the artifact, schedule the pins. The always-on process is built; the falsifiable harness (H1H4, holds/bites pairs, vacuity-guarded) is built; **and it was run.** `evals/l10_always_on/contract.md` §"The measured result" records a **5000-beat soak with a reboot at beat 2500** (landed 2026-07-19 in `aed273b1`) in which all four predicates pass: `versor_condition` flat at `1.389e-07` across all 5000 beats, vault bounded at 6 entries, convergence at beat 1 with the 4999-beat tail at rest, reboot resuming the same life with derived learning intact.
**Evidence:** `M6` + `always-on-process` cards; suite-membership scan. · **Authority:** execution + MV suite ruling + cadence ruling.
What remains owed is the **ceremony**, and it is this assessment's own §8 maintenance contract operating on this assessment's own frontier:
- the result is **prose in a contract file** — no committed machine-readable report, no run SHA, no rerun path;
- `deterministic_digest` is computed by `report.py` and **pinned nowhere**, though the contract's closing line says *"Pin it once the lane is trusted so a regression flips it"*;
- **no cadence** rules the lane, and the local-first/Mac-runner doctrine makes "nightly" itself a ruling rather than a cron line.
A recorded prose result with no pinned digest is **testimony, not evidence** — the same failure mode as the map, the ratchet, and the blueprint.
**Two prior claims in this entry were wrong and are withdrawn.** (a) *"No suite contains any `l10`/`always_on` test"* was a scan artifact: `TEST_SUITES["full"] = ("tests/",)` is a **directory**, so every test file is trivially in a suite. The true statement is that the five `tests/test_l10_*.py` files are in **no curated suite tuple** — reachable only through `full`. (b) *"nothing runs its pins"* is false: CI never invokes `core test --suite`, it runs raw pytest with marker filters, and no L10 file carries `quarantine` or `slow` — so the pins **execute twice a day**, in `full-pytest.yml` (post-merge) and `nightly-full-pytest.yml` (cron `0 2 * * *`). The correct statement is that **nothing runs them on any pre-merge gate**.
**Evidence:** `M6` + `always-on-process` cards; `evals/l10_always_on/contract.md`; `core/cli_test.py:13`; `.github/workflows/{smoke,full-pytest,nightly-full-pytest}.yml`. · **Authority:** execution + the R-9 evidence-standard and cadence ruling.
### G-6 · F-6 — the lived learning loop is half-gated ### G-6 · F-6 — the lived learning loop is half-gated
**Layer:** M6/M5 · **Leverage: 6** **Layer:** M6/M5 · **Leverage: 6**
@ -42,25 +52,37 @@ The daemon forces `consolidate_determinations` but not `accrue_realized_knowledg
## Tier B — Enforcement & instrument debt (capability exists; the guarantee doesn't) ## Tier B — Enforcement & instrument debt (capability exists; the guarantee doesn't)
### G-7 · No orphaned-pin meta-check ### G-7 · No gate-parity pin and no *curated*-suite membership check **[AMENDED N-1/N-3]**
**Layer:** MV · **Leverage: 7** **Layer:** MV · **Leverage: 7** — **the highest-leverage single mechanical change in the repository**
Suite tuples are hand-curated; a test file in zero suites is indistinguishable from one that runs everywhere. This is the *mechanism* by which G-5 happened. A meta-pin — every `tests/**/*.py` belongs to ≥1 suite or an explicit exclusion list — converts the doctrine "a pin in no suite never runs" into a failing test. Likely the highest-leverage *single mechanical change* in the repository. Suite tuples are hand-curated; a test file in zero *curated* suites is indistinguishable from one that runs everywhere.
**Evidence:** `MV` card; the M6 case as the demonstration. · **Authority:** mechanical (small PR); no ruling needed.
### G-8 · No flag-default register **The mechanism this entry originally proposed would not have worked.** "Every `tests/**/*.py` belongs to ≥1 suite or an explicit exclusion list" is **already satisfied**, because `TEST_SUITES["full"] = ("tests/",)` is a directory — the pin would ship green and prove nothing. A hollow gate by the Third-Door criterion, in the entry proposing to abolish hollow gates.
**The correct mechanism is two pins:**
1. **Gate parity**`smoke.yml`'s path set must equal `TEST_SUITES["smoke"]`, parsed from the workflow file, failing on drift in either direction.
2. **Curated-suite membership** — every `tests/**/test_*.py` is in ≥1 suite **excluding `full`**, or in a registered exclusion list with a one-line reason.
Pin 1 exists because the two gates have **already** drifted (see H-12): the local suite is 23 files, `smoke.yml` is 13, and the 10-file delta includes ADR-0265's denial pin, both ADR-governance pins, volume honesty, and curriculum polarity. Measured parity cost: **429 tests in 46s** on the Act runner's own hardware.
**Evidence:** `MV` card; `core/cli_test.py:13`; `.github/workflows/smoke.yml`; H-12. · **Authority:** mechanical (small PR) for the pins; the R-14 ruling for the parity *direction*.
### G-8 · No flag-default register **[AMENDED N-7]**
**Layer:** cross-cut · **Leverage: 8** **Layer:** cross-cut · **Leverage: 8**
Seventeen capability flags default `False`; one is ratified ON (`deduction_serving_enabled`); three are daemon-forced (`persist_session_state`, `consolidate_determinations`, `strict_identity_continuity`). No document states the set, which defaults are deliberate posture vs accumulated hesitancy, or what evidence would flip each. The largest lever in the system, unregistered. The register format already exists in-repo: the ratified-ledger pattern (declare absence policy in the table, not the call site — ADR-0263 Rule 5). `RuntimeConfig` is a single frozen dataclass with **32 boolean fields: 28 default `False`, 4 default `True`** (`allow_cross_language_recall`, `use_salience`, `discourse_planner`, `deduction_serving_enabled` — the last ratified ON by ADR-0256). Three are daemon-forced (`persist_session_state`, `consolidate_determinations`, `strict_identity_continuity`). No document states the set, which defaults are deliberate posture vs accumulated hesitancy, or what evidence would flip each. The largest lever in the system, unregistered. The register format already exists in-repo: the ratified-ledger pattern (declare absence policy in the table, not the call site — ADR-0263 Rule 5).
**Evidence:** `core/config.py` scan (Phase 2); daemon trio (Phase 3). · **Authority:** documentation PR + per-flag evidence bars set by ruling.
*(This entry originally said "seventeen." The verified count is 28 default-off. That two counts of the same set differ by eleven is itself the finding: nothing in the repository distinguishes a capability flag from a policy or deployment flag, which is exactly the classification the register must make.)*
**Evidence:** `core/config.py` at `ed06dd64` (32 `bool` fields, 28 `= False`, 4 `= True`); daemon trio (Phase 3). · **Authority:** documentation PR + per-flag evidence bars set by ruling (R-4).
### G-9 · Enforcement pins unverified for three doctrine-level prohibitions ### G-9 · Enforcement pins unverified for three doctrine-level prohibitions
**Layer:** M1 / MG · **Leverage: 9** **Layer:** M1 / MG · **Leverage: 9**
(a) No verified failing pin for the no-approximate-recall law (would a cosine ranker actually fail a test?); (b) no pin that fails when a layer *bypasses* governance entirely (as distinct from governance working when called); (c) safety-pack non-swappability not verified as mechanically enforced. All three are law in `AGENTS.md`; law-enforced-by-review is weaker than law-enforced-by-test. (a) No verified failing pin for the no-approximate-recall law (would a cosine ranker actually fail a test?); (b) no pin that fails when a layer *bypasses* governance entirely (as distinct from governance working when called); (c) safety-pack non-swappability not verified as mechanically enforced. All three are law in `AGENTS.md`; law-enforced-by-review is weaker than law-enforced-by-test.
**Evidence:** `M1`/`MG` cards (flagged, not resolved, in Phase 23). · **Authority:** verification pass, then mechanical PRs. **Evidence:** `M1`/`MG` cards (flagged, not resolved, in Phase 23). · **Authority:** verification pass, then mechanical PRs.
### G-10 · Curriculum SERVE is fully blocked by one engineering item, and its ledger doesn't exist ### G-10 · Curriculum SERVE is blocked by **one ruling**, and its ledger doesn't exist **[AMENDED N-5]**
**Layer:** M5 · **Leverage: 10** **Layer:** M5 · **Leverage: 10**
ADR-0264 §4.1: the 16-premise compilation cap holds every band to ≤16 entailed cases, so **no curriculum band can earn SERVE until query-scoping lands** — an engineering blocker gating a content problem that is itself quantified at 24×73× under-fed. Downstream, `chat/data/curriculum_serve_ledger.json` is absent (the one honest `missing_ok=True` in production), and a committed ledger is necessarily an *earning* one — the outcome-mix ruling remains the binding constraint. **The engineering blocker this entry named does not exist.** ADR-0264 §4.1 carries a banner directly beneath its own heading: *"**SELF-SUPERSEDED by this ADR's own R5, discharged 2026-07-26. The heading is no longer true of the running system.** … R5 removed that: compilation is query-scoped, so a family of any size answers. With the cap gone, **four bands would earn SERVE the moment a ledger is sealed**`physics·causal`, `systems_software·causal`, and `philosophy_theology·{modal,contrast}`."* Query-scoping already landed; this entry quoted a superseded heading past its own correction.
**Evidence:** `M5` card; ADR-0264 §4.1; `chat/curriculum_serve_license.py:46`. · **Authority:** engineering (scoping) + outcome-mix ruling.
What survives, and is now the **whole** of the blockage: reliability is commitment precision and a correct UNKNOWN *is* a commitment, so a band clears θ_SERVE **on non-commitments alone** (`conservative_floor(660,660) = 0.990046`). The licensable evidence is 99.099.98% non-entailed and **max entailed volume in any band is 9**. `chat/data/curriculum_serve_ledger.json` is deliberately absent (the one honest `missing_ok=True` in production) and `core proposal-queue reseal` refuses a license without `--allow-new-licenses`, so nothing is licensed while the outcome-mix ruling is unmade. A committed ledger is necessarily an *earning* one.
**Evidence:** `M5` card; ADR-0264 §4.1 supersession banner + §5 (open); `docs/research/curriculum-practice-producer-2026-07-26.md` §1; `chat/curriculum_serve_license.py:40-52`. · **Authority:** the outcome-mix ruling (R-8) **alone** — no engineering prerequisite remains.
### G-11 · Identity enforcement has no stated authorization bar ### G-11 · Identity enforcement has no stated authorization bar
**Layer:** MG · **Leverage: 11** **Layer:** MG · **Leverage: 11**
@ -90,7 +112,7 @@ Own `use_salience`, the two underived constants, the self-narrowing budget feedb
- **G-16 · ADR-0265's defect class survives in `_inflect_predicate`'s aspect arms** (`generate/templates.py:79`) — 10,530/16,146 template points, *not reachable today*. Latent, recorded from the prior arc; becomes live if aspect arms become reachable. **Authority:** the widening program (G-3) must clear it first. - **G-16 · ADR-0265's defect class survives in `_inflect_predicate`'s aspect arms** (`generate/templates.py:79`) — 10,530/16,146 template points, *not reachable today*. Latent, recorded from the prior arc; becomes live if aspect arms become reachable. **Authority:** the widening program (G-3) must clear it first.
- **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling. - **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling.
- **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR. - **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR.
- **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Authority:** ADR amendment + re-count. - **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Note (Phase 6):** the exposure is **already pinned in-repo**`tests/test_volume_honesty.py` (ADR-0264 R9, in the local smoke gate) pins the 21-of-25 shortfall *"in BOTH directions"* and calls its inventory *"an EXPOSURE INVENTORY, not an approved baseline."* So the open work is **applying** the demotions, not discovering them, and that pin moves in the same PR. **Authority:** ADR amendment + re-count, authorized by R-13.
- **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain). - **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain).
--- ---

View file

@ -5,6 +5,8 @@
Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS.md protocol — not by ease. Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS.md protocol — not by ease.
**Amended:** 2026-07-27 at `ed06dd64` (Opus 5, Phase 6). **H-12 added** (the two smoke gates have diverged). **H-8 gains a fourth instance**, located inside the source rather than the documents. **H-1 gains a note**: its exposure is already pinned. Derivations in [`50-execution-plan.md`](50-execution-plan.md) §0.
--- ---
## H-1 · License evidence counted on an independence assumption replay violates ## H-1 · License evidence counted on an independence assumption replay violates
@ -12,7 +14,8 @@ Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS
**Evidence:** Wilson lower-bound licensing (θ_SERVE=0.99, ADR-0175 lineage) assumes independent trials; a replay of the same sealed case is one trial observed again, not a new one. Measured consequence recorded in the prior arc: **21 of 25 ratified bands fall short** of their floor when replays are deduplicated. **Evidence:** Wilson lower-bound licensing (θ_SERVE=0.99, ADR-0175 lineage) assumes independent trials; a replay of the same sealed case is one trial observed again, not a new one. Measured consequence recorded in the prior arc: **21 of 25 ratified bands fall short** of their floor when replays are deduplicated.
**Why it hinders:** the entire earned-license architecture — CORE's mechanism for *deserving* to serve — rests on the evidence count. An overstated count grants licenses the evidence doesn't support, which is precisely the failure the mechanism exists to prevent. The gate is right; the arithmetic feeding it is not. **Why it hinders:** the entire earned-license architecture — CORE's mechanism for *deserving* to serve — rests on the evidence count. An overstated count grants licenses the evidence doesn't support, which is precisely the failure the mechanism exists to prevent. The gate is right; the arithmetic feeding it is not.
**Better home:** distinct-evidence counting at the seal boundary (count distinct cases; a replay refreshes, never increments), declared in the ledger schema the way ADR-0263 Rule 5 declares absence policy — in the table, not the call site. **Better home:** distinct-evidence counting at the seal boundary (count distinct cases; a replay refreshes, never increments), declared in the ledger schema the way ADR-0263 Rule 5 declares absence policy — in the table, not the call site.
**Authority:** ADR amendment (0175/0263 lineage) + a re-count of the 25 bands. The re-count may demote licenses; that is the mechanism working. **Note (Phase 6):** the exposure is **already pinned in-repo**. `tests/test_volume_honesty.py` (ADR-0264 R9, in the local smoke gate) pins *"21 of 25 bands do not clear θ_SERVE on distinct evidence"* **in both directions**, measured 2026-07-25 at `6ada6f7a` with every band at 720 committed decisions, and its own comment calls the inventory *"an EXPOSURE INVENTORY, not an approved baseline."* The audit source is `docs/research/distinct-evidence-audit-2026-07-25.md`. The open work is therefore **applying** the shortfall, not discovering it — and that pin moves in the same PR.
**Authority:** ADR amendment (0175/0263 lineage) + a re-count of the 25 bands, authorized by R-13. The re-count may demote licenses; that is the mechanism working.
## H-2 · Decoration in the runtime constructor — objects built and never read ## H-2 · Decoration in the runtime constructor — objects built and never read
**Verdict:** decoration (fails the sabotage test) · **Layers:** M6/M3 **Verdict:** decoration (fails the sabotage test) · **Layers:** M6/M3
@ -58,10 +61,10 @@ Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS
## H-8 · The record contradicts the code at three load-bearing points ## H-8 · The record contradicts the code at three load-bearing points
**Verdict:** `wrong-solution` as *record-keeping* — divergence that reasoners inherit · **Layers:** governance **Verdict:** `wrong-solution` as *record-keeping* — divergence that reasoners inherit · **Layers:** governance
**Evidence:** (a) ADR-0146 rejects the daemon shape; an unowned daemon ships. (b) ADR-0252's headline "34 organs" has no reproducible basis (18 entry organs at the ratification commit itself; ~32 modules). (c) `architecture-assessment-verification-2026-07-25.md` asserts accrual "is enabled by the production L10 process"; the flag set says otherwise. **Evidence:** (a) ADR-0146 rejects the daemon shape and places *"cross-process file locking, daemon synchronization, and signal handling"* out of scope; `chat/always_on_daemon.py` ships all three (`fcntl.flock` single-instance lock, SIGINT/SIGTERM, load-time identity guard) and is unowned. (b) ADR-0252's headline "34 organs" has no reproducible basis (18 `resolve_promotable_*` entry organs at the ratification commit *and* at `ed06dd64`; ~32 modules). (c) `architecture-assessment-verification-2026-07-25.md` asserts accrual "is enabled by the production L10 process"; the flag set says otherwise. **(d) — added Phase 6, and it is inside the code.** `core/config.py`'s docstring for `accrue_realized_knowledge` states *"the production L10 process enables it alongside `persist_session_state`"*, and the docstring for `consolidate_determinations` states the same about *"`accrue_realized_knowledge` + `persist_session_state`"*. `CONTINUOUS_LIFE_CONFIG_FLAGS` (`chat/always_on_daemon.py:45-49`) contains neither claim's subject: it forces `persist_session_state`, `consolidate_determinations`, `strict_identity_continuity`, and **not** `accrue_realized_knowledge`.
**Why it hinders:** demonstrated, not hypothetical — this assessment's own Phase 0 inherited a stale-record error, and the 2026-07-25 doc (itself a *corrective* document) introduced one. Every divergence is a future wrong analysis. **Why it hinders:** demonstrated, not hypothetical — this assessment's own Phase 0 inherited a stale-record error, and the 2026-07-25 doc (itself a *corrective* document) introduced one. Every divergence is a future wrong analysis. Instance (d) is the sharpest: it is one layer *below* the documents, where a reader checking the code against the docs would stop and believe.
**Better home:** three one-paragraph amendments (ADR-0146 addendum owning the daemon or superseding the rejection; ADR-0252 basis sentence; a correction note on the 07-25 doc). **Better home:** four amendments — ADR-0146 addendum owning the daemon (drafted in `50-rulings.md` R-12a); ADR-0252 basis footnote (drafted, R-12b); a correction note on the 07-25 doc; and the two docstrings corrected or the flag added, per R-3.
**Authority:** docs PRs + ruling signatures. **Authority:** docs PRs + ruling signatures (R-12 for the two ratified ADRs, R-3 for the docstrings).
## H-9 · Dead instruments still standing as if live ## H-9 · Dead instruments still standing as if live
**Verdict:** `superseded-in-place` (unratified) · **Layers:** MV/governance **Verdict:** `superseded-in-place` (unratified) · **Layers:** MV/governance
@ -83,6 +86,14 @@ Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS
**Better home:** count the swallow (a telemetry field on `IdleTickResult`/turn accrual), not a behavior change. **Better home:** count the swallow (a telemetry field on `IdleTickResult`/turn accrual), not a behavior change.
**Authority:** mechanical PR. **Authority:** mechanical PR.
## H-12 · The two smoke gates have diverged by ten files, and one comment says they cannot
**Verdict:** `strained` — a real asymmetry, *smaller* than it first looks · **Layers:** MV
**Evidence:** `TEST_SUITES["smoke"]` is **23** files; `.github/workflows/smoke.yml` is 8 path patterns expanding to **13**. The ten local-only files are `test_audit_ledger_r7`, `test_cli_runner_contract`, `test_pack_draft_serve_boundary`, `test_workbench_deduction_provenance`, `test_prior_surface_deduction_binding`, `test_negation_survives_articulation`, `test_adr_status_governance`, `test_adr_index`, `test_volume_honesty`, `test_curriculum_polarity`.
**State the counter-evidence first.** AGENTS.md §CI/CD is explicit that `.github/workflows/*.yml` are *"secondary observability only — never a substitute for local gates"*: the merge gate **is** the in-worktree run, and the in-worktree run is the **superset**. Eight of the ten files carry in-code comments saying they belong *"on the pre-push gate"* — which is exactly where they are. **Under doctrine, nothing is unguarded**, and any reading of this entry as "ten pins run nowhere" is wrong.
**Why it hinders, precisely:** (1) **two independent places assert the parity** — the audio block in `core/cli_test.py` says it is listed explicitly *"so the local-first pre-push gate (AGENTS.md protocol) **equals** the CI gate rather than silently narrowing it,"* accurate about the six audio files and false as the statement about the suite a reader will take it for; and `scripts/hooks/pre-push`, the *automation of the AGENTS.md protocol itself*, describes its own step 1 as *"the `smoke` suite — **exact CI-gate parity**"* while running 23 files against CI's 13. A claim made twice inside the enforcement tooling is the strongest available evidence that the drift was never intended; (2) the real exposure is the **push that skipped the local gate** — from a cloud session, another machine, or an agent — for which CI is the only automatic check, and for which ADR-0265's denial pin, both ADR-governance pins, volume honesty, and curriculum polarity currently run nowhere before merge. Six of these files were promoted *after* real silent-regression incidents (#136, #113, the 2026-07-20..24 register-axis drift), so the failure mode they exist to catch is demonstrated, not hypothetical.
**Better home:** the gate-parity pin in G-7 — the drift is only invisible because nothing measures it. Measured parity cost: **429 tests in 46s** on the Act runner's own hardware (`ubuntu-latest:host` = native macOS host).
**Authority:** R-14 sets the direction (raise CI, or lower local, or keep them different and correct the comment); the pin itself is mechanical.
--- ---
## Explicitly examined and cleared ## Explicitly examined and cleared

View file

@ -2,6 +2,7 @@
**Phase 5 synthesis · Fable 5 · 2026-07-27 · verified at `forgejo/main` @ `8927c563`** **Phase 5 synthesis · Fable 5 · 2026-07-27 · verified at `forgejo/main` @ `8927c563`**
**Method:** `docs/conceptualizing_engineering_mastery.md`, applied per `00-scope-and-method.md`. Everything below is traceable to a card or register entry; nothing below is new evidence. **Method:** `docs/conceptualizing_engineering_mastery.md`, applied per `00-scope-and-method.md`. Everything below is traceable to a card or register entry; nothing below is new evidence.
**Amended:** 2026-07-27 at `ed06dd64` (Opus 5, Phase 6) — sizing the execution plan produced seven corrections, four of which touch statements made here. They are applied inline and marked **[AMENDED N-x]**; the derivations are in [`50-execution-plan.md`](50-execution-plan.md) §0. The chain in §8 continues: *nothing in it was ever caught by re-reading documents.*
--- ---
@ -22,7 +23,7 @@ From the evidence-bearing stage-coverage audit (Phase 2, corrected by Phase 3):
| **articulate** | covered | Selection-not-rewrite, disclosed estimates, typed refusal — served as an empty string (H-3). | | **articulate** | covered | Selection-not-rewrite, disclosed estimates, typed refusal — served as an empty string (H-3). |
| **learn** | covered, throttled | The single reviewed path is proven by pinned lanes; volume is 24×73× under the floor and every curriculum band is capped at 16 entailed cases by one engineering item. | | **learn** | covered, throttled | The single reviewed path is proven by pinned lanes; volume is 24×73× under the floor and every curriculum band is capped at 16 entailed cases by one engineering item. |
| **replay** | covered | Eleven SHA-pinned lanes failing CI on drift; the strongest sustained discipline in the repository. | | **replay** | covered | Eleven SHA-pinned lanes failing CI on drift; the strongest sustained discipline in the repository. |
| ***the runner of the cycle*** | **built, unproven** | The continuous life exists as code and has never been observed living longer than a test. Its learning loop is half-gated (F-6). Its guardian tests are orphaned (G-5/G-7). | | ***the runner of the cycle*** | **built, evidenced as prose** | The continuous life exists as code and **has been observed over 5000 idle beats with a reboot at 2500, all four falsifiable gates passing** — recorded as prose in a contract file, with no committed artifact and no pinned digest. Its learning loop is half-gated (F-6), and no pre-merge gate runs its pins. **[AMENDED N-4]** |
## 3. What is excellent — the standard the rest should be held to ## 3. What is excellent — the standard the rest should be held to
@ -41,8 +42,8 @@ The layer table (Phase 2, with Phase 3 corrections applied): **no layer is `wron
Three structural facts dominate the macro picture: Three structural facts dominate the macro picture:
- **Built-and-dark.** Seventeen capability flags default off; one is ratified on; three are daemon-forced. The gap between what CORE *is* and what CORE *does by default* is the widest gap in the system, and it is a governance artifact, not an engineering one (G-8). - **Built-and-dark.** **Twenty-eight** of `RuntimeConfig`'s 32 boolean flags default off; one of the four on is ratified on; three of the off ones are daemon-forced. The gap between what CORE *is* and what CORE *does by default* is the widest gap in the system, and it is a governance artifact, not an engineering one (G-8). *(This section originally said seventeen; two counts of one set differing by eleven is the finding — nothing distinguishes a capability flag from a policy flag.)* **[AMENDED N-7]**
- **Capability outruns proof.** The daemon, Shape B+ persistence, the soak harness, the SME scaffolding — all built; none carried to verdict. The pattern is consistent enough to be cultural: *this project finishes machinery and defers ceremonies.* The mastery framework's step 4 (accelerate cycle time) applies to evidence loops, not just build loops. - **Capability outruns proof — and the ceremony, not the run, is what is missing.** The daemon, Shape B+ persistence, the soak harness, the SME scaffolding — all built. The soak was even *run*, at 5000 beats, and passed; what it never got was a committed artifact and a pinned digest. The SME experiment was never run at all. The pattern is consistent enough to be cultural: *this project finishes machinery and defers ceremonies.* The mastery framework's step 4 (accelerate cycle time) applies to evidence loops, not just build loops. **[AMENDED N-4]**
- **The record decays faster than the code.** Five articulations, three record/code contradictions, two dead registers, one stale map, and an assessment (this one) that had to correct itself twice using the only method that works. The cure is not more documents — it is `verified_at` stamps, failing pins for laws, and instruments that supersede rather than accumulate (H-8, H-9, G-7, G-9). - **The record decays faster than the code.** Five articulations, three record/code contradictions, two dead registers, one stale map, and an assessment (this one) that had to correct itself twice using the only method that works. The cure is not more documents — it is `verified_at` stamps, failing pins for laws, and instruments that supersede rather than accumulate (H-8, H-9, G-7, G-9).
## 5. The five frontiers ## 5. The five frontiers
@ -52,8 +53,8 @@ Everything separating CORE-as-built from CORE-as-intended reduces to five named
1. **The reading** — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier. 1. **The reading** — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier.
2. **The verdict** — run ADR-0252 §5 (G-1). One experiment, already authorized, already scaffolded, NO-GO defined as full credit. It decides the *shape* of frontier 1 and retires or redeems the 18 condemned organs. Highest leverage in the project. 2. **The verdict** — run ADR-0252 §5 (G-1). One experiment, already authorized, already scaffolded, NO-GO defined as full credit. It decides the *shape* of frontier 1 and retires or redeems the 18 condemned organs. Highest leverage in the project.
3. **The chooser** — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention. 3. **The chooser** — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention.
4. **The proof of life**run the soak, record the artifact, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness, awaiting execution. 4. **The proof of life**the soak *passed* at 5000 beats; commit the artifact, pin the digest, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness and answered in prose that nothing can regress against. **[AMENDED N-4]**
5. **The throughput**curriculum query-scoping, the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). The learning engine is sound and starved; this frontier is volume with integrity. 5. **The throughput** — the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). Curriculum query-scoping **already landed** (ADR-0264 R5, discharged 2026-07-26), so this frontier has no engineering prerequisite left: four bands earn SERVE the moment one content-policy ruling is made. The learning engine is sound and starved; this frontier is volume with integrity. **[AMENDED N-5]**
## 6. The recommended attack order ## 6. The recommended attack order
@ -82,7 +83,7 @@ Soak cadence under a ruled schedule; flag profiles flipped per their registered
**What is the metadata?** The card schema (`03-card-schema.md`) — liveness ⊥ fitness, design ⊥ build, evidence with the would-fail-if-absent bit, capacity with ceilings, `verified_at` stamps — plus 17 filled cards and two registers. This directory is the instrument you asked for: each layer and component now has a philosophical intent, a functional contract, an implementation status with evidence, and a fitness judgment, in one greppable place that travels with the repository. **What is the metadata?** The card schema (`03-card-schema.md`) — liveness ⊥ fitness, design ⊥ build, evidence with the would-fail-if-absent bit, capacity with ceilings, `verified_at` stamps — plus 17 filled cards and two registers. This directory is the instrument you asked for: each layer and component now has a philosophical intent, a functional contract, an implementation status with evidence, and a fitness judgment, in one greppable place that travels with the repository.
**What is hindering us?** Eleven audited entries (H-1…H-11), each with evidence, better home, and authority — headlined by the license-counting basis, decoration-as-testimony, and record/code divergence — plus five candidates examined and *cleared*, so the audit's negative space is as deliberate as its findings. No ratified ADR was found to be a wrong decision; three were found to have wrong *records*. **What is hindering us?** Twelve audited entries (H-1…H-12), each with evidence, better home, and authority — headlined by the license-counting basis, decoration-as-testimony, and record/code divergence — plus five candidates examined and *cleared*, so the audit's negative space is as deliberate as its findings. No ratified ADR was found to be a wrong decision; three were found to have wrong *records*, and a fourth divergence was later found **inside the code** (H-8d). **[AMENDED N-3/N-6]**
## 8. Method, and what it earned ## 8. Method, and what it earned
@ -90,4 +91,6 @@ Four phases, three self-corrections, one direction: **Phase 0 trusted a map and
That is also the maintenance contract for this directory: a card whose `verified_at` falls behind a load-bearing arc is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint. The registers supersede the dead instruments only for as long as they are kept live. The cheapest way to keep them live is Wave 2's mechanical enforcement; the most expensive way is another assessment like this one. That is also the maintenance contract for this directory: a card whose `verified_at` falls behind a load-bearing arc is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint. The registers supersede the dead instruments only for as long as they are kept live. The cheapest way to keep them live is Wave 2's mechanical enforcement; the most expensive way is another assessment like this one.
**Phase 6 continued the chain, and it is worth naming what caught what.** Seven further corrections (`50-execution-plan.md` §0) came from reading `core/cli_test.py`, three workflow files, `core/config.py`, `evals/l10_always_on/contract.md`, and one ADR's own supersession banner. The banner case is the instructive one: ADR-0264 had already corrected itself, in place, directly under the superseded heading — and Phase 4 quoted the heading anyway. **A document that corrects itself only helps a reader who reads past the heading.** That is an argument for the failing pin over the amended paragraph, everywhere it is available.
*— End of assessment. All deliverable sets complete: scope/method, ground truth, taxonomy, schema, 9 layer cards, 8 component cards, both registers, this synthesis.* *— End of assessment. All deliverable sets complete: scope/method, ground truth, taxonomy, schema, 9 layer cards, 8 component cards, both registers, this synthesis.*

View file

@ -14,8 +14,8 @@ A read-only, evidence-bearing assessment of CORE's cognitive-cycle design versus
| [`04-phase2-findings.md`](04-phase2-findings.md) | 2 | Stage-coverage audit; corrections to Phase 0; findings F-1…F-5 | | [`04-phase2-findings.md`](04-phase2-findings.md) | 2 | Stage-coverage audit; corrections to Phase 0; findings F-1…F-5 |
| [`20-component-cards/`](20-component-cards/) | 3 | Eight component cards: the four zero-subsystem zones + always-on, derivation organs, surface selection, attention | | [`20-component-cards/`](20-component-cards/) | 3 | Eight component cards: the four zero-subsystem zones + always-on, derivation organs, surface selection, attention |
| [`05-phase3-findings.md`](05-phase3-findings.md) | 3 | Corrections C-1…C-5; findings F-6…F-10; the consolidated Phase-4 seed list | | [`05-phase3-findings.md`](05-phase3-findings.md) | 3 | Corrections C-1…C-5; findings F-6…F-10; the consolidated Phase-4 seed list |
| [`30-gap-register.md`](30-gap-register.md) | 4 | **The live gap register** — 20 entries, 4 tiers, each with evidence + deciding authority (proposes superseding `docs/gaps.md`) | | [`30-gap-register.md`](30-gap-register.md) | 4 | **The live gap register** — 20 entries, 4 tiers, each with evidence + deciding authority (proposes superseding `docs/gaps.md`); four entries amended in Phase 6 |
| [`31-hindrance-audit.md`](31-hindrance-audit.md) | 4 | Eleven hindrances with fitness verdicts and better homes; five candidates examined and cleared | | [`31-hindrance-audit.md`](31-hindrance-audit.md) | 4 | Twelve hindrances with fitness verdicts and better homes; five candidates examined and cleared |
| [`40-assessment.md`](40-assessment.md) | 5 | The synthesis | | [`40-assessment.md`](40-assessment.md) | 5 | The synthesis |
| [`50-execution-plan.md`](50-execution-plan.md) | 6 | **The execution plan** — five waves + five frontier tracks over every G/H entry, with the dependency gates and the risks. §0 carries seven corrections to the assessment found while sizing it | | [`50-execution-plan.md`](50-execution-plan.md) | 6 | **The execution plan** — five waves + five frontier tracks over every G/H entry, with the dependency gates and the risks. §0 carries seven corrections to the assessment found while sizing it |
| [`50-rulings.md`](50-rulings.md) | 6 | **The ruling packet** — R-1…R-14, each with evidence, options, a recommendation, and the exact diff that follows from each choice. Wave 0's deliverable | | [`50-rulings.md`](50-rulings.md) | 6 | **The ruling packet** — R-1…R-14, each with evidence, options, a recommendation, and the exact diff that follows from each choice. Wave 0's deliverable |

View file

@ -1,3 +1,12 @@
> **Correction note — 2026-07-27, at `ed06dd64` (holistic assessment Phase 6).**
> One claim in this document is false and is corrected here rather than silently edited, because this document is itself a *corrective* one and its error was inherited downstream.
>
> **The claim:** that `accrue_realized_knowledge` *"is enabled by the production L10 process."*
> **The code:** `CONTINUOUS_LIFE_CONFIG_FLAGS` (`chat/always_on_daemon.py:45-49`) forces `persist_session_state`, `consolidate_determinations`, and `strict_identity_continuity`**not** `accrue_realized_knowledge`. The daemon therefore consolidates determinations while the only turn-path writer of realized facts stays off, which is finding F-6: *as coded, the continuous life may consolidate an empty set.*
> **Not this document's fault alone:** `core/config.py`'s own docstrings for both flags make the same claim (recorded as H-8d in `docs/assessment/31-hindrance-audit.md`). Whether the flag set is incomplete or the dormancy is intended is ruling R-3 in `docs/assessment/50-rulings.md`; whichever way it goes, either one line of code or three records change.
>
> Everything else in this document was re-verified in the holistic assessment's Phases 23 and stands.
# Independent verification of the Tier-S arc assessment (2026-07-25) # Independent verification of the Tier-S arc assessment (2026-07-25)
**Input:** an external architectural assessment of the deduction-serve generalization arc **Input:** an external architectural assessment of the deduction-serve generalization arc
@ -61,7 +70,7 @@ turns flow through the *deduction* path. That comment has been stale since the f
| T12 weekly-ledger entry is "completely missing" | **False.** `docs/analysis/weekly-audit-2026-07-22-stragglers-todo.md:170` — an open pattern-watch item. The claim came from grepping commit messages only. | | T12 weekly-ledger entry is "completely missing" | **False.** `docs/analysis/weekly-audit-2026-07-22-stragglers-todo.md:170` — an open pattern-watch item. The claim came from grepping commit messages only. |
| `assert_corpus_sound()` has one impl; a second subject must reimplement it | **False.** `evals/curriculum_serve/runner.py:71` is already `assert_corpus_sound(domain, cases)`, and `build_report()` calls it plus `assert_provenance` plus `assert_anti_recall_coverage` unconditionally (`:105-108`). Structural, not discipline-based. | | `assert_corpus_sound()` has one impl; a second subject must reimplement it | **False.** `evals/curriculum_serve/runner.py:71` is already `assert_corpus_sound(domain, cases)`, and `build_report()` calls it plus `assert_provenance` plus `assert_anti_recall_coverage` unconditionally (`:105-108`). Structural, not discipline-based. |
| Manual promotion sweeps need an automated Promotion Oracle | **Already exists.** `docs/research/tier-s-housekeeping-2026-07-24.md` §3: "*not a one-time manual check… The sweep's proof is the passing gate, not a separate pass.*" `ds-mem-0020` surfaced *as* a lane failure. | | Manual promotion sweeps need an automated Promotion Oracle | **Already exists.** `docs/research/tier-s-housekeeping-2026-07-24.md` §3: "*not a one-time manual check… The sweep's proof is the passing gate, not a separate pass.*" `ds-mem-0020` surfaced *as* a lane failure. |
| `accrue_realized_knowledge` is "dark," a closing temporal-consistency window | **Misread.** `core/config.py:316-321` gates **session** memory, is `False` for eval/one-shot runtimes, and is enabled by the production L10 process. A deployment profile, not an epistemic debt clock. | | `accrue_realized_knowledge` is "dark," a closing temporal-consistency window | **Misread** — *and this correction was itself wrong on one point; see the note below.* `core/config.py:316-321` gates **session** memory and is `False` for eval/one-shot runtimes. A deployment profile, not an epistemic debt clock. |
| Cross-subject proof could "emerge accidentally" | **Structurally impossible.** `curriculum_surface.resolve_domain` requires both terms in one subject's vocabulary and refuses `ambiguous_reading` rather than picking. Deduction has no subject notion at all. | | Cross-subject proof could "emerge accidentally" | **Structurally impossible.** `curriculum_surface.resolve_domain` requires both terms in one subject's vocabulary and refuses `ambiguous_reading` rather than picking. Deduction has no subject notion at all. |
## 3. The flagship item is diagnosed at the wrong layer ## 3. The flagship item is diagnosed at the wrong layer