core/docs/assessment/30-gap-register.md
Shay 9efd5a6143 docs: apply the Wave 0 record amendments that need no ruling (N-1…N-7, H-12)
The corrections derived in 50-execution-plan.md §0, applied where the wrong
claims live. Every amendment is marked [AMENDED N-x] inline so the register
carries its own correction history rather than quietly reading correct.

Gap register — four entries amended:
- G-5  the soak RAN and passed (5000 beats, reboot at 2500, all four gates,
       aed273b1). Retitled: what is owed is the committed artifact, the pinned
       deterministic_digest, and a cadence — not the run. Both prior claims
       withdrawn with their causes: "no suite contains any l10 test" was a
       scan artifact (TEST_SUITES["full"] is a directory), and the pins
       execute twice a day (post-merge + nightly cron).
- G-7  respecified. The proposed meta-check was already satisfied and would
       have shipped green proving nothing — a hollow gate in the entry
       proposing to abolish hollow gates. Now two pins: gate parity, and
       curated-suite membership EXCLUDING full.
- G-8  28 default-off flags of 32 RuntimeConfig booleans, not seventeen.
- G-10 the 16-premise cap was removed by ADR-0264 R5 on 2026-07-26; the ADR
       carries a self-supersession banner directly under the heading Phase 4
       quoted. No engineering prerequisite remains — R-8 alone gates it.
- G-19 note: the 21/25 exposure is already pinned in test_volume_honesty.py,
       so the work is applying the demotions, not discovering them.

Hindrance audit:
- H-12 added — the local smoke suite (23) and smoke.yml (13) have diverged by
       ten files. Counter-evidence stated first: AGENTS.md makes the local run
       the merge gate and it is the superset, so nothing is unguarded. What is
       real is one comment claiming an equality that does not hold, and the
       push that skipped the local gate.
- H-8  gains instance (d), inside the code: two flag docstrings assert a
       production profile CONTINUOUS_LIFE_CONFIG_FLAGS does not have.
- H-1  gains the already-pinned note.

Also: correction note on the 2026-07-25 verification doc (a corrective
document that introduced the error at H-8c), and a staleness banner on
MIND-PHYSICS-BLUEPRINT.md pointing at the taxonomy that dissolved it.

The two ratified-ADR amendments (ADR-0146, ADR-0252) are NOT here — they are
drafted in full in 50-rulings.md R-12 and land on that ruling, because editing
a ratified ADR is Shay's authority even when the edit changes no decision.

Docs only. No runtime behavior, no flag, no test changed.
2026-07-27 17:38:12 -07:00

17 KiB
Raw Blame History

The Gap Register

Assessor: Fable 5 (Phase 4) · Verified at: 8927c563 (2026-07-27) Standing: This is CORE's first live gap register since docs/gaps.md closed its 26th entry. Proposal (for ruling): this register supersedes docs/gaps.md, which is marked historical; two dead registers plus a live one is worse than one live one. Discipline: A gap is an absence the telos requires filled with no explicit deferral ruling. Deferred-with-ruling is not a gap (scripture content is the model). Every entry carries evidence, its deciding authority, and a leverage rank. The register decides nothing. Amended: 2026-07-27 at ed06dd64 (Opus 5, Phase 6). Four entries carried claims that code, workflows, and an ADR's own supersession banner falsify. Each amendment is marked [AMENDED N-x] inline and derived in 50-execution-plan.md §0: G-5 (the soak ran and passed; the pins run twice a day), G-7 (the orphan scan was an artifact; the mechanism is respecified), G-8 (28 flags, not 17), G-10 (the engineering blocker was discharged 2026-07-26). G-19 gains a note: its exposure is already pinned in-repo.


Tier A — Frontier-blocking (each blocks a ratified commitment or the telos itself)

G-1 · The ADR-0252 §5 experiment has never returned a verdict

Layer: M3 · Leverage: 1 (highest in the assessment) The ratified governing paradigm's single load-bearing empirical claim — can Cl(4,1) geometry carry relational structure the SME way — sits authorized (§8.4), scaffolded (two unmerged rnd/ worktrees, tip bed29a09 "formalize §5 experiment scaffolding"), and unrun. Until it returns GO or NO-GO, the §6 comprehension correction cannot be authorized, and the 18 condemned organs serve indefinitely with no successor path. A well-controlled NO-GO is defined by the ADR as full credit — the experiment is cheap to finish and expensive to leave open. GSM8K's demotion to diagnostic mispriced this: the paradigm governs all future comprehension, not math. Evidence: ADR-0252 §5/§8; worktree log; M3 card. · Authority: execution (already authorized) + Shay's verdict ruling.

G-2 · The #138 fabrications — measured & pinned, fix held for ADR + ratification

Layer: M3 (locus: generate/meaning_graph/reader.py) → blast radius M4 · Leverage: 2 every dog is a mammalmember(every_dog, mammal); Given: furthermore; p implies q; p.asserted(furthermore) recited back as a served premise. The reader fabricates on 22 constructions beyond its 19-wide verified inventory. The fixes are known — two of the 13 mutations — and are deliberately held because they change what CORE comprehends from user input: serving-path truth behavior, ADR + ratification territory. Entered here pre-labeled per standing instruction; never re-discovered, never fixed by this assessment. Evidence: PR #138 @ c69f9948; realize-phase card (incl. the defensive-gate option: refuse to hold a reading outside the verified inventory). · Authority: the fabrication ADR + Shay's ratification.

G-3 · Reader inventory: 19 constructions against a 1739-construction writer, overlap 6

Layer: M3 · Leverage: 3 The comprehension frontier itself, measured. Standing ruling: close fabrications (G-2) before widening. The widening program after that is the largest single capability gap between CORE and its telos — and its shape depends on G-1's verdict (structure-mapping vs more constructions). Evidence: #138 inventory measurement; M3 card capacity block. · Authority: sequenced rulings (G-2 → G-1 → widening plan).

G-4 · CR-2 — the continuous life has no chooser

Layer: M6 / Candidate Register · Leverage: 4 Confirmed at component depth: drive objects exist (DriveGradientMap — constructed, never read; ExertionMeter — telemetry only); idle mechanisms exist (consolidation, proposal review, contemplation — each flag-gated, each doing one thing); nothing ranks what matters next. The daemon heartbeat advances idle_tick and nothing more ambitious. This is the AGI-grade conceptual absence: everything CORE does is chosen by the operator. Design work, not a flag flip. Evidence: attention-allocation + always-on-process cards; 02-layer-taxonomy.md CR-2. · Authority: design + ruling (does the L10 process own an agenda, governed by what).

G-5 · L10 proof debt — the soak ran and passed; what is owed is the artifact, the pinned digest, and a cadence [AMENDED N-1/N-2/N-4]

Layer: M6 / MV · Leverage: 5 The always-on process is built; the falsifiable harness (H1H4, holds/bites pairs, vacuity-guarded) is built; and it was run. evals/l10_always_on/contract.md §"The measured result" records a 5000-beat soak with a reboot at beat 2500 (landed 2026-07-19 in aed273b1) in which all four predicates pass: versor_condition flat at 1.389e-07 across all 5000 beats, vault bounded at 6 entries, convergence at beat 1 with the 4999-beat tail at rest, reboot resuming the same life with derived learning intact.

What remains owed is the ceremony, and it is this assessment's own §8 maintenance contract operating on this assessment's own frontier:

  • the result is prose in a contract file — no committed machine-readable report, no run SHA, no rerun path;
  • deterministic_digest is computed by report.py and pinned nowhere, though the contract's closing line says "Pin it once the lane is trusted so a regression flips it";
  • no cadence rules the lane, and the local-first/Mac-runner doctrine makes "nightly" itself a ruling rather than a cron line.

A recorded prose result with no pinned digest is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint.

Two prior claims in this entry were wrong and are withdrawn. (a) "No suite contains any l10/always_on test" was a scan artifact: TEST_SUITES["full"] = ("tests/",) is a directory, so every test file is trivially in a suite. The true statement is that the five tests/test_l10_*.py files are in no curated suite tuple — reachable only through full. (b) "nothing runs its pins" is false: CI never invokes core test --suite, it runs raw pytest with marker filters, and no L10 file carries quarantine or slow — so the pins execute twice a day, in full-pytest.yml (post-merge) and nightly-full-pytest.yml (cron 0 2 * * *). The correct statement is that nothing runs them on any pre-merge gate. Evidence: M6 + always-on-process cards; evals/l10_always_on/contract.md; core/cli_test.py:13; .github/workflows/{smoke,full-pytest,nightly-full-pytest}.yml. · Authority: execution + the R-9 evidence-standard and cadence ruling.

G-6 · F-6 — the lived learning loop is half-gated

Layer: M6/M5 · Leverage: 6 The daemon forces consolidate_determinations but not accrue_realized_knowledge; the only turn-path writer of realized facts sits behind the unforced flag. As coded, the continuous life may consolidate an empty set. Incomplete flag set, or intended dormancy — neither is documented, and a prior verification doc asserts the opposite of the code (C-5). Evidence: CONTINUOUS_LIFE_CONFIG_FLAGS (chat/always_on_daemon.py:45-49); determine-phase card gating table. · Authority: ruling (one flag + one sentence, or a documented dormancy rationale).


Tier B — Enforcement & instrument debt (capability exists; the guarantee doesn't)

G-7 · No gate-parity pin and no curated-suite membership check [AMENDED N-1/N-3]

Layer: MV · Leverage: 7the highest-leverage single mechanical change in the repository Suite tuples are hand-curated; a test file in zero curated suites is indistinguishable from one that runs everywhere.

The mechanism this entry originally proposed would not have worked. "Every tests/**/*.py belongs to ≥1 suite or an explicit exclusion list" is already satisfied, because TEST_SUITES["full"] = ("tests/",) is a directory — the pin would ship green and prove nothing. A hollow gate by the Third-Door criterion, in the entry proposing to abolish hollow gates.

The correct mechanism is two pins:

  1. Gate paritysmoke.yml's path set must equal TEST_SUITES["smoke"], parsed from the workflow file, failing on drift in either direction.
  2. Curated-suite membership — every tests/**/test_*.py is in ≥1 suite excluding full, or in a registered exclusion list with a one-line reason.

Pin 1 exists because the two gates have already drifted (see H-12): the local suite is 23 files, smoke.yml is 13, and the 10-file delta includes ADR-0265's denial pin, both ADR-governance pins, volume honesty, and curriculum polarity. Measured parity cost: 429 tests in 46s on the Act runner's own hardware. Evidence: MV card; core/cli_test.py:13; .github/workflows/smoke.yml; H-12. · Authority: mechanical (small PR) for the pins; the R-14 ruling for the parity direction.

G-8 · No flag-default register [AMENDED N-7]

Layer: cross-cut · Leverage: 8 RuntimeConfig is a single frozen dataclass with 32 boolean fields: 28 default False, 4 default True (allow_cross_language_recall, use_salience, discourse_planner, deduction_serving_enabled — the last ratified ON by ADR-0256). Three are daemon-forced (persist_session_state, consolidate_determinations, strict_identity_continuity). No document states the set, which defaults are deliberate posture vs accumulated hesitancy, or what evidence would flip each. The largest lever in the system, unregistered. The register format already exists in-repo: the ratified-ledger pattern (declare absence policy in the table, not the call site — ADR-0263 Rule 5).

(This entry originally said "seventeen." The verified count is 28 default-off. That two counts of the same set differ by eleven is itself the finding: nothing in the repository distinguishes a capability flag from a policy or deployment flag, which is exactly the classification the register must make.) Evidence: core/config.py at ed06dd64 (32 bool fields, 28 = False, 4 = True); daemon trio (Phase 3). · Authority: documentation PR + per-flag evidence bars set by ruling (R-4).

G-9 · Enforcement pins unverified for three doctrine-level prohibitions

Layer: M1 / MG · Leverage: 9 (a) No verified failing pin for the no-approximate-recall law (would a cosine ranker actually fail a test?); (b) no pin that fails when a layer bypasses governance entirely (as distinct from governance working when called); (c) safety-pack non-swappability not verified as mechanically enforced. All three are law in AGENTS.md; law-enforced-by-review is weaker than law-enforced-by-test. Evidence: M1/MG cards (flagged, not resolved, in Phase 23). · Authority: verification pass, then mechanical PRs.

G-10 · Curriculum SERVE is blocked by one ruling, and its ledger doesn't exist [AMENDED N-5]

Layer: M5 · Leverage: 10 The engineering blocker this entry named does not exist. ADR-0264 §4.1 carries a banner directly beneath its own heading: "SELF-SUPERSEDED by this ADR's own R5, discharged 2026-07-26. The heading is no longer true of the running system. … R5 removed that: compilation is query-scoped, so a family of any size answers. With the cap gone, four bands would earn SERVE the moment a ledger is sealedphysics·causal, systems_software·causal, and philosophy_theology·{modal,contrast}." Query-scoping already landed; this entry quoted a superseded heading past its own correction.

What survives, and is now the whole of the blockage: reliability is commitment precision and a correct UNKNOWN is a commitment, so a band clears θ_SERVE on non-commitments alone (conservative_floor(660,660) = 0.990046). The licensable evidence is 99.099.98% non-entailed and max entailed volume in any band is 9. chat/data/curriculum_serve_ledger.json is deliberately absent (the one honest missing_ok=True in production) and core proposal-queue reseal refuses a license without --allow-new-licenses, so nothing is licensed while the outcome-mix ruling is unmade. A committed ledger is necessarily an earning one. Evidence: M5 card; ADR-0264 §4.1 supersession banner + §5 (open); docs/research/curriculum-practice-producer-2026-07-26.md §1; chat/curriculum_serve_license.py:40-52. · Authority: the outcome-mix ruling (R-8) alone — no engineering prerequisite remains.

G-11 · Identity enforcement has no stated authorization bar

Layer: MG · Leverage: 11 identity_wave_gate is off and "not authorized" — a deliberate posture. What's missing is the criterion: no document states what evidence would authorize live refusal. Scoring-without-blocking is an honest state only while the path to blocking is defined. Evidence: MG card; runtime_contracts.md identity contract. · Authority: ruling (set the bar).


Tier C — One-line rulings (cheap to close; expensive only if left silent)

G-12 · CR-3 efferent action — deferred, or out of telos?

No system-level statement exists either way; the only adjacent text is one bench's v1 prohibition. The alignment posture is arguably stronger with action explicitly deferred — the ruling costs one line. Authority: ruling.

G-13 · CR-4 temporal self-location — a stance before the L10 spike

Determinism bans clocks; continuity implies lived time; the 24h+ no-drift requirement cannot be stated precisely without a stance on "now." The spike (G-5) should not be designed with an accidental answer. Authority: ruling (can be a paragraph in the soak's contract).

G-14 · CR-1 attention governance — the one-page ADR

Own use_salience, the two underived constants, the self-narrowing budget feedback, and the InhibitionMask disposition. Mechanism verified live and sound; only the governance is absent. Authority: ADR (one page — the card is its draft evidence base).

G-15 · The daemon's ratifying ADR

chat/always_on_daemon.py is unowned while ADR-0146 explicitly rejected the daemon shape it implements. Whatever the right answer, the record currently contradicts the code (see H-8). Authority: ADR amendment or a new short ADR.


Tier D — Latent & carried-forward (recorded so nothing silently drops)

  • G-16 · ADR-0265's defect class survives in _inflect_predicate's aspect arms (generate/templates.py:79) — 10,530/16,146 template points, not reachable today. Latent, recorded from the prior arc; becomes live if aspect arms become reachable. Authority: the widening program (G-3) must clear it first.
  • G-17 · Non-text ingest — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). Authority: ruling.
  • G-18 · Identity-divergence curriculum may still bypass formation's gates — known gap since 2026-05-17 (teaching_order.md); unverified at this SHA. Authority: Phase-3-style verification pass, then a routing PR.
  • G-19 · Wilson/replay evidence shortfall — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as evidence debt on existing licenses; the counting fix is the hindrance entry. Note (Phase 6): the exposure is already pinned in-repotests/test_volume_honesty.py (ADR-0264 R9, in the local smoke gate) pins the 21-of-25 shortfall "in BOTH directions" and calls its inventory "an EXPOSURE INVENTORY, not an approved baseline." So the open work is applying the demotions, not discovering them, and that pin moves in the same PR. Authority: ADR amendment + re-count, authorized by R-13.
  • G-20 · The refusal_reason materialisation — typed refusal evidence exists and is discarded at the public str boundary; the plumbing for materialisation already landed. Cross-listed as H-3. Authority: small ADR (anticipated by the ADR-0024 chain).

What is not in this register, and why

  • Scripture/theology content — deferred by explicit ruling (2026-07-26); the model case for deferred-is-not-missing.
  • Benchmark wins — excluded by the completeness criterion itself (taxonomy §6): architectural distinctiveness is the target; benchmarks are downstream validation.
  • Sociality, affect, full embodiment — considered and not registered, with reasons, in the taxonomy's Candidate Register; importing them would violate Pillar II.
  • Rust-backend default — an open question with a stated blocker (crates.io unreachable under sandbox), not a gap; the measured case for urgency dissolved (0.22%).