diff --git a/docs/assessment/00-scope-and-method.md b/docs/assessment/00-scope-and-method.md index 60248073..140e2da4 100644 --- a/docs/assessment/00-scope-and-method.md +++ b/docs/assessment/00-scope-and-method.md @@ -44,8 +44,8 @@ The assessment is conducted under `docs/conceptualizing_engineering_mastery.md`, | **1 — Taxonomy & schema** | The macro→micro layer taxonomy and the metadata card schema every card must fill | Fable 5 | | **2 — Macro layer cards** | One card per top-level layer; layer-level verdicts; cross-cutting concerns | Opus 5 | | **3 — Micro component cards** | Per-subsystem descent, depth allocated by load-bearing-ness | Fable 5 | -| **4 — Gap register + hindrance audit** | Two separate registers; evidence-carrying | Opus 5 | -| **5 — Synthesis** | Executive assessment; ranked gaps; ranked hindrances; recommended R&D attack order | Opus 5 | +| **4 — Gap register + hindrance audit** | Two separate registers; evidence-carrying | Fable 5 *(reassigned from Opus 5 by Shay, 2026-07-27)* | +| **5 — Synthesis** | Executive assessment; ranked gaps; ranked hindrances; recommended R&D attack order | Fable 5 *(same reassignment)* | Phase 1 is the keystone. A wrong taxonomy miscategorizes everything downstream, and it is the cheapest phase to correct. diff --git a/docs/assessment/30-gap-register.md b/docs/assessment/30-gap-register.md new file mode 100644 index 00000000..635f55fd --- /dev/null +++ b/docs/assessment/30-gap-register.md @@ -0,0 +1,103 @@ +# The Gap Register + +**Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27) +**Standing:** This is CORE's first *live* gap register since `docs/gaps.md` closed its 26th entry. Proposal (for ruling): this register supersedes `docs/gaps.md`, which is marked historical; two dead registers plus a live one is worse than one live one. +**Discipline:** A gap is an *absence the telos requires filled* with no explicit deferral ruling. Deferred-with-ruling is not a gap (scripture content is the model). Every entry carries evidence, its **deciding authority**, and a leverage rank. The register decides nothing. + +--- + +## Tier A — Frontier-blocking (each blocks a ratified commitment or the telos itself) + +### G-1 · The ADR-0252 §5 experiment has never returned a verdict +**Layer:** M3 · **Leverage: 1 (highest in the assessment)** +The ratified governing paradigm's single load-bearing empirical claim — can Cl(4,1) geometry carry relational structure the SME way — sits authorized (§8.4), scaffolded (two unmerged `rnd/` worktrees, tip `bed29a09` "formalize §5 experiment scaffolding"), and unrun. Until it returns GO or NO-GO, the §6 comprehension correction cannot be authorized, and the 18 condemned organs serve indefinitely with no successor path. A well-controlled NO-GO is *defined by the ADR as full credit* — the experiment is cheap to finish and expensive to leave open. GSM8K's demotion to diagnostic mispriced this: the paradigm governs **all** future comprehension, not math. +**Evidence:** ADR-0252 §5/§8; worktree log; `M3` card. · **Authority:** execution (already authorized) + Shay's verdict ruling. + +### G-2 · The #138 fabrications — *measured & pinned, fix held for ADR + ratification* +**Layer:** M3 (locus: `generate/meaning_graph/reader.py`) → blast radius M4 · **Leverage: 2** +`every dog is a mammal` → `member(every_dog, mammal)`; `Given: furthermore; p implies q; p.` → `asserted(furthermore)` recited back as a served premise. The reader fabricates on 22 constructions beyond its 19-wide verified inventory. The fixes are known — two of the 13 mutations — and are **deliberately held** because they change what CORE comprehends from user input: serving-path truth behavior, ADR + ratification territory. Entered here pre-labeled per standing instruction; never re-discovered, never fixed by this assessment. +**Evidence:** PR #138 @ `c69f9948`; `realize-phase` card (incl. the defensive-gate option: refuse to *hold* a reading outside the verified inventory). · **Authority:** the fabrication ADR + Shay's ratification. + +### G-3 · Reader inventory: 19 constructions against a 1739-construction writer, overlap 6 +**Layer:** M3 · **Leverage: 3** +The comprehension frontier itself, measured. Standing ruling: close fabrications (G-2) **before** widening. The widening program after that is the largest single capability gap between CORE and its telos — and its *shape* depends on G-1's verdict (structure-mapping vs more constructions). +**Evidence:** #138 inventory measurement; `M3` card capacity block. · **Authority:** sequenced rulings (G-2 → G-1 → widening plan). + +### G-4 · CR-2 — the continuous life has no chooser +**Layer:** M6 / Candidate Register · **Leverage: 4** +Confirmed at component depth: drive objects exist (`DriveGradientMap` — constructed, never read; `ExertionMeter` — telemetry only); idle mechanisms exist (consolidation, proposal review, contemplation — each flag-gated, each doing one thing); **nothing ranks what matters next**. The daemon heartbeat advances `idle_tick` and nothing more ambitious. This is the AGI-grade conceptual absence: everything CORE does is chosen by the operator. Design work, not a flag flip. +**Evidence:** `attention-allocation` + `always-on-process` cards; `02-layer-taxonomy.md` CR-2. · **Authority:** design + ruling (does the L10 process own an agenda, governed by what). + +### G-5 · L10 proof debt — the soak has never produced an artifact, and nothing runs its pins +**Layer:** M6 / MV · **Leverage: 5** +The always-on process is built; the falsifiable harness (H1–H4, holds/bites pairs, vacuity-guarded) is built; **no recorded long-horizon artifact exists, no suite contains any `l10`/`always_on` test, no nightly cadence exists** — and the local-first/Mac-runner doctrine makes "nightly" itself need a ruling rather than a cron line. The still-owed ADR-0146 Phase-4 spike, in its modern form: run the soak, record the artifact, schedule the pins. +**Evidence:** `M6` + `always-on-process` cards; suite-membership scan. · **Authority:** execution + MV suite ruling + cadence ruling. + +### G-6 · F-6 — the lived learning loop is half-gated +**Layer:** M6/M5 · **Leverage: 6** +The daemon forces `consolidate_determinations` but not `accrue_realized_knowledge`; the only turn-path writer of realized facts sits behind the unforced flag. As coded, the continuous life may consolidate an empty set. Incomplete flag set, or intended dormancy — **neither is documented**, and a prior verification doc asserts the opposite of the code (C-5). +**Evidence:** `CONTINUOUS_LIFE_CONFIG_FLAGS` (`chat/always_on_daemon.py:45-49`); `determine-phase` card gating table. · **Authority:** ruling (one flag + one sentence, or a documented dormancy rationale). + +--- + +## Tier B — Enforcement & instrument debt (capability exists; the guarantee doesn't) + +### G-7 · No orphaned-pin meta-check +**Layer:** MV · **Leverage: 7** +Suite tuples are hand-curated; a test file in zero suites is indistinguishable from one that runs everywhere. This is the *mechanism* by which G-5 happened. A meta-pin — every `tests/**/*.py` belongs to ≥1 suite or an explicit exclusion list — converts the doctrine "a pin in no suite never runs" into a failing test. Likely the highest-leverage *single mechanical change* in the repository. +**Evidence:** `MV` card; the M6 case as the demonstration. · **Authority:** mechanical (small PR); no ruling needed. + +### G-8 · No flag-default register +**Layer:** cross-cut · **Leverage: 8** +Seventeen capability flags default `False`; one is ratified ON (`deduction_serving_enabled`); three are daemon-forced (`persist_session_state`, `consolidate_determinations`, `strict_identity_continuity`). No document states the set, which defaults are deliberate posture vs accumulated hesitancy, or what evidence would flip each. The largest lever in the system, unregistered. The register format already exists in-repo: the ratified-ledger pattern (declare absence policy in the table, not the call site — ADR-0263 Rule 5). +**Evidence:** `core/config.py` scan (Phase 2); daemon trio (Phase 3). · **Authority:** documentation PR + per-flag evidence bars set by ruling. + +### G-9 · Enforcement pins unverified for three doctrine-level prohibitions +**Layer:** M1 / MG · **Leverage: 9** +(a) No verified failing pin for the no-approximate-recall law (would a cosine ranker actually fail a test?); (b) no pin that fails when a layer *bypasses* governance entirely (as distinct from governance working when called); (c) safety-pack non-swappability not verified as mechanically enforced. All three are law in `AGENTS.md`; law-enforced-by-review is weaker than law-enforced-by-test. +**Evidence:** `M1`/`MG` cards (flagged, not resolved, in Phase 2–3). · **Authority:** verification pass, then mechanical PRs. + +### G-10 · Curriculum SERVE is fully blocked by one engineering item, and its ledger doesn't exist +**Layer:** M5 · **Leverage: 10** +ADR-0264 §4.1: the 16-premise compilation cap holds every band to ≤16 entailed cases, so **no curriculum band can earn SERVE until query-scoping lands** — an engineering blocker gating a content problem that is itself quantified at 24×–73× under-fed. Downstream, `chat/data/curriculum_serve_ledger.json` is absent (the one honest `missing_ok=True` in production), and a committed ledger is necessarily an *earning* one — the outcome-mix ruling remains the binding constraint. +**Evidence:** `M5` card; ADR-0264 §4.1; `chat/curriculum_serve_license.py:46`. · **Authority:** engineering (scoping) + outcome-mix ruling. + +### G-11 · Identity enforcement has no stated authorization bar +**Layer:** MG · **Leverage: 11** +`identity_wave_gate` is off and "not authorized" — a deliberate posture. What's missing is the *criterion*: no document states what evidence would authorize live refusal. Scoring-without-blocking is an honest state only while the path to blocking is defined. +**Evidence:** `MG` card; `runtime_contracts.md` identity contract. · **Authority:** ruling (set the bar). + +--- + +## Tier C — One-line rulings (cheap to close; expensive only if left silent) + +### G-12 · CR-3 efferent action — deferred, or out of telos? +No system-level statement exists either way; the only adjacent text is one bench's v1 prohibition. The alignment posture is arguably *stronger* with action explicitly deferred — the ruling costs one line. **Authority:** ruling. + +### G-13 · CR-4 temporal self-location — a stance before the L10 spike +Determinism bans clocks; continuity implies lived time; the 24h+ no-drift requirement cannot be stated precisely without a stance on "now." The spike (G-5) should not be designed with an accidental answer. **Authority:** ruling (can be a paragraph in the soak's contract). + +### G-14 · CR-1 attention governance — the one-page ADR +Own `use_salience`, the two underived constants, the self-narrowing budget feedback, and the `InhibitionMask` disposition. Mechanism verified live and sound; only the governance is absent. **Authority:** ADR (one page — the card is its draft evidence base). + +### G-15 · The daemon's ratifying ADR +`chat/always_on_daemon.py` is unowned while ADR-0146 explicitly rejected the daemon shape it implements. Whatever the right answer, the record currently contradicts the code (see H-8). **Authority:** ADR amendment or a new short ADR. + +--- + +## Tier D — Latent & carried-forward (recorded so nothing silently drops) + +- **G-16 · ADR-0265's defect class survives in `_inflect_predicate`'s aspect arms** (`generate/templates.py:79`) — 10,530/16,146 template points, *not reachable today*. Latent, recorded from the prior arc; becomes live if aspect arms become reachable. **Authority:** the widening program (G-3) must clear it first. +- **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling. +- **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR. +- **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Authority:** ADR amendment + re-count. +- **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain). + +--- + +## What is *not* in this register, and why + +- **Scripture/theology content** — deferred by explicit ruling (2026-07-26); the model case for deferred-is-not-missing. +- **Benchmark wins** — excluded by the completeness criterion itself (taxonomy §6): architectural distinctiveness is the target; benchmarks are downstream validation. +- **Sociality, affect, full embodiment** — considered and not registered, with reasons, in the taxonomy's Candidate Register; importing them would violate Pillar II. +- **Rust-backend default** — an open *question* with a stated blocker (crates.io unreachable under sandbox), not a gap; the measured case for urgency dissolved (0.22%). diff --git a/docs/assessment/31-hindrance-audit.md b/docs/assessment/31-hindrance-audit.md new file mode 100644 index 00000000..97fc10b1 --- /dev/null +++ b/docs/assessment/31-hindrance-audit.md @@ -0,0 +1,96 @@ +# The Hindrance Audit + +**Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27) +**Discipline:** A hindrance is something *present* that works against the goal — a wrong solution for its underlying problem, a responsibility lodged in the wrong owner, a trade-off tuned instead of dissolved, or a record that misleads the next reasoner. Every entry carries evidence, a fitness verdict from the schema vocabulary, a **proposed better home**, and its deciding authority. Per standing rule: settled rulings are constraints — an entry may flag a ratified decision only on evidence, only for ruling, never as a unilateral recommendation to reverse. **This audit decides nothing.** + +Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS.md protocol — not by ease. + +--- + +## H-1 · License evidence counted on an independence assumption replay violates +**Verdict:** `wrong-solution` (the *counting basis*, not the gating) · **Layers:** M5 +**Evidence:** Wilson lower-bound licensing (θ_SERVE=0.99, ADR-0175 lineage) assumes independent trials; a replay of the same sealed case is one trial observed again, not a new one. Measured consequence recorded in the prior arc: **21 of 25 ratified bands fall short** of their floor when replays are deduplicated. +**Why it hinders:** the entire earned-license architecture — CORE's mechanism for *deserving* to serve — rests on the evidence count. An overstated count grants licenses the evidence doesn't support, which is precisely the failure the mechanism exists to prevent. The gate is right; the arithmetic feeding it is not. +**Better home:** distinct-evidence counting at the seal boundary (count distinct cases; a replay refreshes, never increments), declared in the ledger schema the way ADR-0263 Rule 5 declares absence policy — in the table, not the call site. +**Authority:** ADR amendment (0175/0263 lineage) + a re-count of the 25 bands. The re-count may demote licenses; that is the mechanism working. + +## H-2 · Decoration in the runtime constructor — objects built and never read +**Verdict:** decoration (fails the sabotage test) · **Layers:** M6/M3 +**Evidence:** `DriveGradientMap` constructed at `chat/runtime.py:716`, read nowhere. `InhibitionMask`/`InhibitionOperator` exported by `core/physics/__init__.py`, constructed on no path. Deleting either changes no output. +**Why it hinders:** dead structure is not neutral — it is *testimony*. Both objects tell every reader that drive mapping and inhibition masking are live, and Phase 1 of this very assessment initially believed them. Decoration is how architecture lies without anyone lying. +**Better home:** deletion (mastery algorithm step 2: the best part is no part), with their *intents* preserved where they belong — drive in the CR-2 design (G-4), the mask's disposition in the CR-1 ADR (G-14). If a future mechanism needs them, re-adding a deleted class is cheap; un-believing a phantom is not. +**Authority:** mechanical PR + one line each in the CR-1/CR-2 decisions. + +## H-3 · The typed refusal is constructed, then discarded at the public boundary +**Verdict:** `strained` — truth built and unserved · **Layers:** M4 +**Evidence:** `InnerLoopExhaustion` carries reason, region, and per-step rejected-attempt evidence; `respond()`/`arespond()` convert it to `""` for the `str` contract, so a refusing turn serves the empty string with `refusal_reason == ""`. The plumbing to materialise (`CognitiveTurnResult.refusal_reason`, `compute_trace_hash` fold) already landed; `runtime_contracts.md` names it a residual awaiting a future ADR. +**Why it hinders:** the honesty machinery is the product. A system whose refusals are richer than its answers, serving its refusals as nothing, undersells its own thesis on every hard turn. +**Better home:** materialise into `ChatResponse.refusal_reason` (and a minimal honest surface), per the contract's own anticipation. +**Authority:** small ADR — the chain already reserved the seam. + +## H-4 · Composer-arm precedence is ordered branches above a declarative resolver +**Verdict:** `strained` — a solved pattern not yet extended · **Layers:** M4 +**Evidence:** `core/cognition/surface_resolution.py` (494 lines) resolves the pipeline seam by *declared* precedence, self-documenting, with an in-code falsifiable contract for its own regression. Upstream, the composer arms (deduction `:1834`, curriculum `:1850`, pack/narrative/example/relation `:1871–:1936`, determination, estimate, gate, hedge) remain ordered branches across `chat/runtime.py`, in a different package with a different owner — nothing structurally prevents arm N+1 from bypassing the resolver's disciplines. +**Why it hinders:** prospectively — each new serving capability adds an arm ahead of any declared order. With only deduction ON, the live complexity is modest; the time to dissolve the pattern is *before* the next three arms, not after. +**Better home:** extend the resolver's declared-precedence pattern upstream to arm selection — the Third Door here is half-built and proven to fit this codebase. +**Authority:** refactor ADR; low-risk while one arm is live. + +## H-5 · Underived constants at the semantic center of generation +**Verdict:** `strained` — Pillar I violation with an in-repo counterexample · **Layers:** M3/CR-1 +**Evidence:** `salience_top_k=16`, `inhibition_threshold=0.3` gate every token walk's candidate set; no recorded derivation exists for either. The contrast is instructive and in-repo: `admissibility_margin δ=0.4` was derived from the minimum observed margin of a characterization corpus (0.456), declared *falsifiable*, and survived a 20-case stratified attempt — the standard exists two config lines away. +**Why it hinders:** "thresholds tuned for good-enough" at the exact point where the system decides what it may consider. Also the self-narrowing budget feedback (`stream.py:637`) — a real cognitive property nobody has named or justified. +**Better home:** the CR-1 ADR (G-14) with an empirical derivation in the δ=0.4 style. +**Authority:** ADR + a small characterization run. + +## H-6 · The half-forced flag pair gating the lived learning loop +**Verdict:** `misplaced` responsibility — a *set* decision made one flag at a time · **Layers:** M6/M5 +**Evidence:** F-6 (`05-phase3-findings.md`): the daemon forces the consolidator, not the accruer; the loop's writer and its consumer are gated independently, and only the consumer is on. +**Why it hinders:** flags that must be coherent *as a set* are owned nowhere as a set. `CONTINUOUS_LIFE_CONFIG_FLAGS` is the right pattern (a named, documented flag *profile*) applied to the wrong subset. +**Better home:** the flag-default register (G-8) with named profiles (one-shot / eval / continuous-life), each profile ruled as a unit. +**Authority:** ruling on the accrual flag + the register PR. + +## H-7 · The production ingest boundary lacks the trust contract its sibling has +**Verdict:** `strained` — the standard exists and stops one layer short · **Layers:** M2 +**Evidence:** formation declares six boundaries — content-addressed in/out, no floats in hashed payloads, no pickle, an audit record per rejection. `ingest/gate.py`, facing untrusted user text in production, has the versor gate and the AGENTS.md trust-boundary defaults, but no comparable declared table. +**Why it hinders:** asymmetric rigor invites the assumption that the un-tabled boundary is the less important one; it is the opposite. +**Better home:** an M2 trust-boundary table in `runtime_contracts.md`, formation-style; hardening PRs only where the table exposes real deltas. +**Authority:** documentation first; evidence decides whether code follows. + +## H-8 · The record contradicts the code at three load-bearing points +**Verdict:** `wrong-solution` as *record-keeping* — divergence that reasoners inherit · **Layers:** governance +**Evidence:** (a) ADR-0146 rejects the daemon shape; an unowned daemon ships. (b) ADR-0252's headline "34 organs" has no reproducible basis (18 entry organs at the ratification commit itself; ~32 modules). (c) `architecture-assessment-verification-2026-07-25.md` asserts accrual "is enabled by the production L10 process"; the flag set says otherwise. +**Why it hinders:** demonstrated, not hypothetical — this assessment's own Phase 0 inherited a stale-record error, and the 2026-07-25 doc (itself a *corrective* document) introduced one. Every divergence is a future wrong analysis. +**Better home:** three one-paragraph amendments (ADR-0146 addendum owning the daemon or superseding the rejection; ADR-0252 basis sentence; a correction note on the 07-25 doc). +**Authority:** docs PRs + ruling signatures. + +## H-9 · Dead instruments still standing as if live +**Verdict:** `superseded-in-place` (unratified) · **Layers:** MV/governance +**Evidence:** `docs/gaps.md` — 26/26 closed, no entry from any 2026-06+ arc; `substrate-liveness-ratchet` — v5, stale since ~2026-05-24, all OPEN items L10-chained; ~130 analysis docs with no aggregator. The system map — the best macro artifact — is local, gitignored, 48 days stale, and was wrong precisely where the project moved fastest; its phantom "L12" stratum exists nowhere else. +**Why it hinders:** an instrument that *looks* authoritative converts "I should check" into "I already checked." Phase 0's error was this mechanism operating on this assessment. +**Better home:** this register supersedes `docs/gaps.md` (marked historical); the ratchet's 7 OPEN items migrate here (G-5 absorbs their L10 dependency); the map stays a regeneratable local index per D5, with "L12" dropped; the assessment directory becomes the standing ruled record, `verified_at`-stamped. +**Authority:** ruling (one PR). + +## H-10 · The demotion that mispriced the paradigm experiment +**Verdict:** `strained` framing — a correct ruling casting an incorrect shadow · **Layers:** M3/governance +**Evidence:** GSM8K was demoted to diagnostic (correct — the flags-and-benchmarks reasoning stands). The §5 SME experiment lives in GSM8K's neighborhood (`holdout_dev/v1`, math structures), so it inherited the demotion's priority — but its verdict governs the *comprehension paradigm for everything*, per ADR-0252's own §4 conformance bar. +**Why it hinders:** the highest-leverage open item in the project (G-1) has been priced as math-lane housekeeping. +**Better home:** none needed — G-1's execution *is* the fix; this entry exists so the mispricing mechanism is named and not repeated. +**Authority:** already covered by G-1's ruling. + +## H-11 · A silent-failure pinhole inside a typed layer +**Verdict:** `strained` (small, cheap, principled) · **Layers:** M3 +**Evidence:** `_accrue_in_turn`'s broad guard converts any exception in the read→realize→determine chain into a no-op accrual with no telemetry (F-10). Defensible as a backstop; invisible as a signal — in the one layer whose constitution is "failures are typed, never silent" (INV-34). +**Better home:** count the swallow (a telemetry field on `IdleTickResult`/turn accrual), not a behavior change. +**Authority:** mechanical PR. + +--- + +## Explicitly examined and cleared + +For symmetry with the Candidate Register's "considered and not registered" — hindrance candidates this audit *rejects*: + +- **The 18 derivation organs** — condemned but *ruled* to keep serving (`superseded-in-place` by explicit ruling); their continued service is governance working, not failing. The hindrance was their unreproducible count (H-8b), not their existence. +- **Off-serve quarantines** (holographic vault, wave modules, `topological_reasoning`) — capacity that exists and cannot be used *by AST-pinned design*; legitimate research containment with failing-when-violated enforcement. An exit criterion would be nice (M1 card); the quarantine itself is fit. +- **Pure-Python-by-default algebra** — measured as the correct posture: determinism is the product, `versor_condition` is 0.22% of turn time, and the urgency argument for Rust-by-default dissolved under measurement. The open parity question (blocked on network) is a question, not a hindrance. +- **The five unreconciled articulations** — dissolved by the taxonomy (D1), not a standing hindrance; the residue is one stale Draft banner (folded into H-8's amendment batch). +- **Flag-gated conservatism itself** — seventeen dark flags is not inherently hindrance; *unregistered* darkness is (G-8). The posture may be exactly right; the register exists so that judgment can be made deliberately.