Merge pull request 'docs(assessment): the holistic macro→micro assessment — taxonomy, 17 cards, both registers, synthesis' (#139) from docs/holistic-assessment into main
This commit is contained in:
commit
ed06dd64f3
28 changed files with 2372 additions and 0 deletions
78
docs/assessment/00-scope-and-method.md
Normal file
78
docs/assessment/00-scope-and-method.md
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
# CORE Holistic Assessment — Scope and Method
|
||||
|
||||
**Status:** Ratified approach (Shay, 2026-07-27). Phase 0 complete.
|
||||
**Branch:** `docs/holistic-assessment` (worktree `core-wt-assess`, based on `forgejo/main` @ `8927c563`).
|
||||
**Nature:** Read-only investigation. This assessment changes no runtime behavior, fixes no defect, and decides nothing. It produces evidence and judgments for ruling.
|
||||
|
||||
---
|
||||
|
||||
## 1. The question being answered
|
||||
|
||||
Four questions, in order of dependency:
|
||||
|
||||
1. **Where does CORE actually stand** on its cognitive cycle — design articulated versus implementation fulfilled — from the telos down to individual components?
|
||||
2. **Is the layer model itself complete?** Are there layers, sublayers, or components missing from the *design*, not merely from the implementation — things an AGI/ASI-grade system requires that nothing in the current architecture accounts for?
|
||||
3. **What is the metadata for every layer, sublayer, and component?** Philosophical intent, functional contract, design shape, implementation status, evidence, capacity, and role — such that any future dive begins with a clear target rather than a grep.
|
||||
4. **What is hindering us?** Implemented ADRs or designs that are the wrong solution for their underlying problem; responsibilities lodged in the wrong subsystem; trade-offs tuned rather than dissolved.
|
||||
|
||||
Question 4 is not a courtesy pass. It is the question with the highest expected value, because a wrong component that works is more expensive than a missing component that is known to be missing.
|
||||
|
||||
---
|
||||
|
||||
## 2. Governing method
|
||||
|
||||
The assessment is conducted under `docs/conceptualizing_engineering_mastery.md`, applied to the assessment itself and not only to its subject.
|
||||
|
||||
**Pillar I — Semantic Rigor.** Completeness criteria and the metadata schema are defined *before* any component is judged (Phase 1), so that "missing" and "complete" have fixed meanings rather than per-component ones. Every `implemented` verdict requires an evidence pointer: a test, an eval lane, a pinned SHA, or an acceptance packet. "The module exists" is not evidence that it executes; "the flag is threaded" is not evidence that it changes behavior.
|
||||
|
||||
**Pillar II — Mechanical Sympathy.** Components are judged against CORE's own doctrine — deterministic decoding, exact recall, field-as-substrate with intelligence in the wiring, replay-gated learning, `wrong=0`-or-refuse — and not against a generic AGI checklist. A capability only counts as a gap if CORE's own telos requires it. Importing an external architecture's expectations would manufacture false gaps.
|
||||
|
||||
**Pillar III — The Third Door.** The hindrance audit's explicit charter: find the places where a bad trade-off was split rather than dissolved, and the places where deletion beats addition. Per the execution algorithm, deletion (step 2) precedes optimization (step 3); a component that should not exist is never a performance problem.
|
||||
|
||||
**Two standing discipline rules carried in from prior work:**
|
||||
|
||||
- **The sabotage test.** For every claim that a mechanism is live and load-bearing, ask what the measurement would look like with the mechanism removed. If it would look identical, the claim is decoration and is recorded as such. This repository has produced exactly that failure before — a rate reported as evidence of a reader that was `0.0` throughout.
|
||||
- **Identity, not value.** When measuring whether two things are the same thing, measure identity rather than equal-looking values. Source-scanning metrics can move the wrong way on success.
|
||||
|
||||
---
|
||||
|
||||
## 3. Phases and division of labor
|
||||
|
||||
| Phase | Deliverable | Executor |
|
||||
|---|---|---|
|
||||
| **0 — Ground truth** | Canonical-document ingestion; ADR triage; system-map recovery; raw material and open tensions for the taxonomy | Opus 5 — **complete** |
|
||||
| **1 — Taxonomy & schema** | The macro→micro layer taxonomy and the metadata card schema every card must fill | Fable 5 |
|
||||
| **2 — Macro layer cards** | One card per top-level layer; layer-level verdicts; cross-cutting concerns | Opus 5 |
|
||||
| **3 — Micro component cards** | Per-subsystem descent, depth allocated by load-bearing-ness | Fable 5 |
|
||||
| **4 — Gap register + hindrance audit** | Two separate registers; evidence-carrying | Fable 5 *(reassigned from Opus 5 by Shay, 2026-07-27)* |
|
||||
| **5 — Synthesis** | Executive assessment; ranked gaps; ranked hindrances; recommended R&D attack order | Fable 5 *(same reassignment)* |
|
||||
|
||||
Phase 1 is the keystone. A wrong taxonomy miscategorizes everything downstream, and it is the cheapest phase to correct.
|
||||
|
||||
---
|
||||
|
||||
## 4. Deliverables
|
||||
|
||||
```
|
||||
docs/assessment/
|
||||
00-scope-and-method.md # this file
|
||||
01-phase0-ground-truth.md # Phase 0 findings + Phase 1 handoff
|
||||
02-layer-taxonomy.md # Phase 1
|
||||
03-card-schema.md # Phase 1
|
||||
10-layer-cards/ # Phase 2
|
||||
20-component-cards/ # Phase 3
|
||||
30-gap-register.md # Phase 4
|
||||
31-hindrance-audit.md # Phase 4
|
||||
40-assessment.md # Phase 5
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Rules of engagement
|
||||
|
||||
1. **Read-only.** No runtime code is modified. No defect is fixed. Where a fix is obvious, it is recorded as a finding with its evidence, not applied.
|
||||
2. **Verify against code, not against documents.** Documentation in this repository has been measurably wrong about the repository before — most recently `docs/research/architecture-assessment-verification-2026-07-25.md` falsified roughly a third of an external blueprint's work items by reading the implicated code. A claim sourced only from a document is labeled as such.
|
||||
3. **Settled rulings are constraints, not subjects.** The deduction pivot, the scripture-content deferral, the no-merge-automation rule, and ratified ADRs enter as given. An ADR is reopened only on evidence of hindrance, and only as a flag for ruling — never as a unilateral recommendation to reverse.
|
||||
4. **Known-and-held findings are recorded as such.** The PR #138 fabrication findings (`every dog is a mammal` → `member(every_dog, mammal)`; `Given: furthermore; p implies q; p.` reaching served output) are **measured and pinned, not fixed**. The fixes are known — two of the 13 mutations — and are deliberately held out pending ADR and ratification because they change what CORE comprehends from user input, which is serving-path truth behavior. They enter the gap register pre-labeled and are never re-presented as newly discovered.
|
||||
5. **No timelines.** Scope size, phase, and priority only.
|
||||
6. **All wrinkles surfaced.** Technical truths are volunteered unprompted, including ones that complicate the picture or reflect badly on prior work.
|
||||
195
docs/assessment/01-phase0-ground-truth.md
Normal file
195
docs/assessment/01-phase0-ground-truth.md
Normal file
|
|
@ -0,0 +1,195 @@
|
|||
# Phase 0 — Ground Truth and Phase 1 Handoff
|
||||
|
||||
**Executor:** Opus 5, 2026-07-27. **Base:** `forgejo/main` @ `8927c563`.
|
||||
**Charter:** ingest the canonical corpus, recover the existing macro→micro artifacts, triage the decision record, and hand Phase 1 the raw material plus the tensions it must resolve.
|
||||
|
||||
**What Phase 0 deliberately did not do:** build the taxonomy, fill any card, or judge any component. Those are Phases 1–3. Findings below are recorded as *evidence and tension*, not as verdicts. Where a finding looks like a gap or a hindrance, it is flagged for the Phase 4 registers rather than adjudicated here.
|
||||
|
||||
---
|
||||
|
||||
## 1. The canonical corpus — what is authoritative, and as of when
|
||||
|
||||
| Document | Role | Dated | Status |
|
||||
|---|---|---|---|
|
||||
| `AGENTS.md` | Canonical governance. Wins over all provider files. | live | Authoritative |
|
||||
| `docs/specs/runtime_contracts.md` | Frozen runtime contracts; INV-21…INV-34 | live (1089 lines) | Authoritative |
|
||||
| `docs/adr/MASTER-BLUEPRINT-2026-07-20-ADR-MAPPING.md` | Blueprint↔registry collision reconciliation | 2026-07-20 | **Governing** |
|
||||
| `docs/adr/ADR-0252` | **The governing problem-solving paradigm** | 2026-07-19 | **Accepted / ratified** |
|
||||
| `docs/position_paper.md` | Thesis and philosophical intent | — | Reference |
|
||||
| `docs/Whitepaper.md` / `docs/Yellowpaper.md` | Architecture narrative / formal Cl(4,1) spec | — | Reference |
|
||||
| `docs/architecture/MIND-PHYSICS-BLUEPRINT.md` | Three-physics-layer cognitive cycle | 2026-05-12 | **Draft**, never advanced |
|
||||
| `docs/master-plan-post-substrate-audit.md` | Phase plan, W-* cascade | 2026-05-24 (+ amendments) | Largely superseded |
|
||||
| `docs/audit/substrate-liveness-ratchet.md` | Wiring-debt registry (v5) | ~2026-05-24 | 7 OPEN, stale |
|
||||
| `docs/gaps.md` | Capability-gap register | — | **All 26 entries closed** |
|
||||
| `CLAIMS.md` | Machine-generated claim ledger: 5 Tier-1 domains, 11 Tier-2 pinned lanes | auto | Authoritative |
|
||||
| `.system-map/` | 33 zones / 205 subsystems / 246 edges / 10 bands | 2026-06-09 | Local-only, gitignored, 48 days stale |
|
||||
|
||||
Validation surface: 21 CLI suites (`fast, smoke, runtime, cognition, teaching, packs, algebra, sensorium, pulse, formation, proof, refusal, margin, rotor, inner-loop, phase5, phase6, adr-0024, math, deductive, full`).
|
||||
|
||||
Code topology: `generate/` 219 modules, `core/` 171 (largest subpackages `physics/` 37, `capability/` 14, `contemplation/` 13, `cognition/` 13), `sensorium/` 59, `packs/` 48, `teaching/` 46, `chat/` 29, `formation/` 23, `workbench/` 22, `algebra/` 9, `recognition/` 7, `vault/` 5, `field/` 4. Tests 881, evals 375.
|
||||
|
||||
---
|
||||
|
||||
## 2. Finding 0-A — CORE has five different macro articulations of its own cognitive cycle, and none reconcile to the others
|
||||
|
||||
This is the central Phase 0 finding and the reason Phase 1 exists. Each articulation below is authoritative inside its own document. No document maps any of them onto any other.
|
||||
|
||||
**(1) The north star — `AGENTS.md`, 7 stages:**
|
||||
`listen → comprehend → recall → think → articulate → learn from reviewed correction → replay deterministically`
|
||||
|
||||
**(2) The live path — `AGENTS.md`, 9 steps:**
|
||||
`CognitiveTurnPipeline → tokenize / OOV policy / inject → intent classification → PropositionGraph → ArticulationTarget → deterministic realizer → telemetry/trace → reviewed teaching capture → deterministic replay/eval/calibration`
|
||||
|
||||
**(3) The three physics layers — `MIND-PHYSICS-BLUEPRINT.md`, Draft 2026-05-12:**
|
||||
`FieldState → Allocation Physics (Salience/Attention/Inhibition, ADR-0008) → Compositional Physics (Binding/Digest/Trajectory/ArticulationPlanner, ADR-0009) → Identity Physics (IdentityCheck/DriveGradientMap/ExertionMeter, ADR-0010) → Renderer (TBD, "ADR-0011 planned")`
|
||||
|
||||
**(4) The governing paradigm — `ADR-0252`, Accepted 2026-07-19, five stages:**
|
||||
`Perceive → core primitives · Comprehend → Structure-Map · Reconstruct → Quinian bootstrap · Solve → predictive processing / means-ends · Select → overlapping waves`
|
||||
|
||||
**(5) The system map — `.system-map/`, 10 bands / 33 zones:**
|
||||
`L0 algebra → L1 field → L2 vault → L3 packs → L4 recognition → L5 cognition → L6 chat-runtime → L7 teaching → L8 memory/contemplation → L9 epistemic verdicts → L10/L11 runtime+identity (unbuilt)`, plus cross-cutting bands for ingest, lexical substrate, infra, reasoning, learning, afferent, governance, tooling.
|
||||
|
||||
**Why this matters, concretely.** Articulation (3) names an entire **allocation/attention layer** — salience, attention budget, inhibition — that appears nowhere in (1), (2), or the system map's telos ordering. Either attention allocation is a real layer that the live path silently omits, or it is a 2026-05 draft that was superseded and never retracted. Phase 2 must determine which; the answer changes whether CORE is missing a layer or carrying a stale blueprint. `MIND-PHYSICS-BLUEPRINT.md` also lists a Renderer as "TBD / ADR-0011 planned" while `generate/realizer.py` has been the shipping renderer for months — that line alone is stale.
|
||||
|
||||
Similarly, articulation (4) is **ratified and governing** but is a paradigm for *problem-solving*, while (2) is a *turn pipeline*. Nothing states how ADR-0252's five stages map onto the nine pipeline steps. Phase 1 must decide whether the taxonomy is built on the pipeline, on the paradigm, on the band/zone layering, or on a reconciliation of all three — and must say explicitly which articulations the chosen spine subsumes and which it retires.
|
||||
|
||||
**Recorded for Phase 4:** an architecture with five unreconciled self-descriptions cannot answer "is anything missing" for any of them, because a component absent from one is present in another and no document adjudicates. This is a design-layer gap, not an implementation gap.
|
||||
|
||||
---
|
||||
|
||||
## 3. Finding 0-B — the system map is the best existing macro→micro artifact, and it is stale, invisible, and partially hollow
|
||||
|
||||
`.system-map/` (2026-06-09) is a 33-zone / 205-subsystem / 246-edge layered map with per-zone greppable cards, built by a four-wave multi-agent sweep with coverage critics. It is the closest thing CORE has to the artifact this assessment is meant to produce, and Phase 1 should treat it as a prior rather than start from scratch.
|
||||
|
||||
Three qualifications:
|
||||
|
||||
- **48 days stale.** Built before the entire deduction-serve arc (ADR-0256 through ADR-0265), the curriculum-serving arc, the two-grammars work, and PRs #98–#138. Its liveness labels predate all of it.
|
||||
- **Local-only and gitignored** by deliberate choice. It is not a project artifact, is invisible to every fresh clone and every agent that does not know to look, and cannot be reviewed in a PR. Whether the assessment's own output should inherit that property is a Phase 1 question.
|
||||
- **Four zones carry zero subsystems** — `comprehend-organ`, `determine-phase`, `realize-phase`, `sensorium-falsification`. These were added in the final wave as zone headers without descent. Three of the four sit directly on the serving path, which makes them the least-mapped parts of the most load-bearing region.
|
||||
|
||||
**Zone liveness rollup (2026-06-09, verify before relying on any row):**
|
||||
|
||||
| Liveness | Count | Zones |
|
||||
|---|---|---|
|
||||
| `live-serving` | 6 | L0-algebra, vocab-manifold, governance-identity-safety, comprehend-organ, determine-phase, realize-phase |
|
||||
| `live-internal` | 5 | morphology, alignment-resonance, capability, engine-state, sensorium-falsification |
|
||||
| `partial-wiring-debt` | **18** | L1-field, L2-vault, L3-packs, L4-recognition, L5-cognition, L6-chat-runtime, L7-teaching, L8-memory-contemplation, L9-epistemic-verdicts, L10-11-runtime-identity, ingest-boundary, ingest-compiler, reasoning-deductive, gsm8k-math, reliability-calibration, formation-curriculum, evals-determinism, tooling-cli-workbench-rs |
|
||||
| `spike` | 1 | core-protocol-ctp |
|
||||
| `inert` | 2 | edge-sync, sensorium-afferent |
|
||||
| `research-negative` | 1 | field-wedge-research |
|
||||
|
||||
The map's own recorded caveat is important and should be carried forward: **a zone inherits its weakest honest label**, so "partial" at zone level is compatible with most subsystems being live. Zone-level liveness must never be quoted as a system-level completion rate.
|
||||
|
||||
---
|
||||
|
||||
## 4. Finding 0-C — two load-bearing questions are open, and one of them gates the governing paradigm
|
||||
|
||||
### 4.1 The ADR-0252 §5 experiment has never returned a verdict
|
||||
|
||||
ADR-0252 is the ratified governing paradigm. Its ruling record is explicit:
|
||||
|
||||
> **NOT authorized.** The §6 build (the structure-map comprehension layer and any serving change) is **not** authorized by this ratification. It is gated on §5 returning GO.
|
||||
|
||||
§5 is a controlled experiment on exactly one empirical claim: **can Cl(4,1) geometry carry a problem's relational structure faithfully enough that `conformal_procrustes` aligns same-deep-structure problems and separates different-structure ones — driven by relations, invariant to surface attributes?** GO requires separability with margin *and* attribute-invariance *and* structure-sensitivity. A well-controlled NO-GO is defined as full credit.
|
||||
|
||||
**Status: unrun.** Two unmerged worktrees exist — `rnd/structure-mapping-experiment` @ `fc9d0c14` ("add SME feasibility research and experimental script") and `rnd/sme-experiment-v2` @ `bed29a09` ("formalize §5 experiment scaffolding — corpus extractor + single-pair probe"). Neither has reached `main`; neither reports a verdict.
|
||||
|
||||
The consequence is structural, not procedural. ADR-0252 diagnoses CORE's single architectural error as *"a **novice** comprehender (34 bespoke surface organs) bolted onto an **expert** substrate it never used for comprehension"* — and rules that the 34 surface organs keep serving until a proven replacement exists. So the diagnosis is ratified, the correction is designed, the replacement is gated on an experiment that has not run, and the thing diagnosed as the error is what serves today. **This is the highest-leverage open item Phase 0 found**, and it belongs at the top of the Phase 5 synthesis.
|
||||
|
||||
### 4.2 L10 is the single blocking node of the entire wiring-debt registry
|
||||
|
||||
All **7 OPEN** entries in `docs/audit/substrate-liveness-ratchet.md` chain to W-008 (the L10 runtime model): W-003 → recognizer-storage ADR → W-007; W-009 → W-017; W-018; each annotated "sized after W-008 commits." The 2026-05-24 master plan named this correctly — *"L10 is the load-bearing decision… wrong shape locks in wrong architecture for the rest of the cascade"* — and the spike it demanded (prove a long-lived process holds field state, vault, and session continuity across 24+ hours without leak or drift) has never been run.
|
||||
|
||||
The system map states the consequence in the sharpest available terms: field excitation and the T1 vault are **discarded on process exit**, so only the learned-recognizer layer compounds across reboots. **"One continuous life" today is many short lives sharing a checkpoint.** Given that `project-core-is-one-continuous-life` is the foundational telos, this is the largest single distance between stated purpose and built system.
|
||||
|
||||
---
|
||||
|
||||
## 5. Finding 0-D — CORE has no live gap register
|
||||
|
||||
Three registers exist; none is currently tracking the frontier.
|
||||
|
||||
- `docs/gaps.md` — **all 26 entries closed** (`[x]`). A register with nothing open is either a solved system or an abandoned instrument. It is the latter: it tracks pack/threshold gaps from the capability-ledger era and has no entry for anything in the 2026-07 arc.
|
||||
- `docs/audit/substrate-liveness-ratchet.md` — v5, 7 OPEN, all L10-blocked, untouched since ~2026-05-24.
|
||||
- `docs/analysis/` — ~130 dated per-arc lookback and ratification documents. This is where the real findings live, but it is a chronological archive, not a register: nothing aggregates it, and a finding recorded in a June lookback is discoverable only by knowing it exists.
|
||||
|
||||
**Implication for Phase 4:** the gap register this assessment produces will be the first live one, and Phase 1 should decide whether it supersedes `docs/gaps.md` or sits beside it. Two dead registers plus a live one is a worse outcome than one live one.
|
||||
|
||||
---
|
||||
|
||||
## 6. Finding 0-E — the decision corpus has outgrown its own index
|
||||
|
||||
333 files in `docs/adr/`, flat sequential numbering to ADR-0265. Status is not reliably machine-extractable: a scan for status lines yields roughly 200 `accepted` and 72 `proposed` against 333 files, with the remainder carrying malformed, absent, or prose status lines (`status follows`, `status as`, `Phase`, `Delegated`). Any claim of the form "N ADRs are Accepted" is therefore approximate, and Phase 3 must not build on it without per-file verification.
|
||||
|
||||
`INDEX-by-domain.md` is the existing mitigation and is admirably honest about its own limits — it covers only live serving, telemetry, and governance surfaces and states plainly that *"an index asserting coverage it does not have is worse than none."* It was written when the corpus was 312 files; it is now 333, so the index is already 21 files behind and its "index on mint" maintenance rule is not holding.
|
||||
|
||||
Two structural hazards are already documented and should be carried into Phase 4: the ADR-0206/ADR-0256 numbering collision that had to be resolved in prose, and the Master Blueprint's wholesale collision with Accepted ADR-0246–0252, which required a governing mapping document to adjudicate and left **numbers 0254–0261 reserved for Blueprint intents that were never materialised** — while ADR-0254 through ADR-0265 were subsequently minted for entirely different decisions. The reservation table in the governing mapping is therefore contradicted by the live registry.
|
||||
|
||||
---
|
||||
|
||||
## 7. Finding 0-F — the measured performance picture contradicts the intuitive one
|
||||
|
||||
Recorded here because Phase 4's hindrance audit will otherwise re-derive it, and because it is a clean instance of the mastery framework's warning against optimizing what shouldn't exist.
|
||||
|
||||
- **CGA dominates turn time (~73%)** — but through `cga_inner` → `geometric_product` at ~33,986 calls/turn in nearest-neighbour and salience search, **not** through the versor invariant. `versor_condition` measures **0.448 ms against a 200 ms turn — 0.22%**. An earlier assessment claimed it was "~10× the entire proof latency" by multiplying an isolated microbenchmark by a call count and comparing against a single verdict's latency; that was corrected by direct measurement (`docs/research/cga-hot-path-measurement-2026-07-25.md`).
|
||||
- **The Rust backend is off by default.** `algebra/backend.py`: pure Python is the deterministic default; Rust is opt-in via `CORE_BACKEND=rust`. Whether `core_rs` still holds bit-exact parity is **an open question, blocked** — `cargo` could not reach `static.crates.io` under the sandbox network policy on 2026-07-25.
|
||||
- **MLX is in no runtime path.** It appears in exactly two files, both under `benchmarks/`. The Neural Engine claim in `README.md` is a design aspiration explicitly disclaimed by the repo's own measured report.
|
||||
- **The deduction and curriculum serving paths are pure-Python ROBDD and never touch CGA.** FrameVerdict TTFV is 0.151 ms.
|
||||
|
||||
---
|
||||
|
||||
## 8. Finding 0-G — a prior assessment exists, and its lesson is methodological
|
||||
|
||||
`docs/research/architecture-assessment-verification-2026-07-25.md` verified an external architectural assessment against the tree and found **roughly a third of its work items falsified by the codebase**, one flagship item diagnosed at the wrong layer, and the most urgent live defect named in neither document. The falsified items were not careless — they were plausible readings that no one had checked against code.
|
||||
|
||||
Two things follow. First, the precedent validates this assessment's read-the-code rule and raises the bar: a finding without an evidence pointer is a hypothesis. Second, its §7 item 1 (the Workbench provenance falsification, where a proved deduction answer was recorded as `grounding_source = 'none'`) now appears **closed** — `runtime_contracts.md` documents the three-part registration requirement and names the pin `tests/test_workbench_deduction_provenance.py`. Phase 3 should verify rather than assume.
|
||||
|
||||
---
|
||||
|
||||
## 9. Raw material for Phase 1 — the proto card schema
|
||||
|
||||
The system map's node schema is the closest existing thing to what Phase 1 must specify, and it already anticipates several fields a naive schema would miss.
|
||||
|
||||
**Existing fields (17):** `zone_id`, `zone_name`, `layer`, `band`, `classification`, `macro_role`, `what_it_is`, `what_it_does`, `inputs`, `outputs`, `invariants`, `key_files`, `owning_adrs`, `relationships`, `subsystems`, `telos_stage`, `liveness`, plus **`honest_wrinkles`** — a field for what is true about the component that its label would otherwise hide. That field should survive into Phase 1's schema under some name; it is the structural defense against decoration.
|
||||
|
||||
**Existing liveness vocabulary (7 values, ordered):**
|
||||
`live-serving` → `live-internal` → `partial-wiring-debt` → `spike` → `inert` → `unbuilt` → `research-negative`
|
||||
|
||||
This is richer than the four-value scale proposed in the approach plan and distinguishes things that matter: serving versus merely-executing, never-built versus built-and-disconnected, and research that returned a negative result (`field-wedge-research`) — a category a naive schema would misfile as failure rather than as knowledge.
|
||||
|
||||
**What the existing schema lacks, measured against what Shay asked for:**
|
||||
|
||||
| Required dimension | Present? | Note |
|
||||
|---|---|---|
|
||||
| Philosophical intent — what the component is *for* in the cognitive model | Partial | `macro_role` / `what_it_is` are functional, not teleological |
|
||||
| Functional contract — inputs/outputs/invariants | **Yes** | Strongest part of the existing schema |
|
||||
| Design shape vs implementation detail | Partial | `key_files` only; no design-vs-built distinction |
|
||||
| Implementation status | **Yes** | The 7-value liveness vocabulary |
|
||||
| **Evidence pointer** — the test/lane/pin proving the status | **No** | Liveness is asserted, not evidenced. Highest-priority addition |
|
||||
| **Capacity** — to what degree, at what limit | **No** | "To what capacity" has no field at all |
|
||||
| Dependencies | **Yes** | `relationships`, `inputs`, `outputs` |
|
||||
| Governing ADRs | **Yes** | `owning_adrs` |
|
||||
| Known hazards | **Yes** | `honest_wrinkles` |
|
||||
| **Fitness judgment** — is this the right solution here? | **No** | Required by Phase 4; no field exists |
|
||||
|
||||
---
|
||||
|
||||
## 10. Handoff to Phase 1
|
||||
|
||||
**Deliverables:** `docs/assessment/02-layer-taxonomy.md` and `docs/assessment/03-card-schema.md`.
|
||||
|
||||
**The five decisions Phase 1 must make explicitly, with reasoning recorded:**
|
||||
|
||||
1. **Which articulation is the spine?** Choose among the turn pipeline (2), the ratified paradigm (4), the band/zone layering (5), or a reconciliation — and state which of the five articulations the choice subsumes, which it retires, and what happens to the allocation/attention layer from (3) that exists in no other articulation.
|
||||
2. **Adopt, extend, or replace the system map's schema and its 7-value liveness vocabulary?** Extension is the recommended default: the vocabulary encodes distinctions already earned. The four missing dimensions — evidence pointer, capacity, design-vs-built, fitness — must be added regardless.
|
||||
3. **What is the completeness criterion?** Per Pillar I this must be fixed before any card is filled: by what test does a layer get called complete, and by what test does a *missing* layer get identified as missing rather than as out of scope? Without this, question 2 of the charter cannot be answered rigorously.
|
||||
4. **Where do candidate layers live?** The taxonomy needs an explicit section for layers CORE's own documents do not name — beginning with whatever the always-on L10 PROCESS requires that no current zone owns, and whatever an AGI/ASI-grade system requires that CORE's telos implies but no document has articulated. This section is where charter question 2 actually gets answered.
|
||||
5. **Is the assessment output committed or local-only?** The system-map precedent is gitignored. This assessment is currently on a branch intended for a Forgejo PR. These are incompatible defaults; pick one and state why.
|
||||
|
||||
**Reference documents for Phase 1** (keep to these two plus this file; do not re-ingest the corpus): `AGENTS.md` and `docs/adr/ADR-0252-problem-solving-paradigm-consolidation.md`. The system map's zone cards are greppable at `.system-map/zones/*.md` in the **main** worktree (`/Users/kaizenpro/Projects/core`), not in `core-wt-assess` — it is gitignored and does not travel with the branch. Query `data.json` with `jq`; never slurp it.
|
||||
|
||||
**Standing cautions for every downstream phase:**
|
||||
|
||||
- Apply the sabotage test to every `live` claim. `g_args_rate` was `0.0` while a section claimed the reader read it.
|
||||
- Zone-level liveness is a weakest-link rollup and must not be quoted as a completion rate.
|
||||
- Everything in the system map is 48 days old and predates the entire deduction-serve and two-grammars arc.
|
||||
- The PR #138 fabrication findings are **measured and pinned, held for ADR + ratification** — record, never re-discover, never fix.
|
||||
238
docs/assessment/02-layer-taxonomy.md
Normal file
238
docs/assessment/02-layer-taxonomy.md
Normal file
|
|
@ -0,0 +1,238 @@
|
|||
# Phase 1 — The Layer Taxonomy
|
||||
|
||||
**Executor:** Fable 5, 2026-07-27. **Verified against:** `forgejo/main` @ `8927c563`.
|
||||
**Inputs:** `01-phase0-ground-truth.md` (the handoff), `AGENTS.md`, `ADR-0252`, `.system-map/` (2026-06-09 prior — local to the main worktree, gitignored).
|
||||
**Companion:** `03-card-schema.md` (the metadata schema every card in Phases 2–3 must fill).
|
||||
|
||||
This document fixes the decomposition of CORE — macro to micro — that Phases 2 and 3 fill with cards, and Phase 4 audits against. Per Pillar I (Semantic Rigor), the taxonomy and its completeness criteria are fixed *before* any component is judged.
|
||||
|
||||
---
|
||||
|
||||
## 0. Decisions made by this phase
|
||||
|
||||
| # | Decision | Ruling |
|
||||
|---|---|---|
|
||||
| D1 | Which articulation is the spine? | **None of the five "wins."** They answer five different questions and become five axes/attributes of one taxonomy (§1). The functional axis is the AGENTS.md north star; the structural axis is a 7+2 macro-layer grouping of the system map's 33 zones. The taxonomy's job is auditing the *mapping* between the two. |
|
||||
| D2 | Schema: adopt, extend, or replace the system map's? | **Extend.** Keep the 17 fields and the 7-value liveness vocabulary; add the four missing dimensions (evidence, capacity, design-vs-build, fitness) as orthogonal fields, not fatter enums (`03-card-schema.md`). |
|
||||
| D3 | Completeness criterion | Fixed in §6: coverage is *evidence-bearing ownership of a functional stage*, completeness is per-layer and per-stage, and "missing" is distinguished from "explicitly deferred" by the presence of a ruling. |
|
||||
| D4 | Where do candidate layers live? | The Candidate Register (§5) — four registered candidates (attention, agenda/drive, efferent action, temporal self-location), each with its telos derivation, partial existence, and the ruling it needs; plus a considered-and-not-registered line to show the boundary was examined. |
|
||||
| D5 | Committed or local-only? | **Committed** (this branch → Forgejo PR). This artifact is governance-adjacent: it will drive R&D ordering and rulings, so it must be reviewable, versioned, and visible to every future session. The system map stays local as a regeneratable *navigation index*; the assessment is the *ruled record*. Staleness is handled by discipline, not by hiding: every card stamps the SHA at which its claims were verified. |
|
||||
|
||||
---
|
||||
|
||||
## 1. D1 — the reconciliation: five articulations, five different questions
|
||||
|
||||
Phase 0's Finding 0-A: CORE describes its own cognitive cycle five different ways, and no document maps any onto any other. The resolution is not to crown one. Read closely, the five are not competing answers to one question — they are answers to **five different questions**, and the apparent conflict dissolves once each is assigned to the axis it actually describes. (This is the Third Door applied to the taxonomy itself: the trade-off "which self-description do we keep?" is not split; it is dissolved.)
|
||||
|
||||
| Articulation | The question it answers | Kind | Disposition in this taxonomy |
|
||||
|---|---|---|---|
|
||||
| **North star** (`AGENTS.md`, 7 stages) | *What is the system for?* | Functional / teleological | **Adopted as the functional axis** (§2). Candidate extensions are registered separately (§5), never silently merged into it. |
|
||||
| **Live path** (`AGENTS.md`, 9 steps) | *What happens on a turn today?* | Realized pathway | **Subsumed.** It is the serving traversal M2 → M3 → M4 → MV (§3). It becomes the spine of the M3/M4 layer cards — a pathway *through* the structure, not a decomposition *of* it. |
|
||||
| **Mind-physics blueprint** (`docs/architecture/MIND-PHYSICS-BLUEPRINT.md`, Draft 2026-05-12) | *By what mechanisms?* | Mechanism proposal | **Recommend ruling: mark historical / partially superseded**, with the per-element disposition in §1.1. Its one orphan — Allocation Physics — moves to the Candidate Register as CR-1. |
|
||||
| **ADR-0252 paradigm** (Accepted 2026-07-19) | *By what competence is problem-solving judged?* | Governing competence model | **Adopted as the governing competence model within M3.** Its five stages decompose the problem-solving pathway inside M3; its §4 conformance bar (deep structure, generalization ratio > 1) becomes a *fitness criterion* for M3 components in Phases 3–4. |
|
||||
| **System map** (`.system-map/`, 2026-06-09) | *What is built, and where?* | Structural inventory | **Adopted as the sublayer stratum.** All 33 zones are retained and regrouped under 7+2 macro layers (§3–4); the 7-value liveness vocabulary is retained unchanged. |
|
||||
|
||||
The two-axis model that results:
|
||||
|
||||
- **Functional axis** — the stages of the cognitive cycle (what the organism does).
|
||||
- **Structural axis** — macro layers → zones → components (what exists, in containment order).
|
||||
|
||||
Every structural card carries the functional stages it serves (`telos_stages`, already a system-map field). The assessment's central audit — run in Phase 2 with evidence, not here — is the **stage-coverage audit**: every functional stage must be owned by at least one live structural element (else a gap), and every structural element must serve at least one stage (else a fitness question). Neither axis can perform this audit alone; that is why neither "wins."
|
||||
|
||||
### 1.1 Disposition of the mind-physics blueprint, element by element
|
||||
|
||||
The blueprint is a Draft that was never advanced and never retracted. Its elements did not fail uniformly, so a uniform disposition would be dishonest. Recommended for ruling (this assessment records, it does not enact):
|
||||
|
||||
| Blueprint element | What became of it | Disposition |
|
||||
|---|---|---|
|
||||
| Identity Physics (IdentityCheck, ADR-0010) | Landed and hardened: wave-only fail-closed identity scoring (ADR-0244 §3, INV-32), identity manifold, identity packs | **Superseded by stronger implementations** — retire the blueprint's version of the claim |
|
||||
| Compositional Physics (Binding/Digest/Trajectory/ArticulationPlanner, ADR-0009) | Landed as the proposition-graph lineage: `PropositionGraph → ArticulationTarget → realizer` | **Superseded** — same |
|
||||
| Renderer ("TBD, ADR-0011 planned") | `generate/realizer.py` has been the shipping renderer for months | **Stale line** — retire |
|
||||
| DriveGradientMap / ExertionMeter | Never built; no successor articulation anywhere | Moves to Candidate Register **CR-2** (agenda/drive) |
|
||||
| **Allocation Physics** (SalienceOperator / AttentionOperator / InhibitionOperator, ADR-0008) | Never landed *as a layer*. Fragments exist under other names — see CR-1 | Moves to Candidate Register **CR-1** (attention/allocation) |
|
||||
|
||||
---
|
||||
|
||||
## 2. The functional axis
|
||||
|
||||
### 2.1 The ratified cycle (unchanged, from `AGENTS.md`)
|
||||
|
||||
```text
|
||||
listen → comprehend → recall → think → articulate → learn from reviewed correction → replay deterministically
|
||||
```
|
||||
|
||||
This is the canonical vocabulary for `telos_stages`. The system map's telos mapping (which zones serve which stage) is adopted as the *claimed* ownership baseline; Phase 2 verifies claims against evidence.
|
||||
|
||||
### 2.2 Claimed stage ownership (baseline — unverified until Phase 2)
|
||||
|
||||
| Stage | Primary owner (macro layer) | Supporting |
|
||||
|---|---|---|
|
||||
| listen | M2 Afferent Boundary | M1 (lexical substrate) |
|
||||
| comprehend | M3 Comprehension & Reasoning | M1 (packs, vocabulary) |
|
||||
| recall | M1 Knowledge & Memory | M3 (in-turn recall), M0 (exact CGA distance) |
|
||||
| think | M3 Comprehension & Reasoning | M0 (the medium) |
|
||||
| articulate | M4 Expression & Serving | M1 (packs), MG (governance of the surface) |
|
||||
| learn | M5 Learning & Growth | M1 (promotion target), M6 (what persists) |
|
||||
| replay | MV Verification & Evidence | *property enforced everywhere; apparatus lives in MV* |
|
||||
| *(the cycle's runner)* | M6 Continuity & Process | **unbuilt at its center** — L11 process |
|
||||
|
||||
Candidate functions (attend, want/agenda, act, temporal self-location) are deliberately **not** rows in this table. They live in §5 until ruled into the cycle or ruled out.
|
||||
|
||||
---
|
||||
|
||||
## 3. The structural axis — seven macro layers, two cross-cuts
|
||||
|
||||
Macro layers are groupings of the 33 zones, carved at the joints the system's own contracts already respect (serve boundaries, trust boundaries, the INV regime). Each gets one card in Phase 2 (`10-layer-cards/`). One-paragraph teleology for each; full intent belongs on the card.
|
||||
|
||||
**M0 — Substrate.** The physical medium: Cl(4,1) algebra, the versor invariant, the field. Per the governing mental model, the field is the *electricity*, not the intelligence — M0 supplies closure, exactness, and replayability to everything above it, and is forbidden from containing cognition-specific policy. Zones: `L0-algebra`, `L1-field`.
|
||||
|
||||
**M1 — Knowledge & Memory.** What is known, at rest: the vault (exact recall), compiled packs, the vocabulary manifold, the lexical substrate (morphology, alignment), and memory's consolidation surface. The epistemic-status regime (SPECULATIVE/COHERENT/…) governs everything here. Zones: `L2-vault`, `L3-packs`, `vocab-manifold`, `morphology`, `alignment-resonance`, `L8-memory-contemplation` *(straddles M5 — consolidation is learning in motion)*.
|
||||
|
||||
**M2 — Afferent Boundary.** World → field: the ingest gate and compiler, the sensorium's afferent track, and environmental falsification. Every entry point is a trust boundary; nothing crosses without construction-boundary normalization. Zones: `ingest-boundary`, `ingest-compiler`, `sensorium-afferent`, `sensorium-falsification`.
|
||||
|
||||
**M3 — Comprehension & Reasoning.** The wiring that thinks: recognition, the cognition pipeline, the comprehend/determine/realize organs, the deduction flagship, curriculum-grounded reasoning, the math reader, and reasoning research. ADR-0252 governs this layer's competence bar; the two-grammars frontier (reader ≠ writer) lives here. Zones: `L4-recognition`, `L5-cognition`, `comprehend-organ`, `determine-phase`, `realize-phase`, `reasoning-deductive`, `gsm8k-math`, `field-wedge-research`.
|
||||
|
||||
**M4 — Expression & Serving.** Field → world: the chat runtime, surface selection policy, the register axis, response governance, and the epistemic-verdict surface. This layer owns *serving-path truth behavior* — what the user actually reads — and is therefore where `wrong=0` lives or dies. Zones: `L6-chat-runtime`, `L9-epistemic-verdicts` *(straddles MG — verdicts are governance made visible)*.
|
||||
|
||||
**M5 — Learning & Growth.** Controlled mutation: the reviewed teaching loop, formation/curriculum, reliability calibration and earned licenses, and the capability ledger. The typed learning boundary (durable-reviewed vs provisional-typed, INV-21…24/29/30) is this layer's constitution. Zones: `L7-teaching`, `formation-curriculum`, `reliability-calibration`, `capability` *(straddles MV — the ledger is also evidence)*.
|
||||
|
||||
**M6 — Continuity & Process.** The life itself: the always-on process (unbuilt), engine-state checkpointing (built), edge-sync, the turn protocol, session continuity, async HITL. The telos ("one continuous life") is this layer's charter, and its center is the single largest distance between stated purpose and built system. Zones: `L10-11-runtime-identity`, `engine-state`, `edge-sync`, `core-protocol-ctp`.
|
||||
|
||||
**MG — Governance & Identity** *(cross-cutting)*. Identity manifold and packs, safety pack (never-swappable, fail-closed), ethics packs, refusal taxonomy, trust boundaries, the INV regime as a body of law. Cross-cutting because its writ runs everywhere; a governance mechanism that only one layer obeys is a bug. Zone: `governance-identity-safety`.
|
||||
|
||||
**MV — Verification & Evidence** *(cross-cutting)*. Evals, lanes, pinned SHAs, CLAIMS.md, replay/determinism apparatus, telemetry/trace, the CLI test surface, workbench-as-auditor. Cross-cutting for the same reason; replay is a property of every layer and an apparatus of this one. Zones: `evals-determinism`, `tooling-cli-workbench-rs`.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
W((world)) --> M2[M2 Afferent]
|
||||
M2 --> M3[M3 Comprehension & Reasoning]
|
||||
M3 <--> M1[M1 Knowledge & Memory]
|
||||
M3 --> M4[M4 Expression & Serving]
|
||||
M4 --> W2((world))
|
||||
M4 --> M5[M5 Learning & Growth]
|
||||
M5 --> M1
|
||||
M0[M0 Substrate] -.the medium.- M1 & M3
|
||||
M6[M6 Continuity & Process] -.hosts the cycle.- M2 & M3 & M4 & M5
|
||||
MG[MG Governance & Identity]:::cc -.governs all.- M4
|
||||
MV[MV Verification & Evidence]:::cc -.witnesses all.- M4
|
||||
classDef cc stroke-dasharray: 3 3;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. The complete zone mapping
|
||||
|
||||
All 33 zones, mapped once. Liveness is the map's 2026-06-09 label — **48 days stale, claimed not verified**; Phase 2 re-verifies. Flags: ⚑ = zero-subsystem zone (Phase 3 must descend it first); ✱ = straddle (noted above).
|
||||
|
||||
| Zone | Macro layer | 2026-06-09 liveness | Note |
|
||||
|---|---|---|---|
|
||||
| L0-algebra | M0 | live-serving | |
|
||||
| L1-field | M0 | partial-wiring-debt | |
|
||||
| L2-vault | M1 | partial-wiring-debt | T1 discarded on exit (→ M6) |
|
||||
| L3-packs | M1 | partial-wiring-debt | |
|
||||
| vocab-manifold | M1 | live-serving | |
|
||||
| morphology | M1 | live-internal | |
|
||||
| alignment-resonance | M1 | live-internal | |
|
||||
| L8-memory-contemplation ✱ | M1 ↔ M5 | partial-wiring-debt | consolidation straddle |
|
||||
| ingest-boundary | M2 | partial-wiring-debt | |
|
||||
| ingest-compiler | M2 | partial-wiring-debt | |
|
||||
| sensorium-afferent | M2 | inert | |
|
||||
| sensorium-falsification | M2 | live-internal | map labels it "L12" — a stratum no other document uses; flag for ruling |
|
||||
| L4-recognition | M3 | partial-wiring-debt | |
|
||||
| L5-cognition | M3 | partial-wiring-debt | |
|
||||
| comprehend-organ ⚑ | M3 | live-serving | zero subsystems mapped |
|
||||
| determine-phase ⚑ | M3 | live-serving | zero subsystems mapped |
|
||||
| realize-phase ⚑ | M3 | live-serving | zero subsystems mapped; serve seam → M4 |
|
||||
| reasoning-deductive | M3 | partial-wiring-debt | flagship; predates ADR-0256–0265 arc |
|
||||
| gsm8k-math | M3 | partial-wiring-debt | demoted to diagnostic |
|
||||
| field-wedge-research | M3 | research-negative | negative result = knowledge, not failure |
|
||||
| L6-chat-runtime | M4 | partial-wiring-debt | |
|
||||
| L9-epistemic-verdicts ✱ | M4 ↔ MG | partial-wiring-debt | |
|
||||
| L7-teaching | M5 | partial-wiring-debt | |
|
||||
| formation-curriculum | M5 | partial-wiring-debt | |
|
||||
| reliability-calibration | M5 | partial-wiring-debt | |
|
||||
| capability ✱ | M5 ↔ MV | live-internal | |
|
||||
| L10-11-runtime-identity | M6 | partial-wiring-debt | center unbuilt (L11 process) |
|
||||
| engine-state | M6 | live-internal | the built footing |
|
||||
| edge-sync | M6 | inert | |
|
||||
| core-protocol-ctp | M6 | spike | |
|
||||
| governance-identity-safety | MG | live-serving | |
|
||||
| evals-determinism | MV | partial-wiring-debt | |
|
||||
| tooling-cli-workbench-rs | MV | partial-wiring-debt | |
|
||||
| sensorium-falsification ⚑ | — | — | *(also zero-subsystem; listed once above)* |
|
||||
|
||||
Calibration to carry forward: at the **subsystem** stratum the 2026-06-09 map records 58 `live-serving` + 79 `live-internal` of 205 total (with 28 partial, 22 inert, 10 spike, 4 unbuilt, 4 research-negative). The organism is substantially more built than zone-level labels imply; the genuinely-unbuilt mass is concentrated in M6. Zone liveness is a weakest-link rollup and must never be quoted as a completion rate.
|
||||
|
||||
---
|
||||
|
||||
## 5. The Candidate Register (D4)
|
||||
|
||||
Functions or layers that **no ratified document names**, but that the telos arguably implies. Registration here asserts nothing except *this deserves a ruling*. Each entry: derivation → partial existence → ruling needed → risk if left unregistered.
|
||||
|
||||
**CR-1 — Attention / allocation (in-turn).**
|
||||
*Derivation:* any bounded cognitive system must select what to process; the blueprint articulated it (ADR-0008) and no successor document owns it; the north star is silent between "listen" and "comprehend."
|
||||
*Partial existence:* admissibility threshold and margin gates (ADR-0024/0026), rotor admissibility (ADR-0025), and salience/nearest-neighbour search — which is precisely the measured hot path (~73% of turn time through `cga_inner`/`geometric_product`, Finding 0-F). The blueprint's `InhibitionMask` appears never to have been built.
|
||||
*Ruling needed:* is attention a first-class layer, or an emergent property of admissibility that should stay distributed?
|
||||
*Risk:* the hottest path in the system — computationally and semantically — has no owner, no card, and no governance. Optimization and correctness work on it currently has nowhere to attach.
|
||||
|
||||
**CR-2 — Agenda / drive (between-turn).**
|
||||
*Derivation:* "one continuous life" must decide what to do when it is not serving. A process with nothing to want is a heartbeat, not a life.
|
||||
*Partial existence:* idle consolidation (CLOSE), read-only proposal review, the contemplation loop (flag-gated), discovery-yield telemetry, `core/epistemic_questions/`. All are *mechanisms without a chooser* — each does one thing when its flag is on; nothing ranks what matters next, and the blueprint's `DriveGradientMap`/`ExertionMeter` were never built.
|
||||
*Ruling needed:* does the L10 process own an agenda; by what policy is it governed (identity packs? curriculum priorities? operator queue?); and is agenda-formation itself subject to the typed learning boundary?
|
||||
*Risk:* L10 lands as an always-on process that idles — the telos's letter without its spirit. This is the largest conceptual absence for an AGI-grade system: everything CORE does is currently chosen by the operator.
|
||||
|
||||
**CR-3 — Efferent action.**
|
||||
*Derivation:* AGI-grade generality ordinarily implies acting on the world; CORE's telos ends at articulate/learn/replay — deliberately text-first.
|
||||
*Partial existence:* typed deterministic tool operators folded into `trace_hash` (ADR-0018); the environmental-falsification contract (ADR-0211) *explicitly forbids* motor/efferent units in v1.
|
||||
*Ruling needed:* an explicit scope declaration — is action **deferred** (like scripture content: a ruling exists) or **out of telos**? Today neither is stated anywhere, which is the gap.
|
||||
*Risk:* silent scope ambiguity. Note honestly: the alignment posture (position paper §6) is arguably *stronger* with action explicitly deferred — this register entry is about making the boundary ruled, not about advocating efferents.
|
||||
|
||||
**CR-4 — Temporal self-location.**
|
||||
*Derivation:* a continuous life experiences sequence and duration; determinism bans clocks from cognition. Both commitments are correct, and their intersection is undesigned: how does the L10 process represent "now," "before," and "how long" without breaking replay?
|
||||
*Partial existence:* session-context ordering, engine-state `turn_count`, idle ticks as pseudo-time, the Fibonacci recency constants schedule (τ_n, ADR-0242 — a constants schedule, not a clock).
|
||||
*Ruling needed:* a stance on lived time for the L10 spike — the 24h+ no-drift requirement cannot even be *stated* precisely without one.
|
||||
*Risk:* the L10 spike gets designed with an implicit, accidental answer to a question nobody asked out loud.
|
||||
|
||||
**Considered and not registered** (recorded so the boundary is visibly examined, per Pillar I):
|
||||
- *Sociality / other-minds* — multi-actor comprehension is a comprehension capacity inside M3 (the ADR-0174 pronoun hazard is its live trace), not a layer. Revisit only if the telos expands to multi-party life.
|
||||
- *Emotion / affect* — no derivation from the telos under decoding-not-generating; disposition is already carried structurally (identity packs, hedging, refusal taxonomy, register axis). Registering it would import an external architecture's expectations — precisely what Pillar II forbids.
|
||||
- *Full embodiment* — subsumed by the CR-3 ruling.
|
||||
- *Epistemic self-governance as a first-class layer* — CORE's most distinctive machinery (earned licenses, calibration, ledgers, typed refusal) exists and is live, but is split across M5/MG/MV. This is an **organizational** question, not a missing function; noted for the Phase 5 synthesis rather than registered as a candidate.
|
||||
|
||||
---
|
||||
|
||||
## 6. Completeness criteria (D3)
|
||||
|
||||
Fixed now, applied in Phases 2–4. All terms below are used in their `03-card-schema.md` senses.
|
||||
|
||||
**A functional stage is covered** iff at least one structural component owns it with liveness `live-serving` (or `live-internal` for internal-only stages such as replay) **and** at least one evidence pointer that would fail if the component were deleted. Ownership without such evidence is *claimed coverage* and is recorded as uncovered.
|
||||
|
||||
**A layer is complete** iff:
|
||||
1. every functional stage it claims is covered at the capacity its card declares (not merely "at all");
|
||||
2. every zone and component within it has a card with no mandatory field left `undetermined`;
|
||||
3. every invariant it declares has failing-when-violated enforcement — a pin that lives in a suite that actually runs, verified by the suite-count check, not by the pin's existence.
|
||||
|
||||
**A missing layer or component is identified** when any of:
|
||||
- (i) a ratified functional stage has no covered owner;
|
||||
- (ii) a governing document's mechanism has no structural home (the Allocation-Physics test);
|
||||
- (iii) a telos-implied function has no owner and no Candidate-Register entry — this clause is the AGI-grade audit, and it is why §5 exists: once registered, a candidate is a *ruling item*, not a silent absence.
|
||||
|
||||
**Deferred is not missing.** An absent capability is out-of-scope iff an explicit ruling defers it (scripture content; motor efferents in falsification v1). Absent-with-no-ruling is a gap. The cure for a gap of this kind may simply be a one-line ruling — the register makes that cheap.
|
||||
|
||||
**The system is complete** — never claimable today, stated so the target is fixed — iff the north-star cycle runs end-to-end under the M6 process with every stage covered at declared capacity, every invariant enforced, and zero `undetermined` fitness verdicts. Note what this criterion deliberately omits: benchmark scores. Per the master plan's own distinction, architectural distinctiveness is the target; benchmark wins are downstream validation.
|
||||
|
||||
---
|
||||
|
||||
## 7. Phase 2 work order
|
||||
|
||||
**Deliverable:** nine layer cards in `10-layer-cards/` — `M0-substrate.md`, `M1-knowledge-memory.md`, `M2-afferent-boundary.md`, `M3-comprehension-reasoning.md`, `M4-expression-serving.md`, `M5-learning-growth.md`, `M6-continuity-process.md`, `MG-governance-identity.md`, `MV-verification-evidence.md` — each schema-compliant per `03-card-schema.md`.
|
||||
|
||||
**Priority order** (by leverage, not ease): **M6** (telos-critical; the L10 question), **M3** (serving-path truth; the two-grammars frontier; ADR-0252's home), **M4** (what the user reads), **M5** (the learning boundary), then M1, M2, M0, MG, MV.
|
||||
|
||||
**Obligations:**
|
||||
1. Re-verify every liveness label quoted from the 2026-06-09 map before writing it into a card; stamp `verified_at` with the SHA actually inspected. The map is a prior, never a source.
|
||||
2. Run the **stage-coverage audit** (§2.2 table, with evidence this time) and record it in each card's `stage_coverage` block. This is where "is a stage uncovered" gets its verdict.
|
||||
3. Apply the sabotage test to every `live-*` claim inherited or newly made.
|
||||
4. The four ⚑ zero-subsystem zones (`comprehend-organ`, `determine-phase`, `realize-phase`, `sensorium-falsification`) are unmapped-and-load-bearing: the M3/M2 cards must scope them honestly (what is *known* vs *unmapped*) and queue them first for Phase 3 descent.
|
||||
5. Where a card touches the PR #138 fabrication findings: they are measured-and-pinned, held for ADR + ratification — record, never re-discover, never fix.
|
||||
6. Straddle zones (✱) are cited on both cards but *owned* by one (the table's left column); the owning card carries the full entry, the other a cross-reference. No double-counting in any rollup.
|
||||
192
docs/assessment/03-card-schema.md
Normal file
192
docs/assessment/03-card-schema.md
Normal file
|
|
@ -0,0 +1,192 @@
|
|||
# Phase 1 — The Card Schema
|
||||
|
||||
**Executor:** Fable 5, 2026-07-27. **Companion:** `02-layer-taxonomy.md`.
|
||||
**Applies to:** every card written in Phases 2 (`10-layer-cards/`) and 3 (`20-component-cards/`), and every claim consumed by Phase 4's registers.
|
||||
|
||||
This is D2 executed: the system map's 17-field node schema and 7-value liveness vocabulary are **extended, not replaced**. The map's vocabulary encodes distinctions it earned over 205 subsystems — serving vs merely-executing, never-built vs built-and-disconnected, research that returned a negative result. What it lacks are the four dimensions Phase 0 measured as missing: **evidence, capacity, design-vs-build, fitness**. Those arrive as orthogonal fields, not fatter enums.
|
||||
|
||||
---
|
||||
|
||||
## 1. Design principles
|
||||
|
||||
**P1 — Orthogonal axes, not fatter enums.** *Liveness* answers "is it running?"; *fitness* answers "should it be?" These are independent. The governing example: ADR-0252 condemns the 34 surface organs while ruling they keep serving until a proven replacement exists — liveness `live-serving`, fitness `superseded-in-place`. A single merged status could not express that state, and that state is load-bearing. The same separation holds for *design* (what documents articulate) versus *build* (what code does): a component can be fully designed and unbuilt (`L11` process), or built with no surviving design authority (fragments of allocation physics).
|
||||
|
||||
**P2 — Evidence or it is a claim.** Every `live-serving` / `live-internal` verdict requires at least one evidence pointer **that would fail if the mechanism were deleted** (the sabotage test). Evidence that would look identical with the mechanism absent is decoration and does not count toward liveness — this repository has produced exactly that failure (`g_args_rate` was 0.0 while prose claimed the reader read it). A card may cite decoration, but must label it as such.
|
||||
|
||||
**P3 — Stamped verification.** Every card carries `verified_at`: the commit SHA actually inspected and the date. A claim inherited from the 2026-06-09 system map without re-verification is written as `claimed (map 2026-06-09)`, never as fact. Staleness must be visible on the card, not discoverable by archaeology.
|
||||
|
||||
**P4 — Pins must run.** An invariant "enforced by a test" is only enforced if the pin lives in a suite that executes (a pin registered in no suite never runs). Citing a pin requires naming its suite and confirming the suite's membership count moved when the pin was added, or that the suite demonstrably collects it today.
|
||||
|
||||
**P5 — Identity, not value.** When a card asserts two things are the same mechanism (shared, not duplicated), the evidence must be identity-grade (same object, same call path), not equal-looking values or name-greps.
|
||||
|
||||
---
|
||||
|
||||
## 2. Field specification
|
||||
|
||||
### 2.1 Identity block
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `id` | ✓ | Stable slug; zones keep their map `zone_id` unchanged for greppability |
|
||||
| `kind` | ✓ | `layer` \| `zone` \| `component` |
|
||||
| `parent` | ✓ | Containment: component → zone → macro layer (per taxonomy §4) |
|
||||
| `verified_at` | ✓ | `<SHA> (<date>)` — the tree actually inspected (P3) |
|
||||
| `assessor` | ✓ | Who filled the card |
|
||||
|
||||
### 2.2 Teleology block *(the "philosophical understanding" dimension)*
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `philosophical_intent` | ✓ | What this element is **for** in the cognitive model — one paragraph, teleological, answerable to the north star. Distinct from `what_it_does`. If no document states an intent, write `undetermined` — that is itself a finding |
|
||||
| `telos_stages` | ✓ | Stages served, from the ratified vocabulary; candidate functions cited as `candidate:CR-n`, never bare |
|
||||
| `macro_role` | ✓ | Kept from map: the element's role in its layer's story |
|
||||
|
||||
### 2.3 Description & contract block *(kept from the map — its strongest part)*
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `what_it_is` / `what_it_does` | ✓ | As in the map: essence vs behavior, in prose |
|
||||
| `inputs` / `outputs` | ✓ | Typed where the code types them |
|
||||
| `invariants` | ✓ | Each entry: the invariant, its enforcement pin, **and the suite that runs the pin** (P4). An invariant with no running pin is recorded with enforcement `none` — a finding, not a formality |
|
||||
| `topology_role` | ✓ | The `AGENTS.md` repository-topology classification (runtime boundary / candidate compiler / reviewed data / read-only projection / demo envelope / benchmark artifact / historical note / tooling) — replaces the map's `classification` |
|
||||
|
||||
### 2.4 Design-vs-build block *(new — the design/implementation split Shay asked for)*
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `design_state` | ✓ | Where the design is articulated (doc/ADR + status) and its one-paragraph summary. `none` if the code has no surviving design authority |
|
||||
| `build_state.liveness` | ✓ | The 7-value vocabulary, unchanged: `live-serving` → `live-internal` → `partial-wiring-debt` → `spike` → `inert` → `unbuilt` → `research-negative` |
|
||||
| `build_state.key_files` | ✓ | Kept from map |
|
||||
| `build_state.evidence[]` | ✓ | See §3.2 — the field P2 is about |
|
||||
|
||||
### 2.5 Capacity block *(new — "to what capacity" now has a field)*
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `capacity.designed` | ✓ | The envelope the design claims (e.g. "any clause the grammar owns") |
|
||||
| `capacity.measured` | ✓ | The envelope evidence supports, with numbers (e.g. "reader comprehends 19 constructions; writer emits 1739; overlap 6"; "16-premise cap holds a band to ≤16 entailed cases") |
|
||||
| `capacity.ceilings` | ✓ | Known hard limits and what imposes them (`unknown` is a legal and honest value) |
|
||||
|
||||
### 2.6 Dependency & provenance block
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `relationships` | ✓ | Kept: typed edges (`feeds` / `reads` / `depends-on` / `verifies` / …) |
|
||||
| `owning_adrs` | ✓ | Kept; extended to include governing analysis/ratification docs with dates |
|
||||
|
||||
### 2.7 Judgment block *(new — feeds Phase 4)*
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `fitness.verdict` | ✓ | See §3.3 |
|
||||
| `fitness.rationale` | ✓ | Why — with evidence pointers. `undetermined` requires naming what evidence would determine it |
|
||||
| `honest_wrinkles` | ✓ | Kept from the map, deliberately: what is true about this element that its labels would otherwise hide |
|
||||
| `open_questions` | ○ | Questions the card raises but cannot answer; each tagged with who can (Phase 3 / Phase 4 / ruling) |
|
||||
|
||||
### 2.8 Layer cards only
|
||||
|
||||
| Field | Req | Notes |
|
||||
|---|---|---|
|
||||
| `stage_coverage` | ✓ | For each `telos_stage` the layer claims: covered / claimed-only / uncovered, with the evidence that decides it (taxonomy §6) |
|
||||
| `zone_roster` | ✓ | The layer's zones from taxonomy §4, with re-verified liveness |
|
||||
| `rollup_note` | ✓ | Mandatory restatement: zone liveness is weakest-link; the subsystem-level picture; never a completion rate |
|
||||
|
||||
---
|
||||
|
||||
## 3. Vocabularies
|
||||
|
||||
### 3.1 Liveness (unchanged, 7 values, ordered)
|
||||
|
||||
`live-serving` · `live-internal` · `partial-wiring-debt` · `spike` · `inert` · `unbuilt` · `research-negative`
|
||||
|
||||
`research-negative` is knowledge, not failure — a mechanism investigated and honestly refuted (field-wedge). Do not "clean it up" into `inert`.
|
||||
|
||||
### 3.2 Evidence entry shape
|
||||
|
||||
```text
|
||||
{ claim, kind: test | lane | pin | acceptance-packet | measurement | code-read,
|
||||
pointer: <path or doc + location>,
|
||||
would_fail_if_absent: yes | no | unknown }
|
||||
```
|
||||
|
||||
Only `would_fail_if_absent: yes` counts toward liveness (P2). `code-read` is legitimate evidence for *existence and wiring* (per the read-the-code doctrine) but never for *behavior* — behavior needs a test, lane, or measurement.
|
||||
|
||||
### 3.3 Fitness (new, 6 values)
|
||||
|
||||
| Verdict | Meaning |
|
||||
|---|---|
|
||||
| `fit` | Right mechanism, right owner, serving its intent |
|
||||
| `strained` | Right place, but shape or capacity no longer matches the load (e.g. an index that stopped scaling) |
|
||||
| `misplaced` | Sound mechanism, wrong owner — belongs to another layer/component whose design parameters (possibly extended) suit it |
|
||||
| `wrong-solution` | The mechanism does not serve the underlying problem; a different approach is needed |
|
||||
| `superseded-in-place` | Condemned by ruling but deliberately still serving until a proven replacement exists |
|
||||
| `undetermined` | Not yet judged — must name the deciding evidence |
|
||||
|
||||
`misplaced` and `wrong-solution` entries are the direct feed for Phase 4's hindrance audit (`31-hindrance-audit.md`); every such verdict must name its **evidence** and its **proposed better home** — and decides nothing (rulings are Shay's).
|
||||
|
||||
---
|
||||
|
||||
## 4. Card template
|
||||
|
||||
```markdown
|
||||
# <id> — <name>
|
||||
|
||||
**Kind:** <layer|zone|component> · **Parent:** <parent> · **Assessor:** <who>
|
||||
**Verified at:** `<SHA>` (<date>)
|
||||
**Liveness:** `<value>` · **Fitness:** `<verdict>` · **Topology role:** <role>
|
||||
|
||||
> <philosophical_intent — one paragraph>
|
||||
|
||||
**Telos stages:** <stages; candidates as candidate:CR-n>
|
||||
**Macro role:** <one line>
|
||||
|
||||
## What it is / What it does
|
||||
<prose, two short paragraphs>
|
||||
|
||||
## Contract
|
||||
- Inputs: …
|
||||
- Outputs: …
|
||||
- Invariants: <invariant> — pin: <test> — suite: <suite> — status: <running|none>
|
||||
|
||||
## Design vs build
|
||||
- Design: <doc/ADR + status> — <summary or `none`>
|
||||
- Build: <liveness>; key files: …
|
||||
- Evidence:
|
||||
- <claim> — <kind> — <pointer> — would-fail-if-absent: <yes|no|unknown>
|
||||
|
||||
## Capacity
|
||||
- Designed: … · Measured: … · Ceilings: …
|
||||
|
||||
## Dependencies & provenance
|
||||
- Relationships: …
|
||||
- Owning ADRs / governing docs: …
|
||||
|
||||
## Judgment
|
||||
- Fitness rationale: …
|
||||
- Honest wrinkles: …
|
||||
- Open questions: … (→ Phase 3 / Phase 4 / ruling)
|
||||
```
|
||||
|
||||
Layer cards append the `stage_coverage` table, `zone_roster`, and `rollup_note` (§2.8).
|
||||
|
||||
---
|
||||
|
||||
## 5. Anti-patterns (each has already occurred in this repository's history)
|
||||
|
||||
1. **Exists ≠ live.** A module's presence proves nothing about execution. (Dormant readback rules; `explain.py`.)
|
||||
2. **Threaded ≠ behavioral.** A flag reaching a call site is not the flag changing output; pair every wiring claim with a behavior sweep. (The ignored-flag lesson.)
|
||||
3. **Green ≠ meaningful.** A measurement that would look identical with the mechanism removed measures nothing. (`g_args_rate` = 0.0.)
|
||||
4. **Pinned ≠ enforced.** A pin in no suite never runs; check the count moved. (The suite-membership lesson.)
|
||||
5. **Documented ≠ true.** Roughly a third of a plausible external blueprint was falsified by reading the code (Finding 0-G). Documents are testimony; code is evidence.
|
||||
6. **Zone label ≠ completion rate.** Weakest-link rollups understate; quoting them as progress metrics misleads in both directions.
|
||||
7. **Acceptance ≠ comprehension.** Never quote an acceptance metric without its identity metric; use faithful / fabricating / refused — never two buckets.
|
||||
|
||||
---
|
||||
|
||||
## 6. Depth allocation for Phase 3
|
||||
|
||||
Full-depth descent (every component gets a full card): serving-path truth behavior (M4 + the M3 serve seam), the two-grammars frontier and the four ⚑ zero-subsystem zones, the learning boundary (M5's INV regime), M6's built footings and unbuilt center, and everything carrying a `fitness` verdict other than `fit`.
|
||||
|
||||
Cite-and-summarize (short card pointing at the acceptance packet): closed arcs with ratified acceptance evidence (deduction bands v1–v6, cohesion ADR-0241/0242, GSM8K sealed-lane machinery), unless Phase 2 surfaces a contradiction — in which case they promote to full depth.
|
||||
|
||||
The rule's purpose is honesty about attention, not economy for its own sake: a short card must say *why* it is short and what would reopen it.
|
||||
108
docs/assessment/04-phase2-findings.md
Normal file
108
docs/assessment/04-phase2-findings.md
Normal file
|
|
@ -0,0 +1,108 @@
|
|||
# Phase 2 — Findings, Corrections, and the Phase 3 Handoff
|
||||
|
||||
**Executor:** Opus 5, 2026-07-27. **Verified against:** `forgejo/main` @ `8927c563`.
|
||||
**Deliverable:** nine layer cards in `10-layer-cards/` (M0, M1, M2, M3, M4, M5, M6, MG, MV), each schema-compliant per `03-card-schema.md`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Corrections forced by re-verification
|
||||
|
||||
Phase 2's charter was to verify rather than inherit. Three inherited claims did not survive. Recording them prominently, because the *pattern* matters more than any single correction: **every falsified claim came from a document, and every falsification came from reading code.**
|
||||
|
||||
### 1.1 The L11 always-on process is BUILT (Phase 0 Finding 0-C was wrong)
|
||||
|
||||
`chat/always_on.py::run_continuous` (279 lines), `chat/always_on_daemon.py::run_daemon` (195 lines, single-instance lock, SIGINT/SIGTERM-wired stop, load-time identity guard), and CLI command `core always-on` (`core/cli.py:244`) all exist. They landed **2026-06-14** — five days *after* the `.system-map/` snapshot of 2026-06-09 that declared "no forever entrypoint exists." Phase 0 inherited the map's claim.
|
||||
|
||||
The real M6 finding is different and more actionable: the process is built, has a complete falsifiable soak harness (`evals/l10_always_on`, four predicates with `*_holds`/`*_bites` pairs), **has never been run to a recorded long-horizon artifact**, and is **enforced by no suite at all**.
|
||||
|
||||
### 1.2 Candidate CR-1 (attention/allocation) is not a missing layer — it is a live, undocumented one
|
||||
|
||||
Phase 1 registered attention as a possibly-missing function whose blueprint articulation (ADR-0008) had no successor. **The mechanism is live on the serving path.** `generate/stream.py::_attention_candidates` (`:255-263`) runs `SalienceOperator().compute(...)` then `AttentionOperator(inhibition_threshold).plan(...)`, and the resulting `allowed_indices` are intersected with language candidates at `:329` — replacing them outright at `:331` when language candidates are absent. `use_salience` defaults **True** (`core/config.py:35`).
|
||||
|
||||
Note also that `generate/salience.py` *composes* `core.physics.salience.SalienceOperator` (imported as `CurvatureSalienceOperator`) rather than duplicating it — so this is a layering, not the duplicate-implementation hazard it first appeared to be. That correction is recorded here rather than propagated as a false finding.
|
||||
|
||||
**CR-1 is therefore re-characterized:** not "does attention exist?" but "the mechanism that gates every generation step, and consumes ~73% of turn time through `cga_inner`, has no owning ADR, no card, and no layer in any ratified articulation." The blueprint's `InhibitionMask` class exists in `core/physics/inhibition.py` but is imported only by `core/physics/__init__.py` — the live inhibition is a scalar threshold, not the operator. The gap is **governance and documentation, not capability** — which makes it cheaper to close and easier to have missed.
|
||||
|
||||
### 1.3 Candidate CR-2 (agenda/drive) — the components exist; the claim survives and sharpens
|
||||
|
||||
Phase 1 recorded `DriveGradientMap` and `ExertionMeter` as never built. Both are constructed in `chat/runtime.py` (`:716`, `:714`). But:
|
||||
|
||||
- **`DriveGradientMap` is constructed and never read.** No site reads `self._drive_map`. Deleting it changes no output. By the sabotage test this is **decoration** — and it is the cleanest instance the assessment found.
|
||||
- **`ExertionMeter` is exercised** (`record` at `:2912`, `fatigue` at `:2913`) but its output flows only into `drive_summaries` and `fatigue_index` — telemetry. **Fatigue gates no decision.**
|
||||
|
||||
So CR-2's substance is confirmed and stated more precisely: CORE has drive and exertion *objects* and no *chooser*. Nothing ranks what to do next. The mechanisms exist as instrumentation; the function does not exist as agency.
|
||||
|
||||
---
|
||||
|
||||
## 2. The stage-coverage audit
|
||||
|
||||
The taxonomy's central instrument, run with evidence. A stage is **covered** only where a live component owns it *and* carries evidence that would fail if the mechanism were deleted.
|
||||
|
||||
| Stage | Verdict | Owner | Basis |
|
||||
|---|---|---|---|
|
||||
| listen/ingest (text) | **covered** | M2 | `inject` on the live path, closure-checked at the gate |
|
||||
| listen/ingest (non-text) | **uncovered** | M2 | `sensorium/` (59 modules) imports nowhere on the serving path; no projection heads |
|
||||
| comprehend | **covered, narrowly** | M3 | Deduction bands `wrong=0` over 18,000 cases; reader spans 19 constructions vs a 1739-construction writer, fabricates on 22 |
|
||||
| recall | **covered** | M1 | Exact CGA recall live; fabrication-control lane pinned |
|
||||
| think/reason | **covered** | M3 | ROBDD entailment 716/716 against an independent oracle |
|
||||
| articulate | **covered** | M4 | Live serving, ratified `wrong=0` lanes, typed refusal, negation now representable |
|
||||
| learn from reviewed correction | **covered** | M5 | Two pinned loop-closure lanes prove the single reviewed path |
|
||||
| replay deterministically | **covered** | MV | 11 SHA-pinned lanes, CI-failing on drift |
|
||||
| *(the cycle's runner)* | **claimed-only** | M6 | Process exists; no measured horizon, no suite-enforced pin, continuity flags default off |
|
||||
|
||||
**Seven of nine covered; one uncovered by deliberate scoping; one — the runner — claimed-only.** The single uncovered *stage* is non-text ingest. The single unproven *layer* is the one the telos names.
|
||||
|
||||
---
|
||||
|
||||
## 3. Cross-cutting findings
|
||||
|
||||
**F-1 — Built-and-off is the dominant pattern, and it is unaccounted.** Across layers, substantial machinery exists behind flags that default `False`: `unified_ingest`, `curriculum_serving_enabled`, `ask_serving_enabled`, `verified_serving_enabled`, `consolidate_determinations`, `review_pending_proposals`, `review_derived_close_proposals`, `auto_contemplate`, `auto_proposal_enabled`, `vault_promotion_enabled`, `persist_session_state`, `strict_identity_continuity`, `identity_wave_gate`, `identity_action_surface`, `realizer_grounded_authority`, `accrue_realized_knowledge`, `estimation_enabled`. Only `deduction_serving_enabled` is ratified ON. Each default is individually defensible; **no document states the set**, nor what evidence would flip any of them. This is the largest single lever in the system and it has no register.
|
||||
|
||||
**F-2 — Enforcement lags capability, most where it matters most.** M6's soak pins run in no suite; M1's no-approximate-recall prohibition has no verified failing pin; MG's cross-cutting writ has no bypass pin. In each case the *mechanism* is sound and the *guarantee that it stays sound* is doctrinal rather than mechanical. MV's hand-curated suite tuples cannot detect an orphaned pin — the highest-leverage single fix the assessment has found.
|
||||
|
||||
**F-3 — The strongest in-repo standard is not applied where exposure is highest.** M5's formation pipeline declares six trust boundaries, content-addressed in and out, no floats in hashed payloads, no pickle, an audit record for every rejection. M2 — the boundary facing untrusted user text in production — has no comparable declared table. The bar exists; it has not been carried to the hotter surface.
|
||||
|
||||
**F-4 — Documentation debt is now load-bearing, not cosmetic.** `chat/always_on_daemon.py` has no ratifying ADR while the governing ADR-0146 explicitly *rejected* the daemon shape. The live attention mechanism has no ADR. The `MIND-PHYSICS-BLUEPRINT` renderer line is stale. ADR-0252's "34 surface organs" does not reproduce (**18** `resolve_promotable_*`, all in `generate/derivation/`). When the record contradicts the code, every downstream reasoner inherits the error — as Phase 0 did.
|
||||
|
||||
**F-5 — What is excellent, stated plainly.** The typed learning boundary (M5) dissolves the autonomy-versus-safety trade-off rather than splitting it. The selection-not-rewrite surface discipline (M4) preserves the honest artifact even when it is not served. The non-hardening invariant (M1) structurally forbids an axiom flag. Fail-closed unknown lane shapes and NON-CANONICAL run stamping (MV) are the habits of a system that expects to be wrong. Phase 5 should carry these forward as the standard other layers are measured against, not merely as an audit's polite paragraph.
|
||||
|
||||
---
|
||||
|
||||
## 4. Layer verdict summary
|
||||
|
||||
| Layer | Liveness (re-verified) | Fitness | The one-line reason |
|
||||
|---|---|---|---|
|
||||
| M0 Substrate | `live-serving` | `fit` | Does one thing without compromise; optimization has twice aimed at the wrong function |
|
||||
| M1 Knowledge & Memory | `partial-wiring-debt` | `fit` | Exactness and typed standing reinforce each other; what compounds vs resets is invisible in the name |
|
||||
| M2 Afferent Boundary | `partial-wiring-debt` / `inert` | `strained` | 59 modules reach no serving path; the trust-boundary bar exists elsewhere in-repo |
|
||||
| M3 Comprehension & Reasoning | `partial-wiring-debt` | `strained` + `superseded-in-place` | Expert substrate, novice reader — ratified diagnosis, unrun acceptance gate |
|
||||
| M4 Expression & Serving | `live-serving` | `strained` | Excellent discipline; surface precedence accreted one arm per capability |
|
||||
| M5 Learning & Growth | `partial-wiring-debt` | `fit` / `strained` | Best-executed idea in CORE, 24×–73× under-fed |
|
||||
| M6 Continuity & Process | `partial-wiring-debt` | `strained` | Built further than reported, proven less than assumed, enforced by nothing |
|
||||
| MG Governance & Identity | `live-serving` | `fit` / `strained` | Alignment as structure; enforcement gate off and unauthorized |
|
||||
| MV Verification & Evidence | `partial-wiring-debt` | `strained` | Right instincts; coverage is a curation artifact with a hole at the worst spot |
|
||||
|
||||
No layer is `wrong-solution`. Two candidate `wrong-solution` findings are deferred to Phase 4 with evidence: the Wilson/replay independence basis in M5's licensing, and `DriveGradientMap` as constructed decoration.
|
||||
|
||||
---
|
||||
|
||||
## 5. Handoff to Phase 3
|
||||
|
||||
**Deliverable:** component cards in `20-component-cards/`, per `03-card-schema.md` §6 depth allocation.
|
||||
|
||||
**Descent order:**
|
||||
1. **The four ⚑ zero-subsystem zones**, all load-bearing, three on the serving path: `comprehend-organ` (→ `core/comprehension_attempt/`, 6 modules), `determine-phase` (→ `generate/determine/`, 8 modules), `realize-phase` (→ `generate/realize/` + `realizer.py` + `realizer_guard.py`), `sensorium-falsification` (→ `sensorium.environment.falsification`).
|
||||
2. **M6's built half** — `engine_state/`, `chat/always_on.py`, `chat/always_on_daemon.py`, and the `evals/l10_*` harnesses. Establish what the soak *would* prove if run.
|
||||
3. **M3's derivation organs** — resolve the 34-vs-18 count against ADR-0252's diagnosis. This determines whether the debt is being paid down or was measured differently.
|
||||
4. **M4's surface-selection arms** — enumerate every arm and its precedence; this is the input to the Phase 4 Third-Door question.
|
||||
5. **The CR-1 attention mechanism** — `generate/salience.py`, `generate/attention.py`, `core/physics/{salience,attention,inhibition}.py`. It gates every generation step and owns the hot path; it needs a card regardless of how the layer question is ruled.
|
||||
|
||||
**Verification obligations (non-negotiable):**
|
||||
- Stamp `verified_at` with the SHA actually inspected. The system map is a prior; 2026-06-09 is stale by two major arcs.
|
||||
- Apply the sabotage test to every `live-*` claim. `DriveGradientMap` shows the failure mode is present in this tree, not hypothetical.
|
||||
- For every invariant cited, name its pin **and** confirm the suite that runs it. Several Phase 2 cards had to record `suite: none`.
|
||||
- Distinguish *imported* from *constructed* from *read* from *gating a decision*. All four appear in this tree and only the last is load-bearing.
|
||||
|
||||
**Standing:** PR #138's fabrication findings are measured-and-pinned, held for ADR + ratification — record, never re-discover, never fix.
|
||||
|
||||
**Open items for Phase 4 seeded by Phase 2:** the flag-default register (F-1); the orphaned-pin meta-check (F-2); M2 adopting formation's trust-boundary table (F-3); the ADR/code contradictions (F-4); Wilson/replay independence in licensing; `DriveGradientMap` deletion; surface-precedence as a declarative table; the `evals/l10_*` suite assignment.
|
||||
68
docs/assessment/05-phase3-findings.md
Normal file
68
docs/assessment/05-phase3-findings.md
Normal file
|
|
@ -0,0 +1,68 @@
|
|||
# Phase 3 — Findings, Corrections, and the Phase 4 Handoff
|
||||
|
||||
**Executor:** Fable 5, 2026-07-27. **Verified against:** `forgejo/main` @ `8927c563`.
|
||||
**Deliverable:** eight component cards in `20-component-cards/` — the four ⚑ zero-subsystem zones (`comprehend-organ`, `determine-phase`, `realize-phase`, `sensorium-falsification`) plus `always-on-process`, `derivation-organs`, `surface-selection`, `attention-allocation`.
|
||||
|
||||
The assessment's correction chain continued into its third phase, in the same direction every time: **documents (including this assessment's own earlier phases) overstate or understate; code decides.** Phase 2 corrected Phase 0; Phase 3 corrects Phase 2 twice and the system map three more times.
|
||||
|
||||
---
|
||||
|
||||
## 1. Corrections
|
||||
|
||||
**C-1 — Three map `live-serving` labels demoted to `live-internal`.** `comprehend-organ` (`core/comprehension_attempt/` — imported by neither serving entrypoint; its consumers are the flag-gated ask path, proposal review, and eval lanes), `determine-phase`, and `realize-phase` (both reachable only through `_accrue_in_turn` / `idle_tick`, gated by flags that default False). **No default-config serving turn touches any of the three.** The map called all three live-serving; the serving path they were presumed to sit on runs through `proof_chain`/`curriculum_surface`/the realizer instead.
|
||||
|
||||
**C-2 — Phase 2's M6 card overstated ephemerality.** "T1 vault and field discarded on exit by design" is the *default-config* posture. Under the daemon, `persist_session_state=True` activates **Shape B+** persistence ("restored bit-exactly", `chat/always_on.py:9`; sites at `chat/runtime.py:893–952`). The residency mechanism exists, is opt-in, and is daemon-forced; what remains open is its exact coverage and horizon proof. Banner added to the M6 card.
|
||||
|
||||
**C-3 — Phase 2's M4 card understated the resolver.** `core/cognition/surface_resolution.py` (494 lines) *is* a declared-precedence resolver for the pipeline seam (served-bytes-wins base selection, canonical-first truth path, abstention, hedge admissibility, substrate folds). The accretion concern survives only for the upstream composer arms in `chat/runtime.py`. The Third-Door candidate refines from "create a resolver" to "extend the existing resolver's pattern upstream." Banner added to the M4 card.
|
||||
|
||||
**C-4 — The 34-vs-18 organ discrepancy is resolved: basis mismatch, not paydown.** At the ADR-0252 ratification commit itself (`1ccef491`) the tree already had exactly **18** `resolve_promotable_*` entry organs across 33 tree entries. "34" plausibly counted modules (~32–33 including support modules), never entry organs. Consequences: no consolidation has occurred since ratification, and the governing ADR's headline number has no stated basis — a one-sentence amendment fixes it.
|
||||
|
||||
**C-5 — A prior verification document is contradicted at this SHA.** `architecture-assessment-verification-2026-07-25.md` §2 claims `accrue_realized_knowledge` "is enabled by the production L10 process." `CONTINUOUS_LIFE_CONFIG_FLAGS` is exactly `{persist_session_state, consolidate_determinations, strict_identity_continuity}` — accrual is **not** in it.
|
||||
|
||||
---
|
||||
|
||||
## 2. New findings
|
||||
|
||||
**F-6 — The continuous life may consolidate an empty set.** The daemon forces the *consolidator* on (`consolidate_determinations`) but not the *accruer* (`accrue_realized_knowledge`) — and `realize_comprehension`, the only turn-path writer of realized facts, sits behind the accrual flag. As coded, the always-on life's Step-D learning loop may have nothing to work on unless facts are seeded some other way. Either the flag set is incomplete or the dormancy is intended; **neither reading is documented.** This is the sharpest single new finding of Phase 3: the two halves of the lived learning loop are gated by different flags and only one is forced.
|
||||
|
||||
**F-7 — The allocation layer landed, mutated, without governance.** ADR-0008's Allocation Physics is live in composed form: physics curvature kernel → generation-facing `SalienceOperator` → `AttentionOperator` scalar-threshold plan → candidate intersection in the walk, with the salience `budget` fed back so attention *self-narrows* across a walk (`stream.py:637`). `InhibitionMask` is decoration (imported by `__init__` only, never constructed on any path). The tuned constants (`top_k=16`, `threshold=0.3`) have no recorded derivation. CR-1's ruling question is now precise: own the flag, the two constants, the feedback loop, and the mask's disposition — a one-page ADR.
|
||||
|
||||
**F-8 — Two naming traps are load-bearing.** `comprehend-organ` is a *math setup router*, not the chat comprehension organ (that is `generate/meaning_graph/reader.py` — the two-grammars reader, 19 constructions, fabricating on 22, measured/pinned/held). `generate/realize/` (knowledge realization) vs `generate/realizer.py` (surface realization) are unrelated mechanisms with confusable names. Both traps would misdirect a grep-first investigation; both cards carry permanent disambiguation.
|
||||
|
||||
**F-9 — REALIZE is sound; its ceiling is its feeder.** The realize phase holds whatever the reader hands it, at honest SPECULATIVE standing, with full derivation provenance. The truth-defect class is entirely upstream in the reading — consistent with fix-upstream doctrine and with holding the #138 fixes for a serving-truth ADR. A defensive option for that ADR: refuse to *hold* a reading whose construction lies outside the reader's verified inventory, converting inventory truth into a mechanical gate at the holding boundary.
|
||||
|
||||
**F-10 — A silent-failure pinhole in a typed layer.** `_accrue_in_turn` wraps the reader/DETERMINE/REALIZE chain in a broad guard that converts any exception into a no-op accrual with no telemetry. Honest as a backstop; invisible as a failure signal.
|
||||
|
||||
---
|
||||
|
||||
## 3. Phase 4 seed list (consolidated from Phases 2–3)
|
||||
|
||||
**Hindrance-audit candidates (`31-hindrance-audit.md`):**
|
||||
1. Wilson/replay independence in license counting — possible `wrong-solution` in the counting basis; 21/25 bands affected (Phase 2).
|
||||
2. `DriveGradientMap` constructed-never-read; `InhibitionMask` exported-never-constructed — deletion candidates (mastery step 2).
|
||||
3. The empty-string refusal at the public `str` boundary — truth constructed, then discarded (M4/surface-selection).
|
||||
4. Composer-arm precedence: extend `surface_resolution`'s declared-precedence pattern upstream (refined Third Door).
|
||||
5. The F-6 flag-set incoherence (consolidator without accruer).
|
||||
6. Tuned constants without derivation at the semantic center (`top_k=16`, `threshold=0.3`, `admissibility_margin=0.4` has a recorded derivation — the contrast is instructive).
|
||||
7. M2 lacking formation's trust-boundary table (F-3).
|
||||
8. ADR/code contradictions: unowned daemon vs ADR-0146's Shape-A rejection; ADR-0252's unreproducible "34"; the 2026-07-25 doc's accrual claim (C-5).
|
||||
|
||||
**Gap-register candidates (`30-gap-register.md`):**
|
||||
1. The #138 fabrications — pre-labeled *measured & pinned, fix held for ADR + ratification*.
|
||||
2. Reader inventory 19-wide vs writer 1739 (close fabrications before widening, per standing ruling).
|
||||
3. No suite runs any `l10`/`always_on` pin; no soak artifact; no nightly (and the local-first/Mac-runner doctrine makes "nightly" itself need a ruling).
|
||||
4. Non-text ingest uncovered; sensorium's 59 modules disconnected; no entry criterion.
|
||||
5. No flag-default register (F-1: 17 flags, only deduction ON, no document states the set — now nuanced by the daemon's forced trio).
|
||||
6. No orphaned-pin meta-check (F-2 — the mechanism by which M6's harness went unscheduled).
|
||||
7. Curriculum query-scoping (the ADR-0264 §4.1 blocker); the absent curriculum serve ledger.
|
||||
8. CR-2 agenda/drive: mechanisms without a chooser (confirmed at component depth); CR-3 efferent ruling; CR-4 temporal stance.
|
||||
9. ADR-0252 §5 experiment unrun — the top open item, now with stakes note: GSM8K's demotion mispriced the experiment, but the paradigm governs *all* future comprehension.
|
||||
10. The Candidate Register items needing one-line rulings vs real design work — separate them.
|
||||
|
||||
**Verification obligations for Phase 4:** the registers decide nothing — every entry carries evidence pointers and names its deciding authority (ruling vs ADR vs mechanical fix). Rank by leverage, not ease, per the AGENTS.md protocol. Where a hindrance names a better home, name the evidence too.
|
||||
|
||||
---
|
||||
|
||||
## 4. Note on method, for the record
|
||||
|
||||
Three phases in, the correction ledger reads: Phase 0 inherited a stale map claim (always-on "unbuilt"); Phase 2 caught it by reading code, then itself overstated ephemerality and understated the resolver; Phase 3 caught both by reading deeper, and demoted three more map labels. **Nothing in this chain was caught by re-reading documents.** The assessment's value is exactly proportional to how much code each phase actually read — which is the empirical vindication, inside the assessment itself, of `AGENTS.md` protocol step 1 and the mastery framework's demand that constraints trace to verifiable ground.
|
||||
83
docs/assessment/10-layer-cards/M0-substrate.md
Normal file
83
docs/assessment/10-layer-cards/M0-substrate.md
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
# M0 — Substrate
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-serving` · **Fitness:** `fit` · **Topology role:** runtime boundary
|
||||
|
||||
> The medium, not the mind. Per the governing mental model, the field is the *electricity* and the intelligence is in the wiring — M0's entire job is to be a substrate so well-behaved that everything above it can be exact, replayable, and locatable when wrong. Cl(4,1) was chosen for one reason: it is the minimal algebra in which every conformal transformation is a versor, so every cognitive operation is algebraically closed by construction rather than by special-case handling. M0 must contain no cognition-specific policy; the moment it does, the substrate has started deciding.
|
||||
|
||||
**Telos stages:** recall, think/reason, articulate, backbone-runtime, replay/determinism (as the medium of all)
|
||||
**Macro role:** Supplies closure, exactness, and bit-level replayability to every layer above.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`algebra/` (9 modules) implements Cl(4,1) — geometric product, versor apply, reverse, `cga_inner`, the versor condition — with `algebra/backend.py` dispatching between a pure-Python implementation and an opt-in Rust one. `field/` (4 modules) holds `FieldState` and propagation. `core/physics/` (37 modules) sits atop the algebra with the wave/identity/quantity machinery.
|
||||
|
||||
The one non-negotiable invariant is `versor_condition(F) = ‖F·reverse(F) − 1‖_F < 1e-6`, checked at the injection gate and preserved across every transition, which is only ever the sandwich product `F' = V·F·reverse(V)`. Closure of field transitions is owned **solely** by `algebra/versor.py::_close_applied_versor`; no other site may repair it. `AGENTS.md` draws the bright line explicitly: *semantic anchoring* (allowed at named construction boundaries, preserves the condition by construction, expresses a relation in the cognitive model) versus *drift repair* (forbidden — restoring an invariant a prior function should have preserved). Naming may not disguise the distinction.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** injected multivectors from M2's gate; versors from packs and operators.
|
||||
- **Outputs:** `FieldState`, exact CGA distances, closure verdicts.
|
||||
- **Invariants:**
|
||||
- `versor_condition(F) < 1e-6` — never weakened to make code or tests pass — pin: algebra suite (15 files) — status: **running**.
|
||||
- No normalization/closure/repair outside owned boundaries; forbidden in `generate/stream.py`, `field/propagate.py`, `vault/store.py`, logging/telemetry — pin: `tests/test_third_door_cohesion.py` (AST-pinned off-serve quarantine).
|
||||
- Physics hot ops must import from `algebra.backend`, not direct pure-algebra modules — pin: `tests/test_physics_backend_dispatch_hygiene.py`.
|
||||
- f64 wave-residual pins stay on the Python product (Rust f32 GP is not parity-safe for 1e-9 pins).
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** `docs/Yellowpaper.md` (formal Cl(4,1) specification), `docs/position_paper.md` §3, ADR-0241/0242 (wave-field, Fibonacci operators), ADR-0245 (CGA unification), ADR-0243 (lifecycle). `AGENTS.md` carries the invariants as law.
|
||||
- **Build:** `live-serving`. Key files: `algebra/versor.py`, `algebra/backend.py`, `field/state.py`, `field/propagate.py`, `core/physics/`.
|
||||
- **Evidence:**
|
||||
- Versor closure is enforced and the algebra suite runs it (15 test files) — pin+suite — would-fail-if-absent: **yes**.
|
||||
- Pure Python is the deterministic default; Rust is opt-in via `CORE_BACKEND=rust` — code-read — `algebra/backend.py:1-9` — would-fail-if-absent: **yes**.
|
||||
- `versor_condition` costs **0.448 ms against a 200 ms turn — 0.22%** — measurement — `docs/research/cga-hot-path-measurement-2026-07-25.md` — would-fail-if-absent: **n/a (a measurement)**.
|
||||
- CGA is ~73% of turn time via `cga_inner` → `geometric_product` at ~33,986 calls/turn in nearest-neighbour and salience search — measurement — same document.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** algebraically closed over the full conformal group; exact recall as a geometric fact.
|
||||
- **Measured:** closure holds at the 1e-6 gate across all serving lanes; Rust backend records `core_rs import: False`, `using_rust(): False`, status `python_fallback` in the benchmark run.
|
||||
- **Ceilings:** bit-exact determinism forbids JIT float reassociation (blocks MLX fusion without a ratifying ADR) and rules out bf16 (ε ≈ 7.8e-3, four orders coarser than the gate). Rust f32 geometric product is not parity-safe for the f64 residual pins.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**feeds** → every layer; **constrained-by** → MV (bit-exact parity pins, `core-rs/tests/test_crdt_hash_parity.rs`); **owning ADRs:** ADR-0241, ADR-0242, ADR-0243, ADR-0245, ADR-0180.
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| *(medium for)* recall / think / articulate / backbone / replay | **covered** | Closure enforced by a running suite; exact CGA distance is the recall primitive; determinism pinned bit-exact |
|
||||
|
||||
**Zone roster:** `L0-algebra` (live-serving, confirmed), `L1-field` (partial-wiring-debt, not individually re-verified).
|
||||
|
||||
**Rollup note:** weakest-link rollup. M0 is the most solid layer in CORE and the least in doubt.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`.** M0 does exactly one thing and does it without compromise. Notably, the position paper's claim that drift-correction machinery "was deleted because it only exists when the algebra is not closed" is architecturally coherent with `AGENTS.md`'s bright-line rule — the doctrine and the code tell the same story here, which is not true everywhere else in this assessment.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **The optimization target has been repeatedly misidentified**, twice in documented history. Both an external assessment and a hardware blueprint aimed at `versor_condition` (0.22%) rather than `cga_inner`/`geometric_product` (~73%). The lesson generalizes past M0: the invariant is the *most visible* thing in the layer, so it attracts attention the profile does not justify.
|
||||
- **`CORE_BACKEND=rust` is off by default and nobody currently knows whether parity holds.** The verification attempt on 2026-07-25 was blocked — `cargo` could not reach `static.crates.io` under the sandbox network policy. This is an open question with a stated blocker, which is the honest state, but it means a written-and-shipped Rust kernel sits unused and unverified.
|
||||
- The measured hot path (`cga_inner` in salience/nearest-neighbour search) is M0 *compute* driven by an M3/CR-1 *mechanism* that has no owning card or ADR. Optimization work here has nowhere to attach architecturally — see the Candidate Register CR-1 revision in the M2/M3 cards and Phase 4.
|
||||
|
||||
**Open questions:**
|
||||
- Does `core_rs` still hold bit-exact parity, and should Rust become the default? (→ ruling; blocked on network access)
|
||||
- Does the ~73% CGA cost in salience search justify an algorithmic change rather than a backend change? (→ Phase 4)
|
||||
90
docs/assessment/10-layer-cards/M1-knowledge-memory.md
Normal file
90
docs/assessment/10-layer-cards/M1-knowledge-memory.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# M1 — Knowledge & Memory
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` · **Fitness:** `fit` · **Topology role:** runtime boundary + reviewed pack data
|
||||
|
||||
> What is known, at rest — and the standing at which it is known. M1 is where CORE's rejection of statistical retrieval becomes concrete: a recall hit is a *geometric fact*, not a probabilistic suggestion, because the CGA inner product **is** Euclidean distance in conformal embedding. Exactness here is not a performance trade-off; it is what makes a recalled result verifiable at all. M1 also carries the epistemic-status regime, so it stores not just claims but their position in the revision graph.
|
||||
|
||||
**Telos stages:** recall (primary); comprehend and articulate (as the source of packs and vocabulary)
|
||||
**Macro role:** Holds and returns knowledge exactly, with its standing attached.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`vault/` (5 modules) is exact CGA recall — `best_match = argmax_i {Q · V_i}` — with no ANN index, no HNSW, no embedding ranking, no tunable similarity threshold; the runtime invariant forbids introducing any. `packs/` (48 modules) holds compiled runtime language packs plus governance/style modality packs. `vocab/`, `morphology/`, `alignment/` form the lexical substrate. `core/contemplation/` (13 modules) is memory in motion — idle consolidation, the wave seam, hypothesis-versus-evidence reconstruction.
|
||||
|
||||
The epistemic surface (ADR-0021) types every claim by its **position in the revision graph**, not by source trust: `COHERENT` (fits current field geometry), `CONTESTED` (incoherent with a reviewed claim, review pending), `SPECULATIVE` (proposed, admissible only as candidate), `FALSIFIED` (incoherent under accumulated evidence — *retained*, eligible for inversion). Two properties deserve emphasis. First, the **non-hardening invariant**: no reviewed claim ever becomes unrevisable; no `final`/`frozen`/`axiom`/`permanent` flag exists or may be added. The only closure in the architecture is mathematical (`versor_condition`), never epistemic. Second, the **curator rule**: status transitions are computed from coherence with the reviewed field, and the curator's *only admissible reasoning is geometric* — source credentials, popularity, and institutional position are explicitly inadmissible as justification.
|
||||
|
||||
The dual-pack serve boundary (ADR-0253 / INV-33) separates compiled runtime packs (`packs/data/<pack_id>/`, the only serve authority, loaded via `packs.compiler.load_pack`) from source/draft language trees (`packs/he`, `packs/grc`, `packs/en`, …), which serve entrypoints must not import as Python packages.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** reviewed teaching applies and certified promotions (M5), compiled pack artifacts, session writes.
|
||||
- **Outputs:** exact recall hits with `epistemic_status`, lexical entries, mounted packs, vocabulary manifold positions, consolidated derived facts with `Derivation` provenance.
|
||||
- **Invariants:**
|
||||
- Exact recall only — no cosine similarity, ANN, HNSW, or embedding ranking as runtime memory truth — pin: architectural invariants; **status: enforced by doctrine + review**, mechanical pin not individually re-verified at this SHA.
|
||||
- INV-21 (vault-writer allowlist), INV-22/23 (unmarked → SPECULATIVE), INV-24 (recall categorization; user-facing evidence COHERENT-only), INV-29 (only `vault/store.py` transitions status) — pins: `tests/test_architectural_invariants.py`, `tests/test_epistemic_invariants.py`.
|
||||
- Non-hardening — pin: `tests/test_epistemic_invariants.py`.
|
||||
- INV-33 dual-pack serve boundary — pin: `tests/test_pack_draft_serve_boundary.py` (static AST + process import probe).
|
||||
- Morphology rows load only from compiled packs, carrying `language`, `source_pack_id`, `source_span` — provenance-complete or not loaded — pin: `tests/test_observed_he_morph_constraint_v0.py` (four-arm ablation).
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** ADR-0021 (epistemic surface), ADR-0054 (exact CGA recall indexing/batching), ADR-0253 (dual-pack boundary), ADR-0180 (Delta-CRDT sharded vault), ADR-0243 (wave-field lifecycle), ADR-0241 (holographic standing-wave storage — **off-serve quarantined**), `docs/position_paper.md` §3.3.
|
||||
- **Build:** `partial-wiring-debt`. Key files: `vault/store.py`, `packs/compiler.py`, `core/contemplation/`, `core/physics/holographic_vault.py` (quarantined).
|
||||
- **Evidence:**
|
||||
- Exact recall is the runtime primitive; `recall`/`recall_batch` in `vault/store.py:224,296` — code-read — would-fail-if-absent: **yes**.
|
||||
- Pack-grounded and teaching-grounded composers serve real turns — lanes — `evals/domain_contract_validation`, `evals/fabrication_control` (phantom endpoints, cross-pack non-bridges, sibling collapses all refuse), both pinned in `CLAIMS.md` — would-fail-if-absent: **yes**.
|
||||
- Off-serve quarantine of the wave/holographic modules is AST-pinned — pin — `tests/test_third_door_cohesion.py` — would-fail-if-absent: **yes**.
|
||||
- Five Tier-1 domains hold ratified packs with zero open gaps — `CLAIMS.md`.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** exact, verifiable recall over a curated manifold, with standing attached to every claim.
|
||||
- **Measured:** 5 ratified domains (2 `reasoning-capable`, 3 `audit-passed`); packs across en / he / grc / el; `docs/gaps.md` shows all 26 historical coverage gaps closed.
|
||||
- **Ceilings:** **the vocabulary manifold is finite and curated by design.** Extending to a new domain requires constructing pack vocabulary, establishing coherence with reviewed claims, and passing an eval lane. The position paper states the honest limit: whether this scales to the breadth of human knowledge is an open question; whether it can be done without confabulation is not. T1 vault contents are discarded on process exit (ADR-0146; see M6).
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**feeds** → M3 (recall, packs, vocabulary), M4 (lexicon, register packs); **written-by** → M5 (the single reviewed path); **built-on** → M0 (CGA distance is the recall primitive); **residency-gated-by** → M6.
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| recall | **covered** | Exact CGA recall serving live; fabrication-control lane pinned |
|
||||
| comprehend / articulate (as source) | **covered** | Compiled packs are the only serve authority (INV-33 pinned) |
|
||||
| *(consolidation — straddles M5)* | **covered, flag-gated** | CLOSE consolidation is proof-gated and defaults OFF |
|
||||
|
||||
**Zone roster:** `L2-vault`, `L3-packs`, `vocab-manifold` (live-serving), `morphology` (live-internal), `alignment-resonance` (live-internal), `L8-memory-contemplation` ✱ (straddles M5; owned here, cross-referenced there).
|
||||
|
||||
**Rollup note:** weakest-link rollup. `vocab-manifold` is confirmed live-serving; the layer label understates it.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`.** The exactness commitment and the epistemic-status regime are mutually reinforcing: exact recall makes a hit verifiable, and typed standing makes it *interpretable*. The non-hardening invariant is a genuinely unusual and, in this assessor's reading, correct design choice — most systems acquire an axiom flag eventually, and forbidding it structurally is what keeps the revision graph honest.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **The knowledge that compounds and the knowledge that resets are different sets.** Packs and reviewed corpora persist; the T1 vault does not survive process exit. So "memory" in M1 means two very different things depending on tier, and the distinction is invisible in the layer name. This is the M6 residency question seen from the other side.
|
||||
- `FALSIFIED` claims are **retained**, not deleted — correct and easy to misread as clutter. Any future cleanup pass must not treat falsified rows as dead data.
|
||||
- The holographic standing-wave vault (ADR-0241) is real, uses the INV-21 writer, and is **hard-quarantined off-serve** by AST pin. It is a substantial built capability that no serving path may touch — legitimate research quarantine, but worth Phase 4 attention as capacity that exists and cannot be used.
|
||||
- The exact-recall prohibition (no ANN/cosine/HNSW) is stated in `AGENTS.md` as law; I did not individually verify a mechanical pin that would *fail* if someone added a cosine ranker. Doctrine-enforced-by-review is weaker than doctrine-enforced-by-test. Flagged for Phase 3.
|
||||
|
||||
**Open questions:**
|
||||
- Is there a failing-when-violated pin for the no-approximate-recall invariant, or is it review-enforced? (→ Phase 3)
|
||||
- Should the holographic vault's quarantine have an exit criterion? (→ ruling)
|
||||
- Does the identity-divergence curriculum still bypass formation's gates? (→ Phase 3, from M5)
|
||||
90
docs/assessment/10-layer-cards/M2-afferent-boundary.md
Normal file
90
docs/assessment/10-layer-cards/M2-afferent-boundary.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# M2 — Afferent Boundary
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` (text) / `inert` (non-text) · **Fitness:** `strained` · **Topology role:** runtime boundary
|
||||
|
||||
> World → field. M2 is where untrusted reality becomes admissible structure, and it is the only layer whose failure mode is *contamination* rather than error: everything downstream inherits whatever M2 admits. Its philosophical charge is that admission is an act with consequences — the gate does not merely parse, it *vouches*. Every entry point is therefore a trust boundary, and normalization is permitted here precisely because this is a declared construction boundary, not a repair site.
|
||||
|
||||
**Telos stages:** listen/ingest
|
||||
**Macro role:** Admits input into the field under a closure-preserving construction boundary, or refuses it.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`ingest/gate.py::inject` is the live text path: it converts tokens into field excitation and is one of the explicitly **allowed normalization boundaries** in `AGENTS.md` — permitted because injection is construction, not drift repair. `core_ingest/` (7 modules) is the ingest compiler. `ingest/` itself is small (2 modules). OOV policy is applied at the same seam via `packs.OOVPolicy`.
|
||||
|
||||
`sensorium/` (59 modules — the largest non-test, non-generate package after `core/`) is the afferent track for non-text modalities, plus `sensorium.environment.falsification` (ADR-0211), a deterministic replay surface comparing expected against actual `ObservationFrame` evidence with a closed two-verdict set (`SUPPORTED` / `FALSIFIED`). Neither verdict promotes anything to reviewed memory or mutates packs, vault, identity, or policy.
|
||||
|
||||
**Verified at this SHA: `sensorium` is imported by neither `chat/runtime.py` nor `core/cognition/pipeline.py`.** The afferent track is off the serving path entirely — consistent with the map's `inert` label for `sensorium-afferent`, and consistent with the position paper's honest statement that vision, audio, and motor modalities are *planned, not built*: the `ProjectionHead` protocol supports them architecturally; the projection heads do not exist.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** raw user text; (designed) non-text modality frames.
|
||||
- **Outputs:** injected `FieldState` satisfying `versor_condition < 1e-6`, OOV decisions, `ObservationFrame` evidence, falsification verdicts.
|
||||
- **Invariants:**
|
||||
- Every accepted field state satisfies the versor condition **at the gate** — pin: algebra/ingest suites — status: running.
|
||||
- Normalization is allowed **only** at declared construction boundaries (`ingest/gate.py` is named explicitly); drift repair is forbidden and may not be disguised by naming.
|
||||
- Falsification bench forbids raw pixels/PCM/event streams/byte payloads/actuator traces in traces; forbids motor-efferent units in v1; forbids learned latents as substrate; forbids probabilistic confidence or tolerance thresholds in verdicts; forbids `generate/*` dependencies and vault mutation — pin: sensorium suite (21 files) — status: running.
|
||||
- Trust-boundary defaults: explicit opt-in for arbitrary execution; reject unsafe paths before filesystem access; no hidden background execution.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** ADR-0007 (ingest → `CandidateGeometricPressure` → `FieldState`), ADR-0211 (environmental falsification), ADR-0243 §multimodal ingress, `AGENTS.md` allowed-normalization-boundaries rule, formation's six trust boundaries as the adjacent standard (M5).
|
||||
- **Build:** text path `partial-wiring-debt`; non-text `inert`.
|
||||
- **Evidence:**
|
||||
- `inject` is the live serving entry — code-read — `chat/runtime.py:115 from ingest.gate import inject` — would-fail-if-absent: **yes**.
|
||||
- Sensorium is absent from both serving entrypoints — code-read (negative) — would-fail-if-absent: **n/a; this is evidence of absence**.
|
||||
- `sensorium` suite exists with 21 test files — code-read — `core/cli_test.py` — would-fail-if-absent: **yes** (the falsification bench is genuinely exercised).
|
||||
- `unified_ingest=False` (`core/config.py:243`) — measurement — the unified ingest path is built and **off by default**.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** any modality admitted through a closure-preserving, trust-bounded gate.
|
||||
- **Measured:** text only, in production. Non-text afferent machinery is substantial (59 modules) and exercised by its own suite, but reaches no serving path.
|
||||
- **Ceilings:** no projection heads exist for vision/audio/motor. The falsification bench is deliberately v1-closed (two verdicts, no probabilistic confidence, no tolerance thresholds) — a scoping decision, not a defect.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**feeds** → M0 (field excitation), M3 (tokens, OOV decisions); **reads** → M1 (packs, OOV policy, vocabulary); **verified-by** → MV.
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| listen/ingest (text) | **covered** | `inject` is on the live path and gate-checks closure |
|
||||
| listen/ingest (non-text) | **uncovered** | Sensorium imports nowhere on the serving path; no projection heads exist |
|
||||
|
||||
**Zone roster:** `ingest-boundary`, `ingest-compiler`, `sensorium-afferent` (inert, **confirmed** by import evidence), `sensorium-falsification` ⚑ (live-internal; zero subsystems mapped).
|
||||
|
||||
**⚑ Zero-subsystem zone.** `sensorium-falsification` is the fourth unmapped zone and the only one outside M3. What is known: it corresponds to `sensorium.environment.falsification` under ADR-0211 with a closed verdict set and an explicit forbidden-list, and its suite runs. What is unmapped: internal decomposition and the boundary against the rest of `sensorium/`. The system map labels it layer **"L12"** — a stratum no other CORE document uses. That label should be either adopted deliberately or dropped; it currently exists only in a local, gitignored artifact.
|
||||
|
||||
**Rollup note:** weakest-link rollup, and here it is genuinely informative rather than misleading: M2's text half serves and its non-text half does not.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained`.** The text gate is sound and correctly privileged as a construction boundary. The strain is proportional: **59 modules of afferent machinery reach no serving path**, which is the largest quantity of built-and-disconnected code the assessment has located. That is not automatically waste — a deliberately staged capability awaiting projection heads is legitimate — but it is a large standing bet whose entry criterion is nowhere stated.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- The single most consequential asymmetry in this layer: **M5's formation pipeline declares six explicit trust boundaries with content-addressed inputs and outputs and an audit record for every rejection; M2 — the boundary that actually faces untrusted user text in production — has no comparable declared table.** Both are "the place the world gets in," and only one of them has a published contract of that rigor. This is a Phase 4 item and, in this assessor's reading, the strongest single hindrance-audit lead outside M6: the standard exists in-repo and has not been applied where the exposure is highest.
|
||||
- `sensorium-falsification`'s "L12" layer label exists in one local artifact and no ratified document. Minor, but it is exactly how a phantom stratum enters an architecture.
|
||||
- `unified_ingest` is built and off; like M5's learning flags, the gap between built and on is undocumented.
|
||||
- The falsification bench's explicit prohibition on motor/efferent units in v1 is the **only** place in the corpus that takes a position adjacent to Candidate CR-3 (efferent action) — and it is a scoping constraint on one bench, not a system-level ruling. CR-3 remains unruled.
|
||||
|
||||
**Open questions:**
|
||||
- Should M2 adopt formation's trust-boundary table format? (→ Phase 4 / ruling)
|
||||
- What is the entry criterion for the sensorium track reaching serving? (→ ruling)
|
||||
- Adopt or drop the "L12" label (→ Phase 1 taxonomy amendment or ruling)
|
||||
93
docs/assessment/10-layer-cards/M3-comprehension-reasoning.md
Normal file
93
docs/assessment/10-layer-cards/M3-comprehension-reasoning.md
Normal file
|
|
@ -0,0 +1,93 @@
|
|||
# M3 — Comprehension & Reasoning
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` · **Fitness:** `strained` (with one `superseded-in-place` sub-region) · **Topology role:** runtime boundary
|
||||
|
||||
> The wiring that thinks. Per the governing mental model — the field is electricity, the intelligence is in the wiring — M3 is where CORE's intelligence actually lives. It takes what M2 admitted and what M1 knows, and produces a decided proposition: what is the case, what follows, what cannot be determined. Its governing competence standard is ADR-0252: comprehension must grasp *deep relational structure* that subsumes a family of problems, not surface features of one.
|
||||
|
||||
**Telos stages:** comprehend, think/reason (primary); recall (in-turn)
|
||||
**Macro role:** Turns admitted input into a decided, evidence-bearing proposition graph — or a typed refusal.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
M3 spans recognition (`recognition/`, teaching-derived structural recognizers via anti-unification), the cognition pipeline (`core/cognition/`, 13 modules, `pipeline.py` at 1381 lines), the comprehension attempt/router (`core/comprehension_attempt/`: `classify.py`, `failure_family.py`, `router.py`, `proposal.py`), the determine phase (`generate/determine/`: `determine.py`, `consolidate.py`, `estimate.py`, `estimation_license.py`, `render.py`), the realize phase (`generate/realize/`, `generate/realizer.py`), the deduction flagship (`generate/proof_chain/` — ROBDD tautology engine under six ratified bands), curriculum-grounded reasoning, the GSM8K math reader (demoted to diagnostic), and the field-wedge research zone (a recorded negative result).
|
||||
|
||||
Operationally, the live path enters M3 after ingest: intent classification → `PropositionGraph` construction → determination or entailment → `ArticulationTarget`. Deduction serving is **ratified ON** (`deduction_serving_enabled=True`, `core/config.py:397`); curriculum serving is **OFF** (`curriculum_serving_enabled=False`, `:412`), as are `ask_serving_enabled` and `verified_serving_enabled`. The deduction engine is `generate/proof_chain/entail.py` — canonicalize `(P1 & … & Pn) → Q`, ask `is_tautology` over a ROBDD. It is a decision procedure, not a derivation: `EntailmentTrace` carries an outcome, a reason, and five canonical BDD node keys (opaque hashes). **There are no intermediate proof steps in the object**, which is why "the renderer drops the proof chain" was previously diagnosed at the wrong layer — nothing is dropped because nothing exists to drop. Multi-step articulation is engine work.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** injected `FieldState`, tokenized/OOV-policed text, mounted packs and vocabulary (M1), session context.
|
||||
- **Outputs:** `PropositionGraph` (now carrying `GraphNode.negated`, ADR-0265), `ArticulationTarget`, `EntailmentTrace` / `Determined` / `Undetermined`, typed refusals, recognizer outcomes, `ContractAssessment`.
|
||||
- **Invariants:**
|
||||
- **INV-30** — open-world `determine()` constructs only `Determined(answer=True)` or refuses; never asserts `False` — pin: `tests/test_architectural_invariants.py` — suite: present in the test tree; **suite membership unverified at this SHA** (flagged).
|
||||
- **INV-31** — closed-world `FrameVerdict` cannot reach the open-world runtime (transitive import containment + single construction allowlist + typed refusal) — pin: `tests/test_architectural_invariants.py`.
|
||||
- **INV-34** — cognition-pipeline failures are typed, never silent; a `None` `ContractAssessment` is itself a violation; unresolvable referents refuse rather than fill; no PASSTHROUGH in the intent ratifier — pin: `tests/test_linguistic_governance_phases.py`.
|
||||
- Kernel no-new-legacy — new derivation must consume `ProblemFrame`/`KernelFacts`; new raw-prose regex requires an explicit `LEGACY_EXCEPTION` — pin: `tests/test_kernel_no_new_legacy_derivation_surfaces.py`.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** **ADR-0252 (Accepted, governing)** — the problem-solving paradigm and its §4 conformance bar; ADR-0251 (halt bespoke per-case regex work; the prohibition governing every math-reader increment); ADR-0249/0250 (reader→Hamiltonian compiler); ADR-0256–0261 (six deduction bands); ADR-0262/0264 (curriculum); ADR-0265 (negation in the proposition graph); ADR-0243 (wave-field lifecycle); ADR-0142 (epistemic taxonomy).
|
||||
- **Build:** `partial-wiring-debt`.
|
||||
- **Evidence:**
|
||||
- Deduction serving decides real arguments end-to-end, `wrong=0` across all splits — lane — `evals/deduction_serve/report.json`, pinned SHA `0b461a5a…` in `CLAIMS.md` — would-fail-if-absent: **yes**.
|
||||
- Propositional entailment scored against an independent truth-table oracle, 716/716 correct, `wrong=0`, `refused=0` — lane — `evals/deductive_logic/report.json`, pinned `97a23094…` — would-fail-if-absent: **yes**.
|
||||
- `deductive` suite exists and carries 20 test files — code-read — `core/cli_test.py` — would-fail-if-absent: **yes**.
|
||||
- 25 sealed bands / 18,000 cases / `wrong=0`; capability index breadth 13, `wrong_total=0` — measurement — `chat/data/deduction_serve_ledger.json`, `evals/capability_index/baseline.json` (confirmed mechanically 2026-07-25).
|
||||
- **Structure-mapping (the ADR-0252 §6 correction) — not built.** `conformal_procrustes` exists in `core/physics/dynamic_manifold.py` but is off-serving and unwired to comprehension. would-fail-if-absent: **no**.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** comprehension that grasps deep relational structure, one canonical structure subsuming a family (generalization ratio > 1).
|
||||
- **Measured — the two-grammars result, and it is the sharpest number in the assessment.** Reader and writer inventories measured against each other (PR #138, `c69f9948`): the **writer emits 1739 constructions; the reader comprehends 19; the overlap is 6**, all set-theoretic. The reader **fabricates on 22 more** — `every dog is a mammal` → `member(every_dog, mammal)`; `Given: furthermore; p implies q; p.` reaching served output with `furthermore` recited back as a premise. *These are measured and pinned, held for ADR + ratification — recorded here, not re-discovered and not fixed.* Prior prose claiming the reader "reads" its arguments was falsified: `g_args_rate` was `0.0` throughout.
|
||||
- **Ceilings:** curriculum bands are capped at ≤16 entailed cases by the 16-premise compilation cap (ADR-0264 §4.1), so **no curriculum band can earn SERVE until query-scoping lands** — an engineering blocker, not a content one. Deduction's ROBDD decides but cannot narrate multi-step derivations. The GSM8K sealed holdout (1,319 cases) has never been opened against a parser with sufficient coverage.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**reads** → M1 (packs, vocabulary, vault recall), M0 (field, CGA); **receives** ← M2; **feeds** → M4 (articulation and serve seam), M5 (refusals and comprehension failures become proposal candidates); **verified-by** → MV; **governed-by** → MG.
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| comprehend | **covered, narrowly** | Deduction bands decide real arguments with `wrong=0`; but the reader spans 19 constructions against a 1739-construction writer and fabricates on 22 — coverage is real and thin |
|
||||
| think/reason | **covered** | ROBDD entailment, 716/716 against an independent oracle; idle consolidation climbs to deductive closure under proof-gating |
|
||||
| recall (in-turn) | **covered** | Exact CGA recall via M1 |
|
||||
|
||||
**Zone roster:** `L4-recognition`, `L5-cognition`, `comprehend-organ` ⚑, `determine-phase` ⚑, `realize-phase` ⚑, `reasoning-deductive`, `gsm8k-math`, `field-wedge-research`.
|
||||
|
||||
**⚑ Zero-subsystem zones — honest scoping.** Three of the four unmapped zones live here and all three sit on the serving path. What is *known* at this SHA: `comprehend-organ` corresponds to `core/comprehension_attempt/` (6 modules: classify, failure_family, router, proposal, model); `determine-phase` to `generate/determine/` (8 modules incl. consolidate, estimate, estimation_license, render); `realize-phase` to `generate/realize/` + `generate/realizer.py` (306 lines) + `generate/realizer_guard.py`. What is *unmapped*: their internal decomposition, per-component liveness, and the boundary between `comprehension_attempt` and `L5-cognition`. **These are Phase 3's first descent targets.**
|
||||
|
||||
**Rollup note:** weakest-link rollup; not a completion rate. M3's deduction region is genuinely strong (ratified, evidenced, `wrong=0`); its comprehension region is measurably narrow. A single layer label cannot express that split, which is why the stage table above separates them.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained`, with the 18 derivation organs `superseded-in-place`.**
|
||||
|
||||
The strain is precisely what ADR-0252 named: an expert substrate driven by a novice reader. The ADR is *ratified*, its diagnosis stands, its §6 correction is designed — and its §5 acceptance gate **has never returned a verdict** (two unmerged worktrees, `rnd/structure-mapping-experiment` and `rnd/sme-experiment-v2`; the latter's tip is "formalize §5 experiment scaffolding"). So the condemned mechanism keeps serving under an explicit ruling that it should, until a proven replacement exists. That is `superseded-in-place`, and the schema's separation of liveness from fitness is what lets this card state it without contradiction.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **The "34 surface organs" count does not reproduce.** At this SHA there are **18** `resolve_promotable_*` functions, all in `generate/derivation/`. ADR-0252 cites 34. Either the count used a different basis, or consolidation has occurred since ratification. The discrepancy is unresolved and matters, because 34→18 would be evidence the debt is already being paid down — or evidence the ADR's diagnosis was calibrated against something else. Phase 3 should resolve it rather than repeat either number.
|
||||
- The single load-bearing empirical claim of the *governing* paradigm ADR is unresolved, and a well-controlled NO-GO is defined as full credit — so the experiment is cheap to finish and expensive to leave open. This is the highest-leverage open item in the assessment.
|
||||
- Deduction's strength and comprehension's narrowness are easy to conflate. `wrong=0` across 18,000 cases is a statement about the *decision procedure*, not about how much English CORE can read.
|
||||
- `field-wedge-research` is `research-negative` — a mechanism honestly refuted. It is knowledge and must not be tidied into `inert`.
|
||||
|
||||
**Open questions:**
|
||||
- Run ADR-0252 §5 to a verdict (→ ruling; highest leverage)
|
||||
- Reconcile the 34-vs-18 organ count (→ Phase 3)
|
||||
- Descend the three ⚑ serving-path zones (→ Phase 3, first)
|
||||
- Is multi-step proof articulation wanted, given it is engine work with a real soundness surface? (→ ruling)
|
||||
92
docs/assessment/10-layer-cards/M4-expression-serving.md
Normal file
92
docs/assessment/10-layer-cards/M4-expression-serving.md
Normal file
|
|
@ -0,0 +1,92 @@
|
|||
# M4 — Expression & Serving
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-serving` · **Fitness:** `strained` · **Topology role:** runtime boundary
|
||||
|
||||
> Field → world. M4 owns the only bytes a user ever sees, which makes it the layer where the project's central discipline — `wrong=0` or refuse — either holds or fails. Everything upstream can be correct and M4 can still serve a falsehood by selecting the wrong surface; everything upstream can refuse and M4 must render that refusal honestly rather than filling it. Its philosophical charge is that *saying* is a distinct act from *knowing*, governed separately.
|
||||
|
||||
**Telos stages:** articulate (primary); replay/determinism (trace emission)
|
||||
**Macro role:** Selects, governs, decorates, and emits the served surface; emits the telemetry that makes the turn auditable.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`chat/runtime.py` (3364 lines) is the layer's centre of mass, with `core/cognition/pipeline.py` (1381) wrapping it for the full cognitive turn, and `generate/realizer.py` + `generate/surface.py` producing the articulation. Around them: the register axis (`chat/register_variation.py`, `chat/register_substantive.py`), response governance (`core/response_governance/`: `govern_response`, `shape_surface`), grounding composers (`chat/pack_grounding.py`, `chat/teaching_grounding.py`, `chat/deduction_surface.py`, `chat/curriculum_surface.py`), the refusal surface (`chat/refusal.py`), safety and ethics checks, and telemetry (`chat/telemetry.py`, `chat/verdicts.py`).
|
||||
|
||||
The load-bearing behavior is **surface selection**, contracted in `runtime_contracts.md`: `surface = determination_surface` when a turn determined an answer over realized knowledge; `= [approximate] estimate` under a genuine SERVE license; `= _UNKNOWN_DOMAIN_SURFACE` when the unknown-domain gate fires; `= articulation_surface` otherwise. `walk_surface` and `articulation_surface` are both *always* retained as evidence — a determination or a gate is a **selection**, never a rewrite. This is the layer's best design property: the honest artifact survives even when it is not what was served.
|
||||
|
||||
Three distinct surfaces exist for three distinct invariants, and conflating them is a documented hazard: `surface` (what the user read), `walk_surface` (manifold/token-walk evidence), and `hash_surface` (the register-invariant truth-path capture that `compute_trace_hash` folds). Because `TurnEvent` carries the served surface but never `hash_surface`, **`trace_hash` is not reconstructable from telemetry** — a contract, not a defect. Verify a hash by replaying the pipeline, never by rebuilding it from a turn record.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** `ArticulationTarget` / `PropositionGraph` from M3, determination or entailment results, mounted register/safety/ethics/identity packs (M1, MG), session context.
|
||||
- **Outputs:** `ChatResponse` (`surface`, `walk_surface`, `articulation_surface`), `TurnEvent`, `TurnVerdicts`, `trace_hash`, `CognitivePipelineRecord`, `grounding_source`, `epistemic_state`.
|
||||
- **Invariants:**
|
||||
- Register must not move `trace_hash` (ADR-0069 inv C) — pin: register lane tests — status: running (register suites present).
|
||||
- Unknown-domain gate honoured — the realizer's fallback must not override the gate's stub — pin: contract tests in the cognition suite.
|
||||
- Realizer slot-type guard / C1 coherence floor (ADR-0075), with a **documented** exemption for deduction's quoted templates at both guard sites (`chat/runtime.py:2371-2375`, `:3007-3010`).
|
||||
- Grounding-source registration is a three-part atomic change (enum + `enum-snapshot.json` + UI badge contract) — pin: `workbench-ui/enumCoverage.test.ts` fails the build on divergence; `tests/test_workbench_deduction_provenance.py`.
|
||||
- A live `/chat/turn` with a trace hash but no `status="recorded"` pipeline record fails before journal append; pre-widening rows must show `missing_evidence`, never a synthesized green pipeline.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** `docs/specs/runtime_contracts.md` (the frozen surface-selection policy, the `trace_hash` contract, audit-ledger R7); ADR-0069/0071/0075/0077 (register axis and guards); ADR-0206 (response-governance bridge); ADR-0039 (audit completeness); ADR-0153 (trace-hash back-stamp); ADR-0254 (grounded-open hedge arm); ADR-0265 (negation reaches the surface); ADR-0024/0025/0026 (typed refusal, rotor and margin admissibility).
|
||||
- **Build:** `live-serving`.
|
||||
- **Evidence:**
|
||||
- Deduction serving is ratified ON and reaches served output — measurement — `core/config.py:397 deduction_serving_enabled=True`; verified by execution 2026-07-25 (`_run_chat_turn` on a modus-ponens prompt returns the entailment surface) — would-fail-if-absent: **yes**.
|
||||
- Surface selection honours the unknown-domain gate — contract + pin — `runtime_contracts.md` §"Unknown-domain gate honour" — would-fail-if-absent: **yes**.
|
||||
- Provenance falsification (a proved deduction answer recorded as `grounding_source='none'`) is **closed** — the registration requirement and its pin are now documented in `runtime_contracts.md` §"Grounding-source registration" — would-fail-if-absent: **yes**.
|
||||
- Negation reaches the surface (ADR-0265, merged `8927c563`): under `realizer_grounded_authority`, "evidence does not support truth" and "evidence supports truth" previously served **byte-identical** surfaces — would-fail-if-absent: **yes**.
|
||||
- Register decoration does not move `trace_hash` — pin — would-fail-if-absent: **yes**.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** every served surface is either grounded, disclosed, or an honest refusal; `wrong=0` or refuse.
|
||||
- **Measured:** `wrong=0` holds across every ratified serving lane (deduction 18,000 cases / 25 bands; curriculum physics 32/32; deductive logic 716/716). Writer-side construction inventory: **1739 constructions emitted**. Refusal is typed and carries a machine-readable reason plus per-step rejection evidence.
|
||||
- **Ceilings:** a real refusing turn still returns `surface == ""` with `refusal_reason == ""` through `ChatRuntime.respond()`'s `str` contract — the typed evidence is **unread between the raise site and the public return**. The plumbing exists on `CognitiveTurnResult` and `compute_trace_hash`; materialisation awaits a future ADR. Curriculum serving is OFF; `ask` and `verified` serving are OFF.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**receives** ← M3; **reads** → M1 (packs, register, lexicon), MG (safety, ethics, identity); **feeds** → M5 (correction capture), MV (telemetry, trace, pipeline record), M6 (checkpoint at turn boundary); **governed-by** → MG (safety is fail-closed and never swappable).
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| articulate | **covered** | Live serving with ratified `wrong=0` lanes; typed refusal; disclosed estimates; negation now representable |
|
||||
| replay/determinism (emission side) | **covered, with a known asymmetry** | `trace_hash` is deliberately register-invariant and deliberately not reconstructable from `TurnEvent`; audit-ledger R7 closed by deferred emission at the serve boundary |
|
||||
|
||||
**Zone roster:** `L6-chat-runtime` (partial-wiring-debt per map → **revised to `live-serving` at layer level**: the serving path demonstrably executes end-to-end), `L9-epistemic-verdicts` ✱ (straddles MG; owned here, cross-referenced there).
|
||||
|
||||
**Rollup note:** the map's `partial-wiring-debt` label on `L6-chat-runtime` reflects internal wiring debt (dormant modules, unmaterialised refusal strings), not a non-serving path. The layer serves; parts of it are unwired. Weakest-link rollup again understating the serving reality.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
> **Phase 3 refinement (C-3, `05-phase3-findings.md`):** "no single place states the precedence order" is half-wrong — `core/cognition/surface_resolution.py::resolve_surface` (494 lines) is a declared-precedence resolver for the pipeline seam. The accretion concern survives only for the upstream composer arms in `chat/runtime.py`; the Third-Door candidate refines to *extending the existing resolver's pattern upstream*. Full arm inventory: `20-component-cards/surface-selection.md`.
|
||||
|
||||
**Fitness: `strained`.** M4's *architecture* is among the strongest in CORE — the selection-not-rewrite discipline, the three-surface separation, the fail-closed typing, the atomic enum/UI coupling. The strain is that its surface-selection policy has accreted one arm per capability (determination, estimate, unknown-domain, deduction, curriculum, hedge, register decoration, logos-morph override), each individually contracted, with no single place that states the precedence order as an executable rule. `runtime_contracts.md` documents it in prose across several sections; `chat/runtime.py` implements it across 3364 lines. That is the classic shape of a trade-off being split rather than dissolved — and a Third Door candidate for Phase 4 (a declarative resolution table with the precedence pinned, rather than ordered branches).
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **A refusing turn serves the empty string.** `respond()`/`arespond()` convert any `ValueError` to `""` for their public `str` contract, so the typed refusal — reason code, blocking region, per-step evidence — is constructed and then discarded before the caller sees it. The system's honesty machinery is real and, on this path, unread.
|
||||
- The layer that owns `wrong=0` is also the layer where the PR #138 fabrications *surface*: the reader's fabricated `member(every_dog, mammal)` and the recited `furthermore` premise reach **served output**. The defect is M3's; the blast radius is M4's. Measured and pinned, held for ADR + ratification — recorded, not fixed.
|
||||
- `epistemic_state` degrades honestly (`epistemic_state_needed`) where hand-copied whitelists degraded dishonestly (`none`). The asymmetry is a reusable design lesson: a coercion that asserts a falsehood is strictly worse than doing nothing.
|
||||
- Surface selection is documented as "current policy… future realizer work may change it" — a standing invitation to accrete another arm.
|
||||
|
||||
**Open questions:**
|
||||
- Materialise the typed refusal into `ChatResponse.refusal_reason` (→ ruling; ADR already anticipated)
|
||||
- Should surface precedence become a declarative table? (→ Phase 4 hindrance audit)
|
||||
- Whether `TurnEvent` should carry `hash_surface` remains an open ruling recorded in `runtime_contracts.md` (→ ruling)
|
||||
86
docs/assessment/10-layer-cards/M5-learning-growth.md
Normal file
86
docs/assessment/10-layer-cards/M5-learning-growth.md
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
# M5 — Learning & Growth
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` · **Fitness:** `fit` (mechanism) / `strained` (throughput) · **Topology role:** runtime boundary + reviewed pack data
|
||||
|
||||
> Controlled mutation. M5 answers the question that separates a cognitive engine from a database: how does the system come to know something it did not know, without acquiring the ability to lie to itself? CORE's answer is a **typed** boundary — durable standing is reviewed or proof-carrying; provisional standing may update autonomously *iff* it is typed, isolated, replayable, and structurally unable to masquerade as ratified truth. This is the layer where the project's epistemology becomes mechanism.
|
||||
|
||||
**Telos stages:** learn from reviewed correction (primary); replay/determinism
|
||||
**Macro role:** Converts served experience and curated material into standing knowledge, at a rate and standing the evidence licenses.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Four regions. **Teaching** (`teaching/`, 46 modules): the reviewed loop — proposals, review, store, replay-equivalence gate, discovery, curriculum premises, ratification, domain chains. **Formation** (`formation/`, 23 modules): the content-addressed data foundry — Mine → Smelt → Forge → Compose → Compile → Run → Ratify → Promote — with six declared trust boundaries, every boundary content-addressed in and out, every rejection producing an audit record. **Reliability calibration** (`core/reliability_gate/`, `core/ratified_ledger.py`): earned SERVE licenses under a Wilson floor (θ_SERVE = 0.99), sealed practice ledgers, hash-verified on load. **Capability** (`core/capability/`, 14 modules): the ledger, the nine-predicate domain contract, lane-shape registry, reviewer registry, expert-demo promotion.
|
||||
|
||||
The boundary is enforced by failing-when-violated invariants rather than convention — INV-21 (vault-writer allowlist), INV-22/23 (unmarked defaults to SPECULATIVE), INV-24 (recall categorization; user-facing evidence is COHERENT-only), INV-29 (only `vault/store.py` transitions `epistemic_status`), INV-30 (open-world `determine()` never asserts False). Autonomous provisional writes — idle CLOSE consolidation, reliability counts, proposal emission, disclosed estimates — all flow through the same `VaultStore.store` path, written SPECULATIVE, rendered `as_told` / `[approximate]` / "proposal". There is no parallel learning path, which is itself an invariant.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** served turns and corrections (M4), comprehension failures and typed refusals (M3), curated source material (formation), curator/HITL rulings, sealed practice results.
|
||||
- **Outputs:** `PackMutationProposal` / `ReviewedTeachingExample` (SPECULATIVE at creation), ratified chain corpora, sealed + SHA-verified ledgers, `LicenseDecision`, capability ledger rows, `MasteryReport` (self-sealing SHA).
|
||||
- **Invariants:** INV-21/22/23/24/29/30 as above; formation's six trust boundaries; content-addressing rules (canonical JSON, sorted keys, **no floats in hashed payloads**, **no pickle** — pickle defeats replay and is a code-execution surface); pack mutation proposal-only until reviewed; identity-manifold mutation by prompt or correction **forbidden**.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** ADR-0021 (epistemic surface + one-mutation-path), ADR-0057 (teaching-chain proposal/review/replay-equivalence), ADR-0175 (calibrated attempt-and-eliminate learning; an engine cannot raise its own bar), ADR-0263 (ratified-ledger bridge: seal → ratify → SHA-verify → serve-gate), ADR-0262/0264 (curriculum + negative curriculum), ADR-0091/0106/0109 (domain contracts, expert-demo promotion, lane-shape registry), ADR-0218 (`apply_certified_promotion`), ADR-0161 (HITL async queue — **scope only**), `docs/teaching_order.md` (five-layer prerequisite-topological doctrine).
|
||||
- **Build:** `partial-wiring-debt`.
|
||||
- **Evidence:**
|
||||
- Miner- and curriculum-sourced proposals route through the *single* reviewed teaching path — lanes — `evals/miner_loop_closure`, `evals/curriculum_loop_closure`, both pinned in `CLAIMS.md` — would-fail-if-absent: **yes**.
|
||||
- Earned-license machinery is real: 25 sealed bands / 18,000 cases with `wrong=0` gate deduction's SERVE — measurement — `chat/data/deduction_serve_ledger.json` — would-fail-if-absent: **yes**.
|
||||
- `formation` suite exists and points at `tests/formation` — code-read — `core/cli_test.py:208` — would-fail-if-absent: **yes**.
|
||||
- Five Tier-1 domains carry ratified status (3 `audit-passed`, 2 `reasoning-capable`), all with zero open gaps — `CLAIMS.md`, mechanically generated.
|
||||
- **Curriculum serve ledger does not exist.** `chat/curriculum_serve_license.py:46` is the single production call site passing `missing_ok=True`, correctly, because `chat/data/curriculum_serve_ledger.json` is absent — code-read.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** an autonomous half that proposes and a reviewed half that ratifies, with volume flowing from curated material through formation into serving licenses.
|
||||
- **Measured — the binding constraint is throughput, and it is quantified.** Curriculum bands are **24×–73× short** of the entailed-bucket floor, and **no served subject has a family close enough to flag as "next."** Compounding this, ADR-0264 §4.1 establishes that the 16-premise compilation cap holds any band to ≤16 entailed cases, so **no curriculum band can earn SERVE until query-scoping lands** — an engineering blocker gating a content problem. Separately, distinct-evidence counting is unsound where replay is treated as independent trials: **21 of 25 ratified bands are short** on that basis (Wilson assumes independence; a replay is one trial).
|
||||
- **Ceilings:** HITL review is CLI-synchronous — the operator cannot review while the engine serves (W-009 open, L10-gated). Vocabulary extension is deliberately manual and slow; the position paper states this plainly as an open scaling question.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**receives** ← M4 (corrections), M3 (failures/refusals); **writes** → M1 (the promotion target; vault and packs); **gated-by** → MG (identity mutation forbidden; safety non-swappable); **evidenced-by** → MV (lanes, ledgers, CLAIMS); **blocked-by** → M6 (async HITL needs the process).
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| learn from reviewed correction | **covered** | Single reviewed path proven by two pinned loop-closure lanes; proposal-only discipline mechanically enforced |
|
||||
| replay/determinism | **covered** | Replay-equivalence gate on teaching chains; content-addressed, float-free, pickle-free hashing |
|
||||
| *(autonomous provisional learning)* | **covered, flag-gated** | CLOSE consolidation climbs to deductive-closure fixed point under proof-gating; **every relevant flag defaults OFF** (`consolidate_determinations`, `review_pending_proposals`, `review_derived_close_proposals`, `auto_contemplate`, `auto_proposal_enabled`, `vault_promotion_enabled`) |
|
||||
|
||||
**Zone roster:** `L7-teaching`, `formation-curriculum`, `reliability-calibration`, `capability` ✱ (straddles MV; owned here).
|
||||
|
||||
**Rollup note:** weakest-link rollup. M5's *mechanisms* are among the most rigorously built in the repository; its *volume* is the constraint. Do not read `partial-wiring-debt` as "the learning loop is unbuilt."
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit` on mechanism, `strained` on throughput.** The typed learning boundary is, in this assessor's reading, CORE's most distinctive and best-executed idea — it dissolves the usual "autonomy versus safety" trade-off rather than splitting it, which is the Third Door done correctly and worth naming as a success rather than only auditing for faults. The strain is entirely on the other side: a ratification pipeline that is architecturally sound and 24×–73× under-fed.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **The autonomous half is built and switched off.** Six or more learning-related flags default `False`. The machinery for a self-improving loop exists, is proof-gated, and does not run by default. Whether that is correct caution or accumulated hesitancy is a ruling, not a finding — but the gap between "built" and "on" is the largest in this layer.
|
||||
- **A committed ledger is necessarily an *earning* one**, which makes the outcome-mix ruling the binding constraint — and the curriculum ledger does not exist yet. The one production `missing_ok=True` is honest and narrow, but it is load-bearing scaffolding.
|
||||
- **The Wilson independence problem is not cosmetic.** If replay is being counted as distinct evidence anywhere a license is earned, licenses are being granted on overstated evidence. 21/25 bands are affected. This deserves Phase 4 attention as a possible `wrong-solution` in the counting, not the gating.
|
||||
- Formation's trust-boundary table is exemplary — content-addressed in and out, no floats in hashed payloads, no pickle, every rejection audited. It should be the template other layers are measured against, and Phase 4 should check whether M2's ingest boundary meets the same bar.
|
||||
- `docs/teaching_order.md` records a known gap since 2026-05-17: the identity-divergence curriculum predates the formation pipeline and still flows through `runner.py` rather than Forge → Compose → Ratify → Promote. Unverified at this SHA; carried forward.
|
||||
|
||||
**Open questions:**
|
||||
- Is the Wilson/replay independence defect actually granting licenses on overstated evidence? (→ Phase 4, then ruling)
|
||||
- Which autonomous-learning flags should be ratified ON, and on what evidence? (→ ruling)
|
||||
- Does M2's ingest boundary meet formation's trust-boundary bar? (→ Phase 2 M2 card / Phase 3)
|
||||
- Curriculum query-scoping: the engineering blocker gating all curriculum SERVE (→ ruling; ADR-0264 §4.1 names it)
|
||||
107
docs/assessment/10-layer-cards/M6-continuity-process.md
Normal file
107
docs/assessment/10-layer-cards/M6-continuity-process.md
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
# M6 — Continuity & Process
|
||||
|
||||
**Kind:** layer · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` · **Fitness:** `strained` · **Topology role:** runtime boundary
|
||||
|
||||
> The life itself. M6 is the layer that makes CORE an organism rather than a function call: a process that persists across time, accumulates capability across reboots, and resumes as *the same life*. Every other layer describes a capability; M6 describes the subject that has them. The foundational telos — "one continuous life" — is this layer's charter, and its condition is the single sharpest measure of the distance between what CORE is for and what CORE is.
|
||||
|
||||
**Telos stages:** replay/determinism, backbone-runtime · hosts every other stage (the cycle's runner)
|
||||
**Macro role:** Holds field, vault, session, and identity continuity over indefinite time; runs the heartbeat that lets the engine learn when nobody is talking to it.
|
||||
|
||||
---
|
||||
|
||||
> **Phase 3 correction (C-2, `05-phase3-findings.md`):** this card's "T1 vault and field excitation are discarded on exit by design" describes the **default config only**. The daemon forces `persist_session_state=True` (Shape B+ — "restored bit-exactly"), so residency machinery exists and is daemon-forced; open items are its exact coverage (`chat/runtime.py:893–952`) and horizon proof. The daemon's forced flag set is exactly `{persist_session_state, consolidate_determinations, strict_identity_continuity}` — see `20-component-cards/always-on-process.md`.
|
||||
|
||||
## ⚠️ Correction to Phase 0 Finding 0-C and to the system map
|
||||
|
||||
**Phase 0 recorded that the L11 always-on process is unbuilt. That is wrong, and the error is instructive.**
|
||||
|
||||
`chat/always_on.py::run_continuous` (279 lines) and `chat/always_on_daemon.py::run_daemon` (195 lines) exist and are reachable from the CLI as `core always-on` (`core/cli.py:244 cmd_always_on`). Both landed **2026-06-14** (`18e25580`, `efd280d4` / PR #758) — **five days after** the `.system-map/` snapshot of 2026-06-09 that declared "no forever entrypoint exists." Phase 0 inherited that claim without re-verification; Phase 2's re-verification obligation is exactly what caught it.
|
||||
|
||||
This is the clearest possible vindication of schema principle P3 (stamped verification) and of the taxonomy's rule that the map is a prior, never a source. It also means the assessment's own worst risk is real: **a 48-day-old map is wrong precisely where the project moved fastest.** Every card in Phases 2–3 must assume the map is stale on anything load-bearing.
|
||||
|
||||
What remains true from Phase 0: the **spike was never run to completion and never recorded**, and the built process is **ungated**. Those are the findings below, and they are different — and more actionable — than "unbuilt."
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
M6 is two halves at very different maturity, joined by a checkpoint.
|
||||
|
||||
**The built half.** `engine_state/EngineStateStore` performs atomic checkpointing (same-dir temp → `fsync` → `os.replace`, mode-bits preserved) of the `RecognizerRegistry`, the `DiscoveryCandidate` working set, and a manifest, written at turn boundaries and reloaded on the next process start. A revision mismatch *warns and never refuses* — "reboot is recovery, not control flow." A `reboot_event` audit line records that a lifetime was lost and regained. On top of this sits `run_continuous`: an unbounded heartbeat loop where each beat advances `idle_tick` (continuous learning), records closure and learning evidence, self-checkpoints on real work, and checkpoints once at exit. `run_daemon` wraps it with a single-instance lock (one life per engine-state dir, lock file deliberately never unlinked), SIGINT/SIGTERM-wired `stop` checked before each beat and interrupting the inter-beat wait, and a load-time identity guard. A `lived_life.json` report feeds the Workbench's Lived Life surface.
|
||||
|
||||
**The unbuilt half — restated correctly.** Not the process: the *proof*, the *gate*, and the *residency*. The 24h+ no-drift soak has a complete falsifiable harness (`evals/l10_always_on`, four predicates H1 closure / H2 bounded-idle / H3 convergence / H4 reboot-resume, each with `*_holds` and `*_bites` test pairs, plus `evals/l10_continuity`) and **has never been run to a recorded artifact**. And per ADR-0146 the T1 vault and field excitation remain deliberately ephemeral, so what compounds across reboots is the recognizer/discovery layer — not the recalled or excited substrate.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** `DerivedRecognizer` objects, `DiscoveryCandidate` objects, `turn_count`, `CORE_ENGINE_STATE_DIR`, git short-revision, prior on-disk checkpoint.
|
||||
- **Outputs:** `engine_state/{recognizers,discovery_candidates}.jsonl`, `manifest.json`, `lived_life.json`, buffered `reboot_event`, restored working set, advisory `RuntimeWarning` on revision mismatch.
|
||||
- **Invariants:**
|
||||
- Atomic checkpoint — crash between write and replace leaves the prior checkpoint intact — pin: `tests/test_adr_0146_engine_state.py` — suite: **none** — status: **not run by any suite**.
|
||||
- Reboot round-trip byte-identity — state after (boot, N turns, reboot, reload) equals state after (boot, N turns) — pin: flagged by the system map as possibly *schema-as-decoration* — status: **unverified at this SHA**.
|
||||
- Versor closure under long-horizon idle (`VERSOR_CEILING`) — pin: `tests/test_l10_always_on_soak.py` — suite: **none** — status: **not run by any suite**.
|
||||
- One life per engine-state dir — pin: `tests/test_l10_always_on_daemon.py` — suite: **none**.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** `docs/adr/L10-runtime-model-scope.md` (scope only — process shape, state partitioning, reboot recovery, async HITL left open), ADR-0146 (Shape B hybrid checkpoint chosen over Shape A daemon / Shape C audit replay), ADR-0156/0157/0158, `L11-hitl-async-queue-scope.md`. **The design record is now behind the code**: ADR-0146 explicitly *rejected* the Shape A daemon, and a daemon was subsequently built. No ADR ratifies `always_on_daemon.py`. That is a governance gap, not merely a documentation lag.
|
||||
- **Build:** `partial-wiring-debt`. Key files: `engine_state/__init__.py`, `chat/always_on.py`, `chat/always_on_daemon.py`, `core/cli.py::cmd_always_on`, `evals/l10_always_on/`, `evals/l10_continuity/`.
|
||||
- **Evidence:**
|
||||
- Always-on loop and daemon exist and are CLI-reachable — code-read — `chat/always_on.py:179`, `core/cli.py:244` — would-fail-if-absent: **yes** (CLI command would not resolve).
|
||||
- Falsifiable soak harness with holds/bites predicate pairs exists — code-read — `evals/l10_always_on/predicates.py`, `tests/test_l10_always_on_soak.py` — would-fail-if-absent: **yes**.
|
||||
- Long-horizon no-drift result — **absent**. No JSON artifact under `evals/l10_always_on/` or `evals/l10_continuity/`. would-fail-if-absent: **n/a — the claim has no evidence**.
|
||||
- Suite enforcement — **absent**. No `l10`/`always_on` test appears in any of the 21 `TEST_SUITES` tuples. would-fail-if-absent: **no — nothing runs these pins**.
|
||||
- Persistence defaults — `persist_session_state=False`, `strict_identity_continuity=False` (`core/config.py:285,293`) — measurement — the continuity machinery is **off by default**.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** an indefinitely-running single process holding field + vault + session without drift, contemplating in slack, resuming as the same life through interruption.
|
||||
- **Measured:** heartbeat loop runs for a bounded `heartbeats` count or until `stop`; checkpoint round-trip works at turn granularity; **no measured horizon exists** — the longest verified run is whatever the short soak tests exercise, and those run in no suite. The compounding surface is the recognizer/discovery layer only.
|
||||
- **Ceilings:** T1 vault and field excitation are discarded on exit **by design** (ADR-0146), so recall and excitation substrate reset every process. Cross-reboot `EngineIdentity` verification is shelved. Async HITL-while-serving is still CLI-synchronous (W-009 open).
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
- **feeds** → M4 (`ChatRuntime` is the sole consumer of the checkpoint store); **recalls-from** → M3 (recognizer registry) and M1/M5 (discovery candidates); **reads** → M1 (vault, deliberately ephemeral) and M0 (field, deliberately ephemeral); **verifies** → MV (round-trip byte-identity is a replay obligation); **depends-on** → MG (same-life identity continuity overlaps identity doctrine, not yet wired to the identity axes).
|
||||
- **Owning ADRs:** ADR-0146, ADR-0156, ADR-0157, ADR-0158, ADR-0220 (identity/build-provenance split); scopes `L10-runtime-model-scope.md`, `L11-hitl-async-queue-scope.md`. **Unowned:** `chat/always_on_daemon.py`.
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| backbone-runtime | **claimed-only** | Process exists and runs; no measured horizon, no suite-enforced pin, continuity flags default off |
|
||||
| replay/determinism | **claimed-only** | Round-trip byte-identity invariant flagged as possible schema-as-decoration; unverified at this SHA |
|
||||
| *(hosting the cycle)* | **covered, degraded** | `idle_tick` genuinely advances learning per beat; but what compounds is recognizers/discovery only — field and vault reset |
|
||||
|
||||
**Zone roster:** `L10-11-runtime-identity` (partial-wiring-debt → **revised: build materially more complete than mapped**), `engine-state` (live-internal, confirmed), `edge-sync` (inert, not re-verified), `core-protocol-ctp` (spike, not re-verified).
|
||||
|
||||
**Rollup note:** zone liveness is a weakest-link rollup and is not a completion rate. M6 is better built than either the map or Phase 0 reported, and *less proven* than the existence of the code suggests. Those are not in tension — they are the layer's actual condition.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained`.** Not `wrong-solution`: Shape B checkpointing is a sound mechanism, and the daemon built on top of it is a reasonable shape. Strained because the layer's *proof obligations* have not kept pace with its *code*, and because its central design document still rejects the architecture that was subsequently built. A layer whose invariants are pinned by tests that no suite runs is protected by nothing.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- The most telos-critical component in CORE is enforced by **zero running tests**. Every `l10`/`always_on` pin sits outside all 21 suites. By the project's own doctrine (a pin in no suite never runs), the continuous life is unguarded.
|
||||
- **`DriveGradientMap` is constructed and never read.** `chat/runtime.py:716` builds `self._drive_map`; no site reads it. Deleting it would change no output. This is textbook decoration and is recorded here rather than fixed.
|
||||
- `ExertionMeter` *is* exercised (`record` at :2912, `fatigue` at :2913) but its output flows only to `drive_summaries` and `fatigue_index` — telemetry. Fatigue **gates no decision**. It is live-internal, not load-bearing.
|
||||
- The daemon has no ratifying ADR while the governing ADR-0146 explicitly rejected the daemon shape. Whatever the right answer, the record currently contradicts the code.
|
||||
- "One continuous life" remains, in substance, *many short lives sharing a recognizer checkpoint* — but the reason is now a **deliberate ADR-0146 scoping decision** about vault/field ephemerality, not an absence of process machinery. That reframes the L10 question entirely: it is no longer "build the process" but "decide what the process is allowed to hold."
|
||||
|
||||
**Open questions:**
|
||||
- Run the long-horizon soak and record the artifact — what horizon does the engine actually survive? (→ ruling: this is the ADR-0146 Phase-4 spike, still owed)
|
||||
- Should `l10`/`always_on` pins enter a suite, and which? (→ Phase 4)
|
||||
- Does the daemon need a ratifying ADR that supersedes ADR-0146's Shape-A rejection? (→ ruling)
|
||||
- Is vault/field ephemerality still the right call now that a real process exists to hold them? (→ ruling; this is the highest-value M6 question)
|
||||
89
docs/assessment/10-layer-cards/MG-governance-identity.md
Normal file
89
docs/assessment/10-layer-cards/MG-governance-identity.md
Normal file
|
|
@ -0,0 +1,89 @@
|
|||
# MG — Governance & Identity (cross-cutting)
|
||||
|
||||
**Kind:** layer (cross-cut) · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-serving` · **Fitness:** `fit` (with one `strained` region) · **Topology role:** runtime boundary + reviewed pack data
|
||||
|
||||
> Who the system is, and what it will not do. MG is cross-cutting because a governance mechanism obeyed by only one layer is not governance — it is a feature. Its philosophical charge is the alignment thesis stated geometrically: keep behavior inside an intended region of possibility space *by construction*, so that alignment is a structural property of the manifold rather than a filter applied to output. The load-bearing asymmetry is that identity and ethics are swappable while safety is not.
|
||||
|
||||
**Telos stages:** articulate (governs the served surface), replay/determinism (verdicts are evidence)
|
||||
**Macro role:** Constrains every layer's behavior to an identity- and safety-consistent region, and refuses in typed form when it cannot.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Four pack families with deliberately different mutability. **Safety** (`packs/safety/`) is **never swappable and fail-closed** — five boundaries present in every manifold; a `SafetyVerdict` violation produces a deterministic typed refusal. **Identity** (`packs/identity/`) is swappable; ship default `default_general_v1`; the manifold is loaded from `packs/identity/<pack_id>.json`. **Ethics** (`packs/ethics/`) is swappable with opt-in typed refusal. **Register/anchor-lens** packs shape expression without moving truth.
|
||||
|
||||
Identity scoring is **wave-only and fail-closed** (ADR-0244 §3, INV-32): `IdentityCheck.check(trajectory, manifold, wave_field=...)` requires an explicit Cl(4,1) wave field; absence raises typed `MissingWaveStateError`, malformed fields raise `ValueError`, and the scalar-L2 dual-mode fallback has been **excised** and may not be reintroduced. Crucially, live *refusal* remains flag-gated (`identity_wave_gate`, default **off**, explicitly *not authorized*) — the flag controls refusal only, never the scoring path.
|
||||
|
||||
The identity contract is protective in both directions: a flagged score must not silently erase useful generation absent an explicit, tested hard-block policy; and **identity-manifold mutation by user prompt or correction is forbidden**, with identity-override attempts *rejected, not learned*. Adversarial override probes are seeded before the concepts they protect (`docs/teaching_order.md` layer 1).
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** trajectory + final wave field from M3/M4, mounted identity/safety/ethics/register packs, session identity path.
|
||||
- **Outputs:** `IdentityScore` (with `flagged`), `SafetyVerdict`, `EthicsVerdict`, typed refusals, hedge injection, `TurnVerdicts`.
|
||||
- **Invariants:**
|
||||
- **INV-32** — identity scoring is wave-only; no scalar-L2 fallback exists — pins: `tests/test_stage2_physics_hardening.py` (excised symbols absent), `tests/test_adr_0244_identity_gate_runtime.py`.
|
||||
- Identity manifold is never mutated by prompt or correction; override attempts rejected, not learned.
|
||||
- Safety pack is never swappable; violation → deterministic typed refusal.
|
||||
- Hedge injection is exclusive with refusal (never both).
|
||||
- No epistemic seal: no `final`/`frozen`/`axiom`/`permanent` flag may exist (shared with M1's non-hardening invariant).
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** ADR-0027 (identity packs), ADR-0244 (wave-field identity manifold and inalienable geometric alignment), ADR-0246 (induced identity action and path integrity), ADR-0010 (identity physics — the mind-physics blueprint element that *did* land), ADR-0220 (engine identity split from build provenance), `docs/identity_packs.md`, `docs/safety_packs.md`, `docs/ethics_packs.md`, `docs/refusal-taxonomy.md`, `docs/position_paper.md` §6.
|
||||
- **Build:** `live-serving`. Key files: `packs/safety/check.py`, `packs/identity/loader.py`, `packs/ethics/check.py`, `core/physics/identity.py`, `core/physics/identity_manifold.py`, `core/physics/identity_action.py`.
|
||||
- **Evidence:**
|
||||
- Safety and identity are constructed and consulted on the live serving path — code-read — `chat/runtime.py:93,104-105` (`load_identity_manifold`, `SafetyCheck`, `load_safety_pack`) — would-fail-if-absent: **yes**.
|
||||
- Scalar-L2 fallback is excised and its absence is pinned — pin — `tests/test_stage2_physics_hardening.py` — would-fail-if-absent: **yes** (a reintroduced symbol fails the pin).
|
||||
- Identity protection under adversarial input is claimed and evidenced in `docs/position_paper.md` §4 with an eval lane (`evals/identity_divergence/`) — lane.
|
||||
- Three merged authority demos (claims / proposed tool actions / epistemic-state assignment) — `docs/position_paper.md`, PRs #687/#688/#690.
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** alignment as a structural property — the system cannot leave the intended region because the geometry does not admit it.
|
||||
- **Measured:** identity scoring is metric-exact geometry (Gram / operator-preservation) on every turn; refusal is typed and taxonomized; five safety boundaries present in every manifold.
|
||||
- **Ceilings:** **live identity *refusal* is off.** `identity_wave_gate` defaults `False` and is documented as "not authorized" — so identity currently *scores and flags* but does not *block*. Likewise `identity_action_surface=False` and `strict_identity_continuity=False`.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**governs** → M4 (surface), M3 (what may be concluded), M5 (identity mutation forbidden); **reads** → M1 (packs), M0 (wave field, Gram geometry); **evidenced-by** → MV; **overlaps** → M6 (same-life continuity across reboot).
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| articulate (governance of) | **covered** | Safety/ethics/identity consulted on the live path; typed refusal; hedge injection |
|
||||
| *(identity enforcement)* | **claimed-only** | Scoring is live and fail-closed; **blocking is flag-gated off and unauthorized** |
|
||||
|
||||
**Zone roster:** `governance-identity-safety` (live-serving, **confirmed** by serving-path imports); `L9-epistemic-verdicts` ✱ (owned by M4, cross-referenced here).
|
||||
|
||||
**Rollup note:** the only macro layer besides M4 that the map rates `live-serving` at zone level, and re-verification supports it.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`, with the enforcement region `strained`.** The design here is the strongest expression of CORE's alignment thesis: making safety non-swappable while identity and ethics are swappable is a genuine structural distinction rather than a policy label, and excising the scalar fallback outright — rather than deprecating it — is the correct response to a dual-mode hazard. The `strained` region is enforcement: the machinery is built, hardened, pinned, and **not switched on**.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **Identity flags but does not block.** The gate is default-off and explicitly *not authorized*, which is an honest and deliberate posture — but it means the strongest claim in the alignment story ("behavior stays in the intended region") is, at runtime today, a *measurement* rather than a *constraint*. The distinction between scoring and refusing must be stated precisely wherever this capability is described externally; `runtime_contracts.md` does state it precisely, and this card records it so no downstream summary blurs it.
|
||||
- Cross-cutting writ is asserted but not mechanically verified. I found no pin that would fail if some layer *bypassed* governance — as distinct from pins that verify governance works when called. For a cross-cutting layer that is the invariant that matters most, and its absence is the analogue of M1's unverified no-approximate-recall pin. Flagged for Phase 3.
|
||||
- Safety's non-swappability is doctrine plus loader design; whether a *test* fails on an attempt to swap the safety pack was not verified at this SHA.
|
||||
- MG is where CORE's most defensible public claims live. That makes precision about flag state a reputational as well as a technical obligation.
|
||||
|
||||
**Open questions:**
|
||||
- Is there a pin that fails when a layer bypasses governance entirely? (→ Phase 3)
|
||||
- Under what evidence should `identity_wave_gate` be authorized live? (→ ruling)
|
||||
- Is safety-pack non-swappability mechanically enforced or loader-conventional? (→ Phase 3)
|
||||
90
docs/assessment/10-layer-cards/MV-verification-evidence.md
Normal file
90
docs/assessment/10-layer-cards/MV-verification-evidence.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# MV — Verification & Evidence (cross-cutting)
|
||||
|
||||
**Kind:** layer (cross-cut) · **Parent:** CORE · **Assessor:** Opus 5 (Phase 2)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `partial-wiring-debt` · **Fitness:** `strained` · **Topology role:** benchmark/eval artifact + tooling surface
|
||||
|
||||
> How CORE knows what it knows about itself. MV is cross-cutting because replay is a *property* of every layer and an *apparatus* of this one. Its philosophical charge follows directly from the thesis: a decoding system's failures are locatable, so the machinery that locates them is not overhead — it is the difference between a structured failure that can be fixed and a stochastic one that can only be regularized. MV is also the layer this assessment most depends on, and therefore the one whose gaps most threaten the assessment's own conclusions.
|
||||
|
||||
**Telos stages:** replay deterministically (primary); learn (evidence for calibration)
|
||||
**Macro role:** Produces the evidence that every other layer's claims are measured against, and fails loudly when a claim drifts.
|
||||
|
||||
---
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`evals/` (375 modules) holds the lane runners and their contracts. `tests/` (881 modules) holds the pins. `core/cli_test.py` defines **21 suites** as hand-curated tuples — `fast, smoke, runtime, cognition, teaching, packs, algebra, sensorium, pulse, formation, proof, refusal, margin, rotor, inner-loop, phase5, phase6, adr-0024, math, deductive, full`. `CLAIMS.md` is machine-generated from in-tree state: Tier 1 = five ratified domains from `core.capability.ledger_report`; Tier 2 = eleven lanes pinned by report SHA-256, where mismatch is a CI failure. `workbench/` (22 modules) is the read-only operator/auditor projection. `scripts/verify_lane_shas.py` is the Tier-2 verifier.
|
||||
|
||||
The validation doctrine is **local-first**: the in-worktree run is the merge bar (`uv run core test --suite smoke -q` pre-push, larger suites pre-merge, `[Verification]:` on the PR). A pre-push hook runs smoke **plus** the `warmed_session` consistency lane pin, deliberately targeted at a regression class smoke does not cover. `scripts/ci/local-ci.sh` reads suite membership *from the CLI* rather than restating it, so the runner cannot drift from the hook. The interpreter contract is fail-closed: `requires-python == "3.12.13"` exactly, and a degraded run stamps itself **NON-CANONICAL** on every line — a degraded run is legitimate, a degraded run reporting itself as the real thing is not.
|
||||
|
||||
---
|
||||
|
||||
## Contract
|
||||
|
||||
- **Inputs:** every other layer's runtime behavior; lane runners; pinned SHAs.
|
||||
- **Outputs:** lane reports (JSON), `CLAIMS.md` rows, `trace_hash`, `CognitivePipelineRecord`, telemetry events, capability ledger, `[Verification]:` evidence.
|
||||
- **Invariants:**
|
||||
- Tier-2 lane report SHA-256 must reproduce byte-for-byte; drift is a CI failure and demotes the claim.
|
||||
- Replay treats `pipeline_record` as **critical evidence** — divergence is evidence against equivalence, not wall-clock noise.
|
||||
- Pre-widening journal rows must show `missing_evidence`, never a synthesized green pipeline.
|
||||
- `derive_evidence_digest` is deterministic in field order; re-running against on-disk lane results reproduces `claim_digest`.
|
||||
- Unknown lane ids **fail closed** (`lane <id> has no registered shape — introduce via ADR amendment`); a broken `reviewers.yaml` yields an empty registry rather than silently granting `audit_passed=true`.
|
||||
|
||||
---
|
||||
|
||||
## Design vs build
|
||||
|
||||
- **Design:** ADR-0092/0093/0096/0098/0099 (registry, contract validation, fabrication control, demo composition, public demo), ADR-0106/0109 (expert-demo promotion, lane-shape registry), ADR-0119.x (GSM8K lane + seal discipline), ADR-0167 (audit-as-evidence), `docs/testing-lanes.md`, `docs/eval_methodology.md`, `AGENTS.md` §Local-First CI Validation Protocol.
|
||||
- **Build:** `partial-wiring-debt`.
|
||||
- **Evidence:**
|
||||
- Eleven Tier-2 lanes carry pinned SHAs verified by CI — `CLAIMS.md` + `.github/workflows/lane-shas.yml` — would-fail-if-absent: **yes**.
|
||||
- Suite membership is read from the CLI by the local runner, so hook and runner cannot diverge — code-read — `scripts/ci/local-ci.sh` — would-fail-if-absent: **yes**.
|
||||
- Fabrication-control lane enforces per-class `refused == n, fabricated == 0` — lane shape — `core/capability/expert_demo.py` — would-fail-if-absent: **yes**.
|
||||
- Sealed holdout discipline: 1,319 GSM8K cases age-encrypted, plaintext never on disk, no CI workflow sets `CORE_HOLDOUT_KEY`, team operates blind — code-read + contract.
|
||||
- **Suite coverage is incomplete in a load-bearing way** — measurement — no `l10`/`always_on` test appears in any of the 21 suite tuples (see M6).
|
||||
|
||||
---
|
||||
|
||||
## Capacity
|
||||
|
||||
- **Designed:** every claim mechanically derived from in-tree state and verified by CI.
|
||||
- **Measured:** 21 suites; 881 test modules; 11 SHA-pinned lanes; 5 Tier-1 domains; `deductive` suite at 20 files, `smoke` at 23, `sensorium` at 21, `algebra` at 15. The full ~12k fast-lane runs async by design so the pre-push hook stays targeted rather than gridlocking the push cycle.
|
||||
- **Ceilings:** the smoke gate is not the full suite — a PR's own tests must be run against the rebased base pre-merge, which is process discipline rather than mechanism. Remote Actions are secondary observability only; GitHub mirror Actions are billing-locked dead signals. Roughly 31 red tests were retired historically. The sealed GSM8K holdout has never been opened against a parser with sufficient coverage.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & provenance
|
||||
|
||||
**witnesses** → every layer; **reads** → M5 (capability ledger, reviewer registry), M4 (telemetry, trace); **constrains** → M0 (bit-exact parity pins).
|
||||
|
||||
---
|
||||
|
||||
## Stage coverage
|
||||
|
||||
| Stage | Verdict | Evidence |
|
||||
|---|---|---|
|
||||
| replay deterministically | **covered** | Tier-2 SHA pinning is mechanical and fails CI on drift; pipeline record is critical replay evidence |
|
||||
| *(evidence for learning)* | **covered** | Sealed practice ledgers, capability ledger, lane-shape registry all fail closed on unknown shapes |
|
||||
| *(coverage of the telos-critical layer)* | **uncovered** | No suite runs any L10/always-on pin |
|
||||
|
||||
**Zone roster:** `evals-determinism`, `tooling-cli-workbench-rs`; `capability` ✱ (owned by M5, cross-referenced here).
|
||||
|
||||
**Rollup note:** weakest-link rollup. MV's *mechanisms* for the things it covers are excellent; its *coverage map* has a hole at precisely the layer with the least other protection.
|
||||
|
||||
---
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained`.** MV's design instincts are consistently right — machine-generated claims, fail-closed registries, SHA pinning, a non-canonical stamp for degraded runs, suite membership read rather than restated. These are the habits of a system that expects to be wrong and wants to find out. The strain is structural rather than qualitative: **suite membership is hand-curated**, so coverage is a curation artifact, and nothing detects the *absence* of a pin from every suite. That is how the most telos-critical layer in CORE came to be enforced by nothing while its tests exist and pass on demand.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- **The hole is where it hurts most.** M6 has a complete falsifiable soak harness with holds/bites predicate pairs — genuinely excellent test design — that no suite runs. The tests are not weak; they are *unscheduled*. A hand-curated suite list has no mechanism to notice.
|
||||
- **There is no meta-pin for coverage.** The project has learned "a pin in no suite never runs" as doctrine, but nothing mechanically enforces it. A test file that belongs to zero suites is currently indistinguishable from a test file that runs everywhere. This is, in this assessor's reading, MV's single highest-leverage improvement and a strong Phase 4 Third-Door candidate: rather than curating harder, make orphaned pins detectable.
|
||||
- The gap registers this layer ought to feed are dead (Phase 0 Finding 0-D): `docs/gaps.md` is 26/26 closed, the liveness ratchet is L10-blocked and stale, and `docs/analysis/` is a chronological archive of ~130 documents with no aggregator. MV produces excellent point-in-time evidence and has no standing instrument that accumulates it.
|
||||
- Long-horizon evidence has no home. The soak harness's `__main__` exists; no results artifact does. Lanes are pinned by SHA precisely because point-in-time reports drift — but a lane nobody runs produces no report to pin.
|
||||
- MV's honesty machinery (`missing_evidence`, NON-CANONICAL stamping, fail-closed unknown lanes) is a genuine model for the rest of the system and should be cited as such in Phase 5 rather than only audited.
|
||||
|
||||
**Open questions:**
|
||||
- Add a meta-pin asserting every `tests/*.py` belongs to ≥1 suite (or an explicit exclusion list)? (→ Phase 4; likely highest-leverage single change in MV)
|
||||
- Which suite should own the L10/always-on pins, given they are soaks? (→ Phase 4 / ruling)
|
||||
- Should the assessment's gap register become the standing instrument MV lacks? (→ Phase 4)
|
||||
43
docs/assessment/20-component-cards/always-on-process.md
Normal file
43
docs/assessment/20-component-cards/always-on-process.md
Normal file
|
|
@ -0,0 +1,43 @@
|
|||
# always-on-process — `chat/always_on.py`, `chat/always_on_daemon.py`, `engine_state/`, `evals/l10_*`
|
||||
|
||||
**Kind:** component (M6's built half) · **Parent:** M6 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal` (CLI-reachable, suite-orphaned) · **Fitness:** `strained` (proof debt, not design debt) · **Topology role:** runtime boundary
|
||||
|
||||
> The process that runs the continuous-life heartbeat. Its design center is a single sentence from the daemon module: a restart is *the same life or it stops — never a silent fork*.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
`run_continuous(runtime, heartbeats, …)` — each beat advances `idle_tick` (continuous learning), records closure + learning evidence, self-checkpoints on real work, checkpoints once at exit; `heartbeats=None` runs unbounded until `stop`, which is checked before each beat *and* interrupts the inter-beat wait. `run_daemon` wraps it with a single-instance `fcntl.flock` lock — kernel-released on process death, so no stale-lock window and no PID-reuse ambiguity; the lock file is deliberately never unlinked (unlinking would let a peer flock a different inode). `lived_life.json` feeds the Workbench Lived Life surface. Landed 2026-06-14 (`18e25580`, `efd280d4`).
|
||||
|
||||
**The forced flag set (exact, from `CONTINUOUS_LIFE_CONFIG_FLAGS`):**
|
||||
|
||||
```python
|
||||
{"persist_session_state": True, # Shape B+ — persist the lived session across reboot
|
||||
"consolidate_determinations": True, # Step D — learn from determined facts each beat
|
||||
"strict_identity_continuity": True} # load-time identity guard — same life or refuse
|
||||
```
|
||||
|
||||
## Correction to the Phase 2 M6 card (Shape B+)
|
||||
|
||||
The M6 layer card carried forward "T1 vault and field excitation are discarded on exit **by design** (ADR-0146)." That is the **default-config** posture only. Under the daemon, `persist_session_state=True` activates Shape B+ persistence — `chat/always_on.py:9`: the lived state is "restored bit-exactly" — with opt-in persistence sites in `chat/runtime.py:893–952`. So the residency question the M6 card called "the highest-value M6 question" (*should the process hold vault/field?*) is already partially answered in code: **the mechanism exists, is opt-in, and is daemon-forced.** What remains open is exactly what Shape B+ covers (session context vs full T1 vault vs field excitation — the persistence sites need a Phase-4 read) and whether it is *proven* at horizon.
|
||||
|
||||
## What the soak would prove if run (`evals/l10_always_on`)
|
||||
|
||||
- **H1 closure** — every observed idle beat is a valid versor, with an explicit vacuity guard: a run where the field never existed cannot pass "by saying nothing."
|
||||
- **H2 bounded idle** — a no-work idle beat adds nothing to the vault (no idle resource leak); flagged consolidation writes are exempt.
|
||||
- **H3 convergence** — a saturated idle life *settles and stays settled*, with a `min_converged_tail` so "settled" is observed, not assumed.
|
||||
- **H4 reboot-resume** — a mid-soak reboot resumes the SAME life; post-reboot closure holds on every segment.
|
||||
|
||||
Each predicate has `*_holds` and `*_bites` test pairs (mutated evidence must fail). This is exemplary falsification design. **No recorded artifact exists; no suite runs any of it** — the long horizon is `python -m evals.l10_always_on`'s job, per its own docstring "run on demand / nightly," and no nightly exists.
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained` — proof debt on a sound design.** The lock discipline, the never-silent-fork rule, the vacuity-guarded predicates, and the bites-pairs are all careful work. What is missing is entirely evidentiary: a recorded soak artifact, a scheduled runner, and an ADR that owns the daemon (ADR-0146 rejected Shape A; the daemon is unowned by any ratifying decision).
|
||||
|
||||
**Honest wrinkles:**
|
||||
- The flag set forces the *consolidator* on but not the *accruer* (`accrue_realized_knowledge` absent) — see the determine-phase card: the continuous life may be consolidating an empty set. Unresolved.
|
||||
- `strict_identity_continuity=True` under the daemon vs `False` default means identity-continuity refusal behavior differs between the daemon and every other entrypoint — correct by design, but nowhere stated outside the flag dict.
|
||||
- The soak evals' own docstring designates a nightly cadence that has never been provisioned. Given the local-first CI doctrine (Mac runner, queue waits while asleep), a nightly soak is architecturally awkward — which may be *why* it never ran. That tension deserves a ruling, not silence.
|
||||
|
||||
**Open questions:** run the soak, record the artifact (→ the still-owed ADR-0146 Phase-4 spike); which suite owns `l10`/`always_on` pins (→ Phase 4); daemon's ratifying ADR (→ ruling); exact Shape B+ coverage (→ Phase 4 read of `runtime.py:893–952`).
|
||||
35
docs/assessment/20-component-cards/attention-allocation.md
Normal file
35
docs/assessment/20-component-cards/attention-allocation.md
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
# attention-allocation — the live CR-1 mechanism
|
||||
|
||||
**Kind:** component (mechanism spanning `generate/salience.py`, `generate/attention.py`, `core/physics/{salience,attention,inhibition}.py`, `generate/stream.py`) · **Parent:** unowned (CR-1) — de facto M3 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-serving` (default ON) · **Fitness:** `strained` — sound mechanism, absent governance · **Topology role:** runtime boundary
|
||||
|
||||
> The mechanism that decides what CORE is allowed to consider at each generation step. It gates every token walk, owns the measured hot path (~73% of turn time through `cga_inner`), and is the only default-ON serving mechanism in the tree with no owning ADR, no zone, and no stage in any ratified articulation. This card gives it the identity it has been operating without.
|
||||
|
||||
## What it is / What it does (372 lines total, all read)
|
||||
|
||||
**`generate/salience.py` (62)** — `SalienceOperator.compute(field, vocab, top_k=16) → SalienceMap{indices, scores, budget}`. The docstring states the lineage plainly: "generation-facing salience from **ADR-0008 field curvature**… the score is now a local curvature magnitude from `core.physics.salience` rather than normalized proximity to the query field." So the mind-physics blueprint's Allocation Physics **did land**, mutated: the physics curvature kernel (`core/physics/salience.py`, 139 lines) is *composed* — imported as `CurvatureSalienceOperator` — by a thin generation-facing adapter. Not a duplicate (Phase 2's correction, confirmed by read).
|
||||
|
||||
**`generate/attention.py` (43)** — `AttentionOperator(inhibition_threshold).plan(salience, vocab) → allowed_indices`. Inhibition is a **scalar threshold** (default 0.3), not the blueprint's mask.
|
||||
|
||||
**Wiring (`generate/stream.py:255–331`)** — `_attention_candidates` runs salience→attention when `use_salience` (default **True**); the walk intersects `language_candidates ∩ salience_candidates` (`:329`), and when language candidates are absent, salience candidates alone gate the step (`:331`). At `:637` the salience `budget` is fed back as the next `salience_top_k` — attention narrows itself as the walk proceeds.
|
||||
|
||||
**The decoration:** `core/physics/inhibition.py` (54 lines, `InhibitionOperator`/`InhibitionMask`) is imported only by `core/physics/__init__.py` and one parity test. No serving or eval path constructs a mask. By the sabotage test: **decoration** — the blueprint's third operator exists as an exported symbol and nothing more.
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- Live and load-bearing — code-read + config — `use_salience=True` default (`core/config.py:35`); deleting `_attention_candidates` would change every walk's candidate set — would-fail-if-absent: **yes**.
|
||||
- Deterministic — pure functions of field state and vocab; parity-pinned (`tests/test_salience_vectorize_parity.py` pins the vectorized rewrite against the original nested loop).
|
||||
- Behavioral pin — `tests/test_salience.py` exercises compute+plan; suite membership not individually confirmed (flagged, like most non-lane pins in this assessment).
|
||||
- Interaction contract with admissibility: salience/attention prune *candidates*; admissibility (threshold/margin/rotor, ADR-0024/0025/0026) then judges them. Two allocation stages, only the second governed by ADRs.
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained` — the gap is governance, and it is now fully characterized.** Mechanism: sound, deterministic, composed cleanly over the physics kernel. Governance: ADR-0008 is a *draft blueprint's* operator spec, never ratified as owning this serving behavior; no ratified articulation names an attention stage; the config knobs (`salience_top_k=16`, `inhibition_threshold=0.3`) are tuned constants with no recorded derivation — precisely the "threshold tuned for good-enough" the mastery framework's Pillar I forbids, sitting at the semantic center of generation.
|
||||
|
||||
**Honest wrinkles:**
|
||||
- The budget feedback loop (`:637`) means attention is *self-narrowing* across a walk — a real cognitive-model property (fatiguing focus? commitment?) that no document names, let alone justifies.
|
||||
- The hot-path cost (~73% of turn time) is this mechanism's nearest-neighbour/salience search — so CORE's compute profile is dominated by a mechanism whose parameters no decision governs.
|
||||
- CR-1's ruling question is now precise: not "is attention a layer" in the abstract, but **who owns `use_salience`, the two constants, the budget feedback, and the InhibitionMask disposition** (ratify the mask's deletion, or build it). A one-page ADR closes all four.
|
||||
|
||||
**Open questions:** the CR-1 ADR (→ ruling); derive or empirically justify `top_k=16` / `threshold=0.3` (→ Phase 4); delete or build `InhibitionMask` (→ ruling; deletion is the mastery-framework default).
|
||||
28
docs/assessment/20-component-cards/comprehend-organ.md
Normal file
28
docs/assessment/20-component-cards/comprehend-organ.md
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
# comprehend-organ — `core/comprehension_attempt/`
|
||||
|
||||
**Kind:** zone→component descent · **Parent:** M3 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal` (**demoted** from the map's `live-serving`) · **Fitness:** `fit`, misleadingly named · **Topology role:** runtime boundary (off the chat serving path)
|
||||
|
||||
> The deterministic multi-organ setup router (N3): when several comprehension organs each attempt a problem setup, admit a setup only when exactly one organ produced an admissible one — or the admitting organs agree by signature. Never pick among disagreeing readings. Its philosophical role is refusal-preserving arbitration: routing must not become a place where a guess can hide.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Six modules, 754 lines: `router.py` (77 — the N3 router), `classify.py` (159 — `classify_cmb/r1/r2/r3`), `failure_family.py` (250), `proposal.py` (148 — comprehension-failure proposals), `model.py` (63 — `ComprehensionAttempt`), `__init__.py`. Routing rule, verbatim from the module doc: exactly one `setup_correct` → routed; zero → `all_refused` (classified downstream); ≥2 agreeing signatures → routed; ≥2 differing → `ambiguous` (refuse — never pick). Cross-organ signatures are produced by different functions and never coincide, so two admitting organs resolve to `ambiguous` in practice. The router never solves and never emits `setup_wrong` — that is an eval-only outcome.
|
||||
|
||||
## The liveness demotion, with evidence
|
||||
|
||||
`comprehension_attempt` is imported by **neither** `chat/runtime.py` nor `core/cognition/pipeline.py`. Its non-test callers are `core/epistemic_disclosure/limitation.py` (→ `chat/ask_runtime.py`, behind `ask_serving_enabled=False`), `core/proposal_review/queue.py`, `core/epistemic_questions/delivery.py`, and `evals/constraint_oracle/verified_producer.py`. Nothing on the default serving path reaches it. The map's `live-serving` label is therefore wrong at this SHA: **`live-internal`** (eval lanes + flag-gated ask path + proposal machinery).
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- Refuse-on-ambiguity is the load-bearing invariant — pins: router tests (`tests/test_cmb_router_contemplation.py`, `tests/test_failure_family.py`, `tests/test_failure_proposal.py`) — present in the tree; suite membership not individually confirmed.
|
||||
- wrong=0 posture: against gold the routed setup must match; the router structurally cannot emit a wrong setup, only a refusal or a routed one.
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit` — but the zone name is a trap.** This is a *math setup* router over the GSM8K-era R1/R2/CMB organs, not "the comprehension organ" of the chat path. Chat comprehension lives in `generate/meaning_graph/reader.py` (see the realize-phase card). A future reader who greps "comprehend-organ" expecting the thing that reads user English will land in the wrong subsystem. Recommend the zone be renamed (e.g. `setup-router`) or its card carry this disambiguation permanently.
|
||||
|
||||
**Honest wrinkles:** the design is deliberately boring and correct — no dynamic scoring, no priority heuristics. Its consumers are all flag-gated or eval-side, so this machinery currently arbitrates for paths that mostly do not serve.
|
||||
|
||||
**Open questions:** should `epistemic_questions`/`disclosure` consumers ever reach serving, the router becomes serving-path arbitration — does it then need a suite-pinned invariant? (→ Phase 4)
|
||||
39
docs/assessment/20-component-cards/derivation-organs.md
Normal file
39
docs/assessment/20-component-cards/derivation-organs.md
Normal file
|
|
@ -0,0 +1,39 @@
|
|||
# derivation-organs — `generate/derivation/`
|
||||
|
||||
**Kind:** component group · **Parent:** M3 (gsm8k-math zone) · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal` (math lanes; not on the chat serving path) · **Fitness:** `superseded-in-place` (by ADR-0252 ruling) · **Topology role:** runtime boundary (math reader)
|
||||
|
||||
> The "novice surface piles" of ADR-0252's diagnosis: per-shape derivation organs that turn recognized problem statements into typed solver state. Ruled superseded by the structure-mapping paradigm — and ruled to keep serving until a proven replacement exists.
|
||||
|
||||
## The 34-vs-18 discrepancy — RESOLVED (basis mismatch, not paydown)
|
||||
|
||||
Phase 2 flagged that ADR-0252's "34 bespoke surface organs" did not reproduce (18 found). Verified against the tree **at the ratification commit itself** (`1ccef491`):
|
||||
|
||||
| Measure | At ratification | Now (`8927c563`) |
|
||||
|---|---|---|
|
||||
| `def resolve_promotable_*` entry organs | **18** | **18** |
|
||||
| `generate/derivation/` tree entries | 33 | 32 (`.py` files) |
|
||||
|
||||
The count was 18 entry organs *on the day ADR-0252 was ratified*. So "34" never counted `resolve_promotable_*` functions — its plausible basis is the module count (~32–33: the 18 entry organs plus support modules — `accumulate`, `clauses`, `comparatives`, `compose`, `extract`, `verify`, `calendar_grounding`, …), or an earlier-generation organ inventory. Two consequences:
|
||||
|
||||
1. **No consolidation has occurred since ratification** (18 → 18; one module removed). The debt is *not* being paid down — Phase 2's alternative hypothesis is eliminated.
|
||||
2. **The governing ADR's central quantitative claim is unreproducible as stated.** The diagnosis stands regardless of whether the pile is 18 or 34 — but a ratified document whose headline number has no stated basis is exactly the documentation-debt class Finding F-4 tracks. The fix is one sentence in an ADR-0252 amendment stating the basis.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
32 modules, 6,765 lines. Each entry organ (`resolve_promotable_affine_fraction_delta`, `…goal_residual`, `…temporal_tariff`, etc.) recognizes one problem *shape* and publishes pre-composed candidates that the registry gates; support modules provide clause extraction, comparatives, composition, and verification. Governed by ADR-0251's standing prohibition: **no new bespoke per-case regex work** — new shape coverage that does not generalize is refused as debt (ADR-0252 §4: generalization ratio > 1 or it is a surface pile).
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- Misparse rate must be zero on the adversarial suite (refusal may be arbitrarily high — the safe failure mode) — contract: `runtime_contracts.md` §Adversarial suite; the `subtle_in_grammar` family (4 cases, all correct) proves the gate is not satisfied by refusing everything.
|
||||
- GSM8K lane shape `wrong == 0` with outcome-accounting completeness — pinned lane `math_teaching_corpus_v1` in `CLAIMS.md`; `math` suite exists.
|
||||
- Serving isolation: these organs feed math lanes; the chat serving path does not import them (deduction and curriculum serve through `proof_chain`/`curriculum_surface`, verified in Phase 2).
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `superseded-in-place`** — the ruling is explicit and this card does not relitigate it. The schema's liveness ⊥ fitness separation exists for exactly this state.
|
||||
|
||||
**Honest wrinkles:** the replacement is gated on the ADR-0252 §5 SME experiment, which has never returned a verdict — so "superseded" currently has no successor timeline at all; the organs are condemned and load-bearing indefinitely. GSM8K's demotion to diagnostic makes this pile low-urgency, which is presumably why the experiment stalled — but the *paradigm* the experiment validates governs all future comprehension, not just math. The stakes are mispriced by the demotion.
|
||||
|
||||
**Open questions:** amend ADR-0252 with the organ-count basis (→ ruling, one sentence); run §5 (→ ruling; the assessment's top open item).
|
||||
41
docs/assessment/20-component-cards/determine-phase.md
Normal file
41
docs/assessment/20-component-cards/determine-phase.md
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
# determine-phase — `generate/determine/`
|
||||
|
||||
**Kind:** zone→component descent · **Parent:** M3 (serve seam → M4) · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal`, flag-gated (**demoted** from the map's `live-serving`) · **Fitness:** `fit` · **Topology role:** runtime boundary
|
||||
|
||||
> The open-world DETERMINE gear: answer a question *from what the engine has realized in this conversation*, affirmatively or not at all. INV-30 is its constitution — it constructs only `Determined(answer=True)` or refuses; absence never refutes. It is the closest thing CORE has to "thinking with what it just learned," and it is off by default.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Eight modules, 1071 lines: `determine.py` (the gear), `consolidate.py` (Step D — one semi-naive layer of member/subset closure per idle tick, every hop verified by the proof_chain ROBDD before `realize_derived` writes it back; `member ∘ member` structurally unreachable; four declared strict-order predicates in `TRANSITIVE_PREDICATES`), `estimate.py` + `estimation_license.py` (Step E — converse-guess under an earned SERVE license, always disclosed `[approximate]`), `derived_close_proposals.py` (PR-2 bridge — eligible derived facts emitted as proposal-only artifacts), `render.py` (42 — honest rendering: SPECULATIVE grounds read "as I was told", never "verified").
|
||||
|
||||
## The gating map (all call sites verified)
|
||||
|
||||
| Entry | Caller | Gate | Default | Daemon forces? |
|
||||
|---|---|---|---|---|
|
||||
| `determine()` / `render_determination` | `chat/runtime.py:1364,1303` via `_accrue_in_turn` / `_surface_determination` | `accrue_realized_knowledge` | **False** | **No** |
|
||||
| `consolidate_once` | `chat/runtime.py:1041` via `idle_tick` | `consolidate_determinations` | False | **Yes** |
|
||||
| `estimate_converse` + `serve_license` | `chat/runtime.py:1420` | `estimation_enabled` | False | No |
|
||||
| derived-close proposal emission | `chat/runtime.py:1054` via `idle_tick` | `review_derived_close_proposals` | False | No |
|
||||
|
||||
A default-config serving turn **never** reaches this package. The map's `live-serving` label is wrong at this SHA: `live-internal`, flag-gated, exercised by lanes (`evals/close_derived_climb`, `evals/determination_closure`, `evals/determination_estimation`).
|
||||
|
||||
## Two discrepancies this table surfaces
|
||||
|
||||
1. **The daemon enables the consolidator but not the accruer.** `CONTINUOUS_LIFE_CONFIG_FLAGS` forces `consolidate_determinations=True` but leaves `accrue_realized_knowledge=False` — yet `realize_comprehension` (the only turn-path writer of realized facts) sits behind the *accrual* flag. Under the continuous-life config as coded, the idle consolidator may have nothing to consolidate unless realized facts arrive some other way. Either the flag set is incomplete, or consolidation under the daemon is deliberately dormant-until-seeded. **Unresolved; needs a ruling or a test that seeds through the lived path.**
|
||||
2. **A prior verification document is contradicted at this SHA.** `docs/research/architecture-assessment-verification-2026-07-25.md` §2 states `accrue_realized_knowledge` "is enabled by the production L10 process." The daemon's flag set does not include it. Either the doc anticipated a change that never landed, or the flag was later removed. Doc/code discrepancy — record, don't guess.
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- INV-30 (True-or-refuse; three visible `Determined` construction sites; typed rejection of forged `ClosedFrame`) — pin: `tests/test_architectural_invariants.py`.
|
||||
- Derived facts stay SPECULATIVE (sound inference never upgrades premise standing); replayable `Derivation` provenance (premise `structure_key`s + rule + `entailed` verdict) — contract: `runtime_contracts.md` §Idle consolidation.
|
||||
- wrong=0 by proof-gating: only `ENTAILED` conclusions are written — lane: `evals/close_derived_climb` (fixed-point climb, monotone, saturated tick consolidates 0).
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`.** The gear is small, typed, proof-gated, and honestly rendered. The INV-30/INV-31 firewall against closed-world machinery is among the cleanest boundaries in the repository.
|
||||
|
||||
**Honest wrinkles:** the entire "think with what you learned this session" capability — arguably the most *alive*-feeling behavior CORE has — is dark on every default path and only partially lit by the daemon. The `_accrue_in_turn` docstring calls DETERMINE and the readers "total (typed results, no raises)" while wrapping them in a broad defensive guard; the guard is honest backstop, but any exception it eats disappears silently into a no-op accrual (`_last_turn_accrual=None`) with no telemetry of the swallow. Flagged for Phase 4 as a small silent-failure surface inside an otherwise typed layer.
|
||||
|
||||
**Open questions:** the daemon flag-set completeness (→ ruling); should accrual-swallowed exceptions be counted in telemetry? (→ Phase 4)
|
||||
37
docs/assessment/20-component-cards/realize-phase.md
Normal file
37
docs/assessment/20-component-cards/realize-phase.md
Normal file
|
|
@ -0,0 +1,37 @@
|
|||
# realize-phase — `generate/realize/` (and the reader that feeds it)
|
||||
|
||||
**Kind:** zone→component descent · **Parent:** M3 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal`, flag-gated (**demoted** from the map's `live-serving`) · **Fitness:** `fit` (the phase) — with its *feeder* under held fabrication findings · **Topology role:** runtime boundary
|
||||
|
||||
> REALIZE — "integrate comprehended structure into the held self" (roadmap Step 3). If M1 is knowledge at rest and DETERMINE is answering from it, REALIZE is the moment comprehension becomes *held*: a reading of the user's words is written into session memory as a typed, SPECULATIVE, as-told record. It is the exact point where a misreading becomes a belief — which is why the fabrication findings matter most here.
|
||||
|
||||
## Disambiguation (a naming trap, recorded permanently)
|
||||
|
||||
`generate/realize/` (**knowledge** realization — this card) and `generate/realizer.py` (**surface** realization — the articulation renderer, M4) are unrelated mechanisms with confusable names. The map's `realize-phase` zone is the former. Any future search for "realizer" work must state which one it means.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Four modules, 645 lines: `realize.py` (425 — `realize_comprehension`, `realize_derived`, `Realized`/`NotRealized`/`RealizedRecord`, `Derivation`), `quantitative.py` (135 — `realize_quantitative`), `recall.py` (62 — `recall_realized`, DETERMINE's read path), `__init__.py`. Writes go through the INV-21-allowlisted vault writer; records carry `epistemic_status="speculative"` and derived records carry full `Derivation` provenance so replay re-derives and re-verifies.
|
||||
|
||||
Call graph (verified): `chat/runtime.py:1367` `_accrue_in_turn` → `realize_comprehension` (gate: `accrue_realized_knowledge=False` default, **not** daemon-forced); `generate/determine/consolidate.py:44` → `realize_derived` (gate: `consolidate_determinations`, daemon-forced); `derived_close_proposals.py:25` → `recall_realized`. Same demotion as determine-phase: no default serving turn reaches this package.
|
||||
|
||||
## The feeder — where the fabrication findings live
|
||||
|
||||
`_accrue_in_turn` comprehends via `generate/meaning_graph/reader.py::comprehend` and `generate/meaning_graph/relational.py::comprehend_relational` — **this is the two-grammars reader**: 19 constructions against the writer's 1739, overlap 6, and **fabricates on 22 more** (`every dog is a mammal` → `member(every_dog, mammal)`; `Given: furthermore; p implies q; p.` → `asserted(furthermore)` recited back as a served premise). Those findings are **measured and pinned in PR #138, held for ADR + ratification** — recorded here because REALIZE is where a fabricated reading would become a held belief, *not* re-discovered, *not* fixed.
|
||||
|
||||
The architecture note that follows from this placement: REALIZE itself is sound — it faithfully holds whatever the reader hands it, at honest SPECULATIVE standing. The truth defect class is entirely upstream (the reading), which is consistent with the project's fix-upstream doctrine and with holding the fixes for a serving-truth ADR.
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- INV-21 (allowlisted writer), INV-22/23 (SPECULATIVE default), INV-29 (only `vault/store.py` transitions status) — pins: `tests/test_architectural_invariants.py`.
|
||||
- Inline realization behavior — `tests/test_inline_realization.py` ("comprehensible declarative accrues; question determines; off-by-default leaves the turn untouched").
|
||||
- Predicate-generality — `relational.py:8`: the spine consumes relational comprehension unchanged; `realize_comprehension` is predicate-general (no per-predicate special case in the writer).
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`.** Small, typed, provenance-complete, honestly gated. The one structural risk is inherited, not local: the phase will hold whatever it is fed, so its integrity ceiling *is* the reader's fidelity — 19-wide today, fabricating on 22.
|
||||
|
||||
**Honest wrinkles:** as with determine-phase, the most life-like capability in the system defaults dark. `realize_quantitative` was not traced to a live caller in this pass (likely lane-side); unverified — Phase 4 should not count it live without a pointer.
|
||||
|
||||
**Open questions:** when the #138 fabrication ADR lands, should REALIZE gain a defensive assertion (refuse to hold a reading whose construction is outside the reader's verified inventory)? That would convert reader-inventory truth into a mechanical gate at the holding boundary. (→ Phase 4 / the fabrication ADR)
|
||||
|
|
@ -0,0 +1,24 @@
|
|||
# sensorium-falsification — `sensorium/environment/`
|
||||
|
||||
**Kind:** zone→component descent · **Parent:** M2 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-internal` (confirmed) · **Fitness:** `fit` · **Topology role:** deterministic replay surface (not a fusion layer, not a world model)
|
||||
|
||||
> Expected-versus-actual environmental falsification: the discipline of saying, before looking, what the world should show — and then checking. Its philosophical role is to keep afferent claims falsifiable without smuggling in a mutable world model, probabilistic confidence, or a learning path. It is deliberately the *narrowest possible* honest instrument.
|
||||
|
||||
## What it is / What it does
|
||||
|
||||
Four modules, 852 lines: `falsification.py` (compares already-compiled afferent units by merge key), `frame.py` (`ObservationFrame` / `ExpectedObservationFrame`), `scenario.py` (338 — scenario construction), `harness.py` (54). The module docstring states its non-goals in the first sentence: it does not compile raw signals, decode motor commands, fuse modalities, mutate Vault state, or create a world model. Verdict set is closed at two values — `SUPPORTED` (every expected slot matched by merge key, nothing unexpected) / `FALSIFIED` (anything missing, changed, or unexpected). Checksums via `sha256_json`. Neither verdict promotes anything to reviewed memory or mutates packs, vault, identity, or policy.
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- ADR-0211 contract frozen in `runtime_contracts.md`, including the forbidden list: no raw pixels/PCM/event streams/byte payloads/actuator traces in traces; no motor/efferent units in v1; no learned latents as substrate; no probabilistic confidence or tolerance thresholds in verdicts; no `generate/*` dependencies; no `ModalityRegistry.decode`.
|
||||
- The `sensorium` suite (21 files) runs — the zone's tests are scheduled, unlike M6's. would-fail-if-absent: yes.
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `fit`.** This is what a v1 should look like: closed verdicts, explicit non-goals in the first paragraph, checksummed evidence, and a hard wall against becoming a world model by accretion. When the sensorium track eventually approaches serving (M2's open entry-criterion question), this bench is the pattern the rest of the afferent stack should be held to.
|
||||
|
||||
**Honest wrinkles:** the map filed this zone under layer "L12" — a stratum that exists in no ratified document (flagged in the M2 card; taxonomy should either mint L12 deliberately or drop it). The bench compares *already-compiled* units, so its guarantees are conditional on the sensorium compiler's honesty — which is unmapped territory (the broader 59-module `sensorium/` package remains the largest built-and-disconnected mass in the tree). Zero-subsystem status in the map was an artifact of the final mapping wave, not of emptiness.
|
||||
|
||||
**Open questions:** none blocking; inherits M2's "entry criterion for serving" ruling.
|
||||
44
docs/assessment/20-component-cards/surface-selection.md
Normal file
44
docs/assessment/20-component-cards/surface-selection.md
Normal file
|
|
@ -0,0 +1,44 @@
|
|||
# surface-selection — the arms and the resolver
|
||||
|
||||
**Kind:** component (mechanism spanning `chat/runtime.py` + `core/cognition/surface_resolution.py`) · **Parent:** M4 · **Assessor:** Fable 5 (Phase 3)
|
||||
**Verified at:** `8927c563` (2026-07-27)
|
||||
**Liveness:** `live-serving` · **Fitness:** `strained` — refined from Phase 2 · **Topology role:** runtime boundary
|
||||
|
||||
> The mechanism that decides which of several candidate surfaces the user actually reads. Every truth property CORE claims converges here.
|
||||
|
||||
## Correction/refinement to the Phase 2 M4 card
|
||||
|
||||
Phase 2 recorded "no single place states the precedence order as an executable rule." **Partially wrong.** `core/cognition/surface_resolution.py` (494 lines) *is* a declared-precedence resolver — `resolve_surface` plus `_base_runtime_surface` ("select the runtime-owned base surface by declared precedence": served `response_surface` always wins; then `pre_decoration_surface`; then `canonical_surface`), `_truth_path_base` (canonical-first, the register-invariant identity folded into `trace_hash` — deliberately the *reverse* preference of the served base, each serving its own invariant), `_abstention_resolution`, `_grounded_open_hedge_resolution` (with an explicit admissibility predicate), and `_substrate_supreme`. The historical note in its docstring says why it exists: "historically these mutated one string in evaluation order."
|
||||
|
||||
**What remains true:** the resolver governs the *pipeline seam* (base/truth-path/abstention/hedge/fold axes). The **composer arms** upstream in `ChatRuntime.respond` are still ordered branches, verified at their call sites:
|
||||
|
||||
| Arm | Site | Gate |
|
||||
|---|---|---|
|
||||
| deduction_grounded_surface | `chat/runtime.py:1834` | `deduction_serving_enabled` (**ON**, ratified) |
|
||||
| curriculum_grounded_surface | `:1850` | `curriculum_serving_enabled` (OFF) |
|
||||
| pack_grounded_comparison | `:1871` | pack-grounding conditions |
|
||||
| narrative_grounded_surface | `:1898` | — |
|
||||
| example_grounded_surface | `:1915` | — |
|
||||
| pack relation-confirmation | `:1936` | — |
|
||||
| determination override (Step B-2) | `_surface_determination` → `:1303` | `accrue_realized_knowledge` (OFF) |
|
||||
| disclosed estimate (Step E) | `_surface_estimate` → `:1312–1340` | `estimation_enabled` (OFF) + earned SERVE license |
|
||||
| unknown-domain gate stub | `:188` `_UNKNOWN_DOMAIN_SURFACE` | gate fires |
|
||||
| grounded-open hedge (ADR-0254) | via resolver admissibility | — |
|
||||
| register decoration (ADR-0071/0077) | post-selection | never moves `trace_hash` |
|
||||
| logos-morph override | pipeline `:526` | answer-authority seam |
|
||||
|
||||
So the true architecture is **two strata**: ordered composer branches choose the *candidate*; the declared-precedence resolver reconciles *runtime vs pipeline vs abstention vs hedge*. The Phase 2 Third-Door candidate refines to: **extend the resolver's declared-precedence pattern upstream to the composer stratum**, rather than "create a resolver" — half the work is already done and proves the pattern fits this codebase.
|
||||
|
||||
## Contract & evidence
|
||||
|
||||
- Served-bytes-wins and its rationale are documented *with the falsifiable contract that catches regression*: preferring canonical would strip the register axis while `trace_hash` stays green — "the register-tour claims are the falsifiable contract that catches it" (`surface_resolution.py:113–116`). This is the sabotage test applied prospectively, in-code — worth naming as a model.
|
||||
- Estimate arm: `govern_response` returns STRICT for an unlicensed class → surface unchanged; `shape_surface` guarantees the `[approximate]` prefix because a converse guess is `UNVERIFIED_POSSIBLE`, never in APPROXIMATE's admissible set — "a wrong estimate is always a DISCLOSED wrong" (`:1312–1340`).
|
||||
- Selection-not-rewrite: determination and gate arms `replace(response, surface=…)`, retaining articulation/walk surfaces as evidence.
|
||||
|
||||
## Judgment
|
||||
|
||||
**Fitness: `strained`** — but more tractably than Phase 2 suggested. The resolver stratum is well-designed and self-documenting; the composer stratum is where arms accrete. With only deduction ON today, the *live* branch complexity is modest; the strain is prospective (each new capability adds an arm ahead of any declared order).
|
||||
|
||||
**Honest wrinkles:** the composer arms and the resolver are in different packages with different owners (`chat/` vs `core/cognition/`), so nothing structurally prevents a new arm from bypassing the resolver's disciplines. The M4 empty-string refusal wrinkle (typed refusal discarded at the public `str` boundary) sits in this same mechanism and remains the sharpest single dishonesty on the serving path — not because anything lies, but because the truth that exists goes unserved.
|
||||
|
||||
**Open questions:** declarative composer-arm precedence (→ Phase 4 Third-Door, refined); materialise `refusal_reason` (→ ruling, plumbing already landed).
|
||||
103
docs/assessment/30-gap-register.md
Normal file
103
docs/assessment/30-gap-register.md
Normal file
|
|
@ -0,0 +1,103 @@
|
|||
# The Gap Register
|
||||
|
||||
**Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27)
|
||||
**Standing:** This is CORE's first *live* gap register since `docs/gaps.md` closed its 26th entry. Proposal (for ruling): this register supersedes `docs/gaps.md`, which is marked historical; two dead registers plus a live one is worse than one live one.
|
||||
**Discipline:** A gap is an *absence the telos requires filled* with no explicit deferral ruling. Deferred-with-ruling is not a gap (scripture content is the model). Every entry carries evidence, its **deciding authority**, and a leverage rank. The register decides nothing.
|
||||
|
||||
---
|
||||
|
||||
## Tier A — Frontier-blocking (each blocks a ratified commitment or the telos itself)
|
||||
|
||||
### G-1 · The ADR-0252 §5 experiment has never returned a verdict
|
||||
**Layer:** M3 · **Leverage: 1 (highest in the assessment)**
|
||||
The ratified governing paradigm's single load-bearing empirical claim — can Cl(4,1) geometry carry relational structure the SME way — sits authorized (§8.4), scaffolded (two unmerged `rnd/` worktrees, tip `bed29a09` "formalize §5 experiment scaffolding"), and unrun. Until it returns GO or NO-GO, the §6 comprehension correction cannot be authorized, and the 18 condemned organs serve indefinitely with no successor path. A well-controlled NO-GO is *defined by the ADR as full credit* — the experiment is cheap to finish and expensive to leave open. GSM8K's demotion to diagnostic mispriced this: the paradigm governs **all** future comprehension, not math.
|
||||
**Evidence:** ADR-0252 §5/§8; worktree log; `M3` card. · **Authority:** execution (already authorized) + Shay's verdict ruling.
|
||||
|
||||
### G-2 · The #138 fabrications — *measured & pinned, fix held for ADR + ratification*
|
||||
**Layer:** M3 (locus: `generate/meaning_graph/reader.py`) → blast radius M4 · **Leverage: 2**
|
||||
`every dog is a mammal` → `member(every_dog, mammal)`; `Given: furthermore; p implies q; p.` → `asserted(furthermore)` recited back as a served premise. The reader fabricates on 22 constructions beyond its 19-wide verified inventory. The fixes are known — two of the 13 mutations — and are **deliberately held** because they change what CORE comprehends from user input: serving-path truth behavior, ADR + ratification territory. Entered here pre-labeled per standing instruction; never re-discovered, never fixed by this assessment.
|
||||
**Evidence:** PR #138 @ `c69f9948`; `realize-phase` card (incl. the defensive-gate option: refuse to *hold* a reading outside the verified inventory). · **Authority:** the fabrication ADR + Shay's ratification.
|
||||
|
||||
### G-3 · Reader inventory: 19 constructions against a 1739-construction writer, overlap 6
|
||||
**Layer:** M3 · **Leverage: 3**
|
||||
The comprehension frontier itself, measured. Standing ruling: close fabrications (G-2) **before** widening. The widening program after that is the largest single capability gap between CORE and its telos — and its *shape* depends on G-1's verdict (structure-mapping vs more constructions).
|
||||
**Evidence:** #138 inventory measurement; `M3` card capacity block. · **Authority:** sequenced rulings (G-2 → G-1 → widening plan).
|
||||
|
||||
### G-4 · CR-2 — the continuous life has no chooser
|
||||
**Layer:** M6 / Candidate Register · **Leverage: 4**
|
||||
Confirmed at component depth: drive objects exist (`DriveGradientMap` — constructed, never read; `ExertionMeter` — telemetry only); idle mechanisms exist (consolidation, proposal review, contemplation — each flag-gated, each doing one thing); **nothing ranks what matters next**. The daemon heartbeat advances `idle_tick` and nothing more ambitious. This is the AGI-grade conceptual absence: everything CORE does is chosen by the operator. Design work, not a flag flip.
|
||||
**Evidence:** `attention-allocation` + `always-on-process` cards; `02-layer-taxonomy.md` CR-2. · **Authority:** design + ruling (does the L10 process own an agenda, governed by what).
|
||||
|
||||
### G-5 · L10 proof debt — the soak has never produced an artifact, and nothing runs its pins
|
||||
**Layer:** M6 / MV · **Leverage: 5**
|
||||
The always-on process is built; the falsifiable harness (H1–H4, holds/bites pairs, vacuity-guarded) is built; **no recorded long-horizon artifact exists, no suite contains any `l10`/`always_on` test, no nightly cadence exists** — and the local-first/Mac-runner doctrine makes "nightly" itself need a ruling rather than a cron line. The still-owed ADR-0146 Phase-4 spike, in its modern form: run the soak, record the artifact, schedule the pins.
|
||||
**Evidence:** `M6` + `always-on-process` cards; suite-membership scan. · **Authority:** execution + MV suite ruling + cadence ruling.
|
||||
|
||||
### G-6 · F-6 — the lived learning loop is half-gated
|
||||
**Layer:** M6/M5 · **Leverage: 6**
|
||||
The daemon forces `consolidate_determinations` but not `accrue_realized_knowledge`; the only turn-path writer of realized facts sits behind the unforced flag. As coded, the continuous life may consolidate an empty set. Incomplete flag set, or intended dormancy — **neither is documented**, and a prior verification doc asserts the opposite of the code (C-5).
|
||||
**Evidence:** `CONTINUOUS_LIFE_CONFIG_FLAGS` (`chat/always_on_daemon.py:45-49`); `determine-phase` card gating table. · **Authority:** ruling (one flag + one sentence, or a documented dormancy rationale).
|
||||
|
||||
---
|
||||
|
||||
## Tier B — Enforcement & instrument debt (capability exists; the guarantee doesn't)
|
||||
|
||||
### G-7 · No orphaned-pin meta-check
|
||||
**Layer:** MV · **Leverage: 7**
|
||||
Suite tuples are hand-curated; a test file in zero suites is indistinguishable from one that runs everywhere. This is the *mechanism* by which G-5 happened. A meta-pin — every `tests/**/*.py` belongs to ≥1 suite or an explicit exclusion list — converts the doctrine "a pin in no suite never runs" into a failing test. Likely the highest-leverage *single mechanical change* in the repository.
|
||||
**Evidence:** `MV` card; the M6 case as the demonstration. · **Authority:** mechanical (small PR); no ruling needed.
|
||||
|
||||
### G-8 · No flag-default register
|
||||
**Layer:** cross-cut · **Leverage: 8**
|
||||
Seventeen capability flags default `False`; one is ratified ON (`deduction_serving_enabled`); three are daemon-forced (`persist_session_state`, `consolidate_determinations`, `strict_identity_continuity`). No document states the set, which defaults are deliberate posture vs accumulated hesitancy, or what evidence would flip each. The largest lever in the system, unregistered. The register format already exists in-repo: the ratified-ledger pattern (declare absence policy in the table, not the call site — ADR-0263 Rule 5).
|
||||
**Evidence:** `core/config.py` scan (Phase 2); daemon trio (Phase 3). · **Authority:** documentation PR + per-flag evidence bars set by ruling.
|
||||
|
||||
### G-9 · Enforcement pins unverified for three doctrine-level prohibitions
|
||||
**Layer:** M1 / MG · **Leverage: 9**
|
||||
(a) No verified failing pin for the no-approximate-recall law (would a cosine ranker actually fail a test?); (b) no pin that fails when a layer *bypasses* governance entirely (as distinct from governance working when called); (c) safety-pack non-swappability not verified as mechanically enforced. All three are law in `AGENTS.md`; law-enforced-by-review is weaker than law-enforced-by-test.
|
||||
**Evidence:** `M1`/`MG` cards (flagged, not resolved, in Phase 2–3). · **Authority:** verification pass, then mechanical PRs.
|
||||
|
||||
### G-10 · Curriculum SERVE is fully blocked by one engineering item, and its ledger doesn't exist
|
||||
**Layer:** M5 · **Leverage: 10**
|
||||
ADR-0264 §4.1: the 16-premise compilation cap holds every band to ≤16 entailed cases, so **no curriculum band can earn SERVE until query-scoping lands** — an engineering blocker gating a content problem that is itself quantified at 24×–73× under-fed. Downstream, `chat/data/curriculum_serve_ledger.json` is absent (the one honest `missing_ok=True` in production), and a committed ledger is necessarily an *earning* one — the outcome-mix ruling remains the binding constraint.
|
||||
**Evidence:** `M5` card; ADR-0264 §4.1; `chat/curriculum_serve_license.py:46`. · **Authority:** engineering (scoping) + outcome-mix ruling.
|
||||
|
||||
### G-11 · Identity enforcement has no stated authorization bar
|
||||
**Layer:** MG · **Leverage: 11**
|
||||
`identity_wave_gate` is off and "not authorized" — a deliberate posture. What's missing is the *criterion*: no document states what evidence would authorize live refusal. Scoring-without-blocking is an honest state only while the path to blocking is defined.
|
||||
**Evidence:** `MG` card; `runtime_contracts.md` identity contract. · **Authority:** ruling (set the bar).
|
||||
|
||||
---
|
||||
|
||||
## Tier C — One-line rulings (cheap to close; expensive only if left silent)
|
||||
|
||||
### G-12 · CR-3 efferent action — deferred, or out of telos?
|
||||
No system-level statement exists either way; the only adjacent text is one bench's v1 prohibition. The alignment posture is arguably *stronger* with action explicitly deferred — the ruling costs one line. **Authority:** ruling.
|
||||
|
||||
### G-13 · CR-4 temporal self-location — a stance before the L10 spike
|
||||
Determinism bans clocks; continuity implies lived time; the 24h+ no-drift requirement cannot be stated precisely without a stance on "now." The spike (G-5) should not be designed with an accidental answer. **Authority:** ruling (can be a paragraph in the soak's contract).
|
||||
|
||||
### G-14 · CR-1 attention governance — the one-page ADR
|
||||
Own `use_salience`, the two underived constants, the self-narrowing budget feedback, and the `InhibitionMask` disposition. Mechanism verified live and sound; only the governance is absent. **Authority:** ADR (one page — the card is its draft evidence base).
|
||||
|
||||
### G-15 · The daemon's ratifying ADR
|
||||
`chat/always_on_daemon.py` is unowned while ADR-0146 explicitly rejected the daemon shape it implements. Whatever the right answer, the record currently contradicts the code (see H-8). **Authority:** ADR amendment or a new short ADR.
|
||||
|
||||
---
|
||||
|
||||
## Tier D — Latent & carried-forward (recorded so nothing silently drops)
|
||||
|
||||
- **G-16 · ADR-0265's defect class survives in `_inflect_predicate`'s aspect arms** (`generate/templates.py:79`) — 10,530/16,146 template points, *not reachable today*. Latent, recorded from the prior arc; becomes live if aspect arms become reachable. **Authority:** the widening program (G-3) must clear it first.
|
||||
- **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling.
|
||||
- **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR.
|
||||
- **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Authority:** ADR amendment + re-count.
|
||||
- **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain).
|
||||
|
||||
---
|
||||
|
||||
## What is *not* in this register, and why
|
||||
|
||||
- **Scripture/theology content** — deferred by explicit ruling (2026-07-26); the model case for deferred-is-not-missing.
|
||||
- **Benchmark wins** — excluded by the completeness criterion itself (taxonomy §6): architectural distinctiveness is the target; benchmarks are downstream validation.
|
||||
- **Sociality, affect, full embodiment** — considered and not registered, with reasons, in the taxonomy's Candidate Register; importing them would violate Pillar II.
|
||||
- **Rust-backend default** — an open *question* with a stated blocker (crates.io unreachable under sandbox), not a gap; the measured case for urgency dissolved (0.22%).
|
||||
96
docs/assessment/31-hindrance-audit.md
Normal file
96
docs/assessment/31-hindrance-audit.md
Normal file
|
|
@ -0,0 +1,96 @@
|
|||
# The Hindrance Audit
|
||||
|
||||
**Assessor:** Fable 5 (Phase 4) · **Verified at:** `8927c563` (2026-07-27)
|
||||
**Discipline:** A hindrance is something *present* that works against the goal — a wrong solution for its underlying problem, a responsibility lodged in the wrong owner, a trade-off tuned instead of dissolved, or a record that misleads the next reasoner. Every entry carries evidence, a fitness verdict from the schema vocabulary, a **proposed better home**, and its deciding authority. Per standing rule: settled rulings are constraints — an entry may flag a ratified decision only on evidence, only for ruling, never as a unilateral recommendation to reverse. **This audit decides nothing.**
|
||||
|
||||
Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS.md protocol — not by ease.
|
||||
|
||||
---
|
||||
|
||||
## H-1 · License evidence counted on an independence assumption replay violates
|
||||
**Verdict:** `wrong-solution` (the *counting basis*, not the gating) · **Layers:** M5
|
||||
**Evidence:** Wilson lower-bound licensing (θ_SERVE=0.99, ADR-0175 lineage) assumes independent trials; a replay of the same sealed case is one trial observed again, not a new one. Measured consequence recorded in the prior arc: **21 of 25 ratified bands fall short** of their floor when replays are deduplicated.
|
||||
**Why it hinders:** the entire earned-license architecture — CORE's mechanism for *deserving* to serve — rests on the evidence count. An overstated count grants licenses the evidence doesn't support, which is precisely the failure the mechanism exists to prevent. The gate is right; the arithmetic feeding it is not.
|
||||
**Better home:** distinct-evidence counting at the seal boundary (count distinct cases; a replay refreshes, never increments), declared in the ledger schema the way ADR-0263 Rule 5 declares absence policy — in the table, not the call site.
|
||||
**Authority:** ADR amendment (0175/0263 lineage) + a re-count of the 25 bands. The re-count may demote licenses; that is the mechanism working.
|
||||
|
||||
## H-2 · Decoration in the runtime constructor — objects built and never read
|
||||
**Verdict:** decoration (fails the sabotage test) · **Layers:** M6/M3
|
||||
**Evidence:** `DriveGradientMap` constructed at `chat/runtime.py:716`, read nowhere. `InhibitionMask`/`InhibitionOperator` exported by `core/physics/__init__.py`, constructed on no path. Deleting either changes no output.
|
||||
**Why it hinders:** dead structure is not neutral — it is *testimony*. Both objects tell every reader that drive mapping and inhibition masking are live, and Phase 1 of this very assessment initially believed them. Decoration is how architecture lies without anyone lying.
|
||||
**Better home:** deletion (mastery algorithm step 2: the best part is no part), with their *intents* preserved where they belong — drive in the CR-2 design (G-4), the mask's disposition in the CR-1 ADR (G-14). If a future mechanism needs them, re-adding a deleted class is cheap; un-believing a phantom is not.
|
||||
**Authority:** mechanical PR + one line each in the CR-1/CR-2 decisions.
|
||||
|
||||
## H-3 · The typed refusal is constructed, then discarded at the public boundary
|
||||
**Verdict:** `strained` — truth built and unserved · **Layers:** M4
|
||||
**Evidence:** `InnerLoopExhaustion` carries reason, region, and per-step rejected-attempt evidence; `respond()`/`arespond()` convert it to `""` for the `str` contract, so a refusing turn serves the empty string with `refusal_reason == ""`. The plumbing to materialise (`CognitiveTurnResult.refusal_reason`, `compute_trace_hash` fold) already landed; `runtime_contracts.md` names it a residual awaiting a future ADR.
|
||||
**Why it hinders:** the honesty machinery is the product. A system whose refusals are richer than its answers, serving its refusals as nothing, undersells its own thesis on every hard turn.
|
||||
**Better home:** materialise into `ChatResponse.refusal_reason` (and a minimal honest surface), per the contract's own anticipation.
|
||||
**Authority:** small ADR — the chain already reserved the seam.
|
||||
|
||||
## H-4 · Composer-arm precedence is ordered branches above a declarative resolver
|
||||
**Verdict:** `strained` — a solved pattern not yet extended · **Layers:** M4
|
||||
**Evidence:** `core/cognition/surface_resolution.py` (494 lines) resolves the pipeline seam by *declared* precedence, self-documenting, with an in-code falsifiable contract for its own regression. Upstream, the composer arms (deduction `:1834`, curriculum `:1850`, pack/narrative/example/relation `:1871–:1936`, determination, estimate, gate, hedge) remain ordered branches across `chat/runtime.py`, in a different package with a different owner — nothing structurally prevents arm N+1 from bypassing the resolver's disciplines.
|
||||
**Why it hinders:** prospectively — each new serving capability adds an arm ahead of any declared order. With only deduction ON, the live complexity is modest; the time to dissolve the pattern is *before* the next three arms, not after.
|
||||
**Better home:** extend the resolver's declared-precedence pattern upstream to arm selection — the Third Door here is half-built and proven to fit this codebase.
|
||||
**Authority:** refactor ADR; low-risk while one arm is live.
|
||||
|
||||
## H-5 · Underived constants at the semantic center of generation
|
||||
**Verdict:** `strained` — Pillar I violation with an in-repo counterexample · **Layers:** M3/CR-1
|
||||
**Evidence:** `salience_top_k=16`, `inhibition_threshold=0.3` gate every token walk's candidate set; no recorded derivation exists for either. The contrast is instructive and in-repo: `admissibility_margin δ=0.4` was derived from the minimum observed margin of a characterization corpus (0.456), declared *falsifiable*, and survived a 20-case stratified attempt — the standard exists two config lines away.
|
||||
**Why it hinders:** "thresholds tuned for good-enough" at the exact point where the system decides what it may consider. Also the self-narrowing budget feedback (`stream.py:637`) — a real cognitive property nobody has named or justified.
|
||||
**Better home:** the CR-1 ADR (G-14) with an empirical derivation in the δ=0.4 style.
|
||||
**Authority:** ADR + a small characterization run.
|
||||
|
||||
## H-6 · The half-forced flag pair gating the lived learning loop
|
||||
**Verdict:** `misplaced` responsibility — a *set* decision made one flag at a time · **Layers:** M6/M5
|
||||
**Evidence:** F-6 (`05-phase3-findings.md`): the daemon forces the consolidator, not the accruer; the loop's writer and its consumer are gated independently, and only the consumer is on.
|
||||
**Why it hinders:** flags that must be coherent *as a set* are owned nowhere as a set. `CONTINUOUS_LIFE_CONFIG_FLAGS` is the right pattern (a named, documented flag *profile*) applied to the wrong subset.
|
||||
**Better home:** the flag-default register (G-8) with named profiles (one-shot / eval / continuous-life), each profile ruled as a unit.
|
||||
**Authority:** ruling on the accrual flag + the register PR.
|
||||
|
||||
## H-7 · The production ingest boundary lacks the trust contract its sibling has
|
||||
**Verdict:** `strained` — the standard exists and stops one layer short · **Layers:** M2
|
||||
**Evidence:** formation declares six boundaries — content-addressed in/out, no floats in hashed payloads, no pickle, an audit record per rejection. `ingest/gate.py`, facing untrusted user text in production, has the versor gate and the AGENTS.md trust-boundary defaults, but no comparable declared table.
|
||||
**Why it hinders:** asymmetric rigor invites the assumption that the un-tabled boundary is the less important one; it is the opposite.
|
||||
**Better home:** an M2 trust-boundary table in `runtime_contracts.md`, formation-style; hardening PRs only where the table exposes real deltas.
|
||||
**Authority:** documentation first; evidence decides whether code follows.
|
||||
|
||||
## H-8 · The record contradicts the code at three load-bearing points
|
||||
**Verdict:** `wrong-solution` as *record-keeping* — divergence that reasoners inherit · **Layers:** governance
|
||||
**Evidence:** (a) ADR-0146 rejects the daemon shape; an unowned daemon ships. (b) ADR-0252's headline "34 organs" has no reproducible basis (18 entry organs at the ratification commit itself; ~32 modules). (c) `architecture-assessment-verification-2026-07-25.md` asserts accrual "is enabled by the production L10 process"; the flag set says otherwise.
|
||||
**Why it hinders:** demonstrated, not hypothetical — this assessment's own Phase 0 inherited a stale-record error, and the 2026-07-25 doc (itself a *corrective* document) introduced one. Every divergence is a future wrong analysis.
|
||||
**Better home:** three one-paragraph amendments (ADR-0146 addendum owning the daemon or superseding the rejection; ADR-0252 basis sentence; a correction note on the 07-25 doc).
|
||||
**Authority:** docs PRs + ruling signatures.
|
||||
|
||||
## H-9 · Dead instruments still standing as if live
|
||||
**Verdict:** `superseded-in-place` (unratified) · **Layers:** MV/governance
|
||||
**Evidence:** `docs/gaps.md` — 26/26 closed, no entry from any 2026-06+ arc; `substrate-liveness-ratchet` — v5, stale since ~2026-05-24, all OPEN items L10-chained; ~130 analysis docs with no aggregator. The system map — the best macro artifact — is local, gitignored, 48 days stale, and was wrong precisely where the project moved fastest; its phantom "L12" stratum exists nowhere else.
|
||||
**Why it hinders:** an instrument that *looks* authoritative converts "I should check" into "I already checked." Phase 0's error was this mechanism operating on this assessment.
|
||||
**Better home:** this register supersedes `docs/gaps.md` (marked historical); the ratchet's 7 OPEN items migrate here (G-5 absorbs their L10 dependency); the map stays a regeneratable local index per D5, with "L12" dropped; the assessment directory becomes the standing ruled record, `verified_at`-stamped.
|
||||
**Authority:** ruling (one PR).
|
||||
|
||||
## H-10 · The demotion that mispriced the paradigm experiment
|
||||
**Verdict:** `strained` framing — a correct ruling casting an incorrect shadow · **Layers:** M3/governance
|
||||
**Evidence:** GSM8K was demoted to diagnostic (correct — the flags-and-benchmarks reasoning stands). The §5 SME experiment lives in GSM8K's neighborhood (`holdout_dev/v1`, math structures), so it inherited the demotion's priority — but its verdict governs the *comprehension paradigm for everything*, per ADR-0252's own §4 conformance bar.
|
||||
**Why it hinders:** the highest-leverage open item in the project (G-1) has been priced as math-lane housekeeping.
|
||||
**Better home:** none needed — G-1's execution *is* the fix; this entry exists so the mispricing mechanism is named and not repeated.
|
||||
**Authority:** already covered by G-1's ruling.
|
||||
|
||||
## H-11 · A silent-failure pinhole inside a typed layer
|
||||
**Verdict:** `strained` (small, cheap, principled) · **Layers:** M3
|
||||
**Evidence:** `_accrue_in_turn`'s broad guard converts any exception in the read→realize→determine chain into a no-op accrual with no telemetry (F-10). Defensible as a backstop; invisible as a signal — in the one layer whose constitution is "failures are typed, never silent" (INV-34).
|
||||
**Better home:** count the swallow (a telemetry field on `IdleTickResult`/turn accrual), not a behavior change.
|
||||
**Authority:** mechanical PR.
|
||||
|
||||
---
|
||||
|
||||
## Explicitly examined and cleared
|
||||
|
||||
For symmetry with the Candidate Register's "considered and not registered" — hindrance candidates this audit *rejects*:
|
||||
|
||||
- **The 18 derivation organs** — condemned but *ruled* to keep serving (`superseded-in-place` by explicit ruling); their continued service is governance working, not failing. The hindrance was their unreproducible count (H-8b), not their existence.
|
||||
- **Off-serve quarantines** (holographic vault, wave modules, `topological_reasoning`) — capacity that exists and cannot be used *by AST-pinned design*; legitimate research containment with failing-when-violated enforcement. An exit criterion would be nice (M1 card); the quarantine itself is fit.
|
||||
- **Pure-Python-by-default algebra** — measured as the correct posture: determinism is the product, `versor_condition` is 0.22% of turn time, and the urgency argument for Rust-by-default dissolved under measurement. The open parity question (blocked on network) is a question, not a hindrance.
|
||||
- **The five unreconciled articulations** — dissolved by the taxonomy (D1), not a standing hindrance; the residue is one stale Draft banner (folded into H-8's amendment batch).
|
||||
- **Flag-gated conservatism itself** — seventeen dark flags is not inherently hindrance; *unregistered* darkness is (G-8). The posture may be exactly right; the register exists so that judgment can be made deliberately.
|
||||
93
docs/assessment/40-assessment.md
Normal file
93
docs/assessment/40-assessment.md
Normal file
|
|
@ -0,0 +1,93 @@
|
|||
# The Assessment
|
||||
|
||||
**Phase 5 synthesis · Fable 5 · 2026-07-27 · verified at `forgejo/main` @ `8927c563`**
|
||||
**Method:** `docs/conceptualizing_engineering_mastery.md`, applied per `00-scope-and-method.md`. Everything below is traceable to a card or register entry; nothing below is new evidence.
|
||||
|
||||
---
|
||||
|
||||
## 1. The verdict, in one paragraph
|
||||
|
||||
CORE is an organism whose **skeleton is real, whose discipline is exceptional, and whose two deepest commitments are unproven in opposite directions**. The deterministic substrate, the typed learning boundary, the earned-license machinery, and the serving-truth discipline are built, enforced, and — where lanes exist — measured at `wrong=0` across every ratified surface. Against that: the *comprehension* the telos requires is measurably narrow (a reader 19 constructions wide feeding a writer 1739 wide, fabricating on 22), and the *continuity* the telos names is built but unproven (a real always-on process whose falsifiable soak has never produced an artifact and whose pins run in no suite). The system's five self-descriptions did not agree until this assessment reconciled them, and its fastest-moving month outran every instrument that was supposed to describe it. The distance to the telos is not mysterious, and it is not large in *kind*: it is five named frontiers (§5), most of which are blocked on rulings and proof-runs rather than on invention.
|
||||
|
||||
## 2. The cognitive cycle, stage by stage
|
||||
|
||||
From the evidence-bearing stage-coverage audit (Phase 2, corrected by Phase 3):
|
||||
|
||||
| Stage | State | The honest sentence |
|
||||
|---|---|---|
|
||||
| **listen** | covered (text) / uncovered (non-text) | The gate is sound and closure-checked; 59 sensorium modules wait disconnected with no entry criterion. |
|
||||
| **comprehend** | covered, *narrow* | Deduction decides real arguments at `wrong=0`; the general reader is 19 constructions wide, fabricates on 22, and its expert replacement is gated on an experiment that has never returned a verdict. |
|
||||
| **recall** | covered | Exact, verifiable, typed by standing; the tier that compounds and the tier that resets are different sets. |
|
||||
| **think** | covered | ROBDD entailment 716/716 against an independent oracle; proof-gated idle consolidation climbs to closure — when anything feeds it. |
|
||||
| **articulate** | covered | Selection-not-rewrite, disclosed estimates, typed refusal — served as an empty string (H-3). |
|
||||
| **learn** | covered, throttled | The single reviewed path is proven by pinned lanes; volume is 24×–73× under the floor and every curriculum band is capped at 16 entailed cases by one engineering item. |
|
||||
| **replay** | covered | Eleven SHA-pinned lanes failing CI on drift; the strongest sustained discipline in the repository. |
|
||||
| ***the runner of the cycle*** | **built, unproven** | The continuous life exists as code and has never been observed living longer than a test. Its learning loop is half-gated (F-6). Its guardian tests are orphaned (G-5/G-7). |
|
||||
|
||||
## 3. What is excellent — the standard the rest should be held to
|
||||
|
||||
Named deliberately, because a system this self-critical earns the right to have its strengths stated as findings (F-5, extended):
|
||||
|
||||
1. **The typed learning boundary** (M5) — durable-reviewed vs provisional-typed *dissolves* the autonomy-versus-safety trade-off; INV-21…30 make it law rather than intention. This is the Third Door executed, and it is CORE's most distinctive idea.
|
||||
2. **Selection, never rewrite** (M4) — the honest artifact survives every override; three surfaces serve three invariants and refuse to be conflated.
|
||||
3. **The non-hardening invariant** (M1) — no axiom flag can exist; the only closure in the architecture is mathematical. Most systems acquire an epistemic seal eventually; CORE structurally cannot.
|
||||
4. **Fail-closed evidence machinery** (MV) — unknown lane shapes refuse, degraded runs stamp themselves NON-CANONICAL, broken registries grant nothing, claims are machine-derived.
|
||||
5. **The falsification bench** (M2) — closed verdicts, first-sentence non-goals, checksummed evidence: what a v1 should look like.
|
||||
6. **In-code prospective sabotage tests** — `surface_resolution.py` documents the regression that would silently pass and names the contract that catches it. Doctrine written where it executes.
|
||||
|
||||
## 4. Where it stands — the macro picture
|
||||
|
||||
The layer table (Phase 2, with Phase 3 corrections applied): **no layer is `wrong-solution`.** M0 and M1 are `fit`. MG and M5 are `fit` with strained enforcement/throughput. M2, M3, M4, M6, MV are `strained` — and every strain decomposes into register entries with named authorities. At subsystem depth the organism is far more built than zone labels imply (137/205 live at the last full sweep); the genuinely unbuilt mass is concentrated where the map said — except that the map's centerpiece claim was stale, and the process it called unbuilt has existed since June 14.
|
||||
|
||||
Three structural facts dominate the macro picture:
|
||||
|
||||
- **Built-and-dark.** Seventeen capability flags default off; one is ratified on; three are daemon-forced. The gap between what CORE *is* and what CORE *does by default* is the widest gap in the system, and it is a governance artifact, not an engineering one (G-8).
|
||||
- **Capability outruns proof.** The daemon, Shape B+ persistence, the soak harness, the SME scaffolding — all built; none carried to verdict. The pattern is consistent enough to be cultural: *this project finishes machinery and defers ceremonies.* The mastery framework's step 4 (accelerate cycle time) applies to evidence loops, not just build loops.
|
||||
- **The record decays faster than the code.** Five articulations, three record/code contradictions, two dead registers, one stale map, and an assessment (this one) that had to correct itself twice using the only method that works. The cure is not more documents — it is `verified_at` stamps, failing pins for laws, and instruments that supersede rather than accumulate (H-8, H-9, G-7, G-9).
|
||||
|
||||
## 5. The five frontiers
|
||||
|
||||
Everything separating CORE-as-built from CORE-as-intended reduces to five named items. Nothing else on the registers is frontier; it is hygiene, enforcement, or ceremony.
|
||||
|
||||
1. **The reading** — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier.
|
||||
2. **The verdict** — run ADR-0252 §5 (G-1). One experiment, already authorized, already scaffolded, NO-GO defined as full credit. It decides the *shape* of frontier 1 and retires or redeems the 18 condemned organs. Highest leverage in the project.
|
||||
3. **The chooser** — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention.
|
||||
4. **The proof of life** — run the soak, record the artifact, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness, awaiting execution.
|
||||
5. **The throughput** — curriculum query-scoping, the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). The learning engine is sound and starved; this frontier is volume with integrity.
|
||||
|
||||
## 6. The recommended attack order
|
||||
|
||||
Sequenced by the mastery algorithm — scrub, delete, simplify, accelerate, automate last. Waves, not dates. Each item names its register entry; nothing here is new.
|
||||
|
||||
**Wave 0 — Scrub & rule** *(rulings, not builds; every later wave gets cheaper after it)*
|
||||
The one-line and one-page rulings: CR-3 efferent (G-12), CR-4 temporal stance (G-13), CR-1 attention ADR (G-14), the daemon's owning ADR (G-15), the F-6 accrual ruling (G-6), the three record/code amendments (H-8), the register supersessions (H-9). Plus the two ADR-track items that unblock frontiers: **run §5 to verdict (G-1)** and **the fabrication ADR (G-2)** — both yours to ratify, both fully staged.
|
||||
|
||||
**Wave 1 — Delete** *(the best part is no part)*
|
||||
`DriveGradientMap`, `InhibitionMask` (H-2); `docs/gaps.md` and the ratchet marked historical (H-9); the map's phantom L12; the stale blueprint banner. Small, but it removes false testimony — after Wave 1, the code stops telling readers things that aren't true.
|
||||
|
||||
**Wave 2 — Simplify & enforce** *(make the guarantees mechanical)*
|
||||
The orphaned-pin meta-check (G-7 — likely the single highest-leverage mechanical change); failing pins for the three unpinned laws (G-9); the flag-default register with named profiles (G-8/H-6); the M2 trust table (H-7); refusal materialisation (H-3/G-20); the accrual-swallow counter (H-11); composer-precedence extension (H-4) before the next serving arm lands, not after.
|
||||
|
||||
**Wave 3 — Accelerate the evidence loops** *(carry built machinery to verdict)*
|
||||
The L10 soak to a recorded artifact (G-5); the Wilson re-count with honest demotions (H-1/G-19); curriculum query-scoping and the earning ledger (G-10); the fabrication fixes landing under their ratified ADR, then the widening program (G-3) in whatever shape §5's verdict dictates.
|
||||
|
||||
**Wave 4 — Automate, last** *(only what Waves 0–3 proved)*
|
||||
Soak cadence under a ruled schedule; flag profiles flipped per their registered evidence bars; the contemplation/proposal machinery lit only once the loop it feeds is whole (F-6 resolved) and the chooser (G-4) exists to steer it. Automating before this point manufactures the mastery framework's "garbage at high speed" — an always-on process consolidating an empty set is precisely that, and CORE came within one flag of it.
|
||||
|
||||
## 7. The charter's four questions, answered
|
||||
|
||||
**Where does CORE stand on its cognitive cycle?** §2's table, evidence-bearing per stage. Seven of nine stages covered; comprehension covered-but-narrow; the runner built-but-unproven. The full decomposition: 9 layer cards, 8 component cards, every claim SHA-stamped.
|
||||
|
||||
**Is the layer model itself complete?** It is *now reconciled* — the five articulations were answering five different questions and are dissolved into the two-axis taxonomy (D1). Four candidate functions the telos implies and no document names are registered with their ruling questions (CR-1 turned out to be live-ungoverned; CR-2 is the real absence; CR-3/CR-4 are one-line rulings). One phantom stratum (L12) is flagged for deletion. Nothing else missing at the layer level survived the completeness criteria.
|
||||
|
||||
**What is the metadata?** The card schema (`03-card-schema.md`) — liveness ⊥ fitness, design ⊥ build, evidence with the would-fail-if-absent bit, capacity with ceilings, `verified_at` stamps — plus 17 filled cards and two registers. This directory is the instrument you asked for: each layer and component now has a philosophical intent, a functional contract, an implementation status with evidence, and a fitness judgment, in one greppable place that travels with the repository.
|
||||
|
||||
**What is hindering us?** Eleven audited entries (H-1…H-11), each with evidence, better home, and authority — headlined by the license-counting basis, decoration-as-testimony, and record/code divergence — plus five candidates examined and *cleared*, so the audit's negative space is as deliberate as its findings. No ratified ADR was found to be a wrong decision; three were found to have wrong *records*.
|
||||
|
||||
## 8. Method, and what it earned
|
||||
|
||||
Four phases, three self-corrections, one direction: **Phase 0 trusted a map and was wrong; Phase 2 read code and corrected it, then overstated twice; Phase 3 read deeper and corrected Phase 2; nothing in the chain was ever caught by re-reading documents.** The assessment's authority rests on exactly this: every liveness claim traces to an import, a call site, a flag default, or a pinned lane, at a named SHA — and where verification stopped short (suite membership of individual pins, Shape B+ exact coverage, the curriculum-formation bypass), the cards say so instead of rounding up.
|
||||
|
||||
That is also the maintenance contract for this directory: a card whose `verified_at` falls behind a load-bearing arc is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint. The registers supersede the dead instruments only for as long as they are kept live. The cheapest way to keep them live is Wave 2's mechanical enforcement; the most expensive way is another assessment like this one.
|
||||
|
||||
*— End of assessment. All deliverable sets complete: scope/method, ground truth, taxonomy, schema, 9 layer cards, 8 component cards, both registers, this synthesis.*
|
||||
23
docs/assessment/README.md
Normal file
23
docs/assessment/README.md
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
# docs/assessment/ — The Holistic Macro→Micro Assessment (2026-07-27)
|
||||
|
||||
A read-only, evidence-bearing assessment of CORE's cognitive-cycle design versus implementation fulfillment, conducted under `docs/conceptualizing_engineering_mastery.md` at `forgejo/main` @ `8927c563`. It changes no runtime behavior, fixes no defect, and decides nothing — it produces evidence and judgments for ruling.
|
||||
|
||||
**Start here:** [`40-assessment.md`](40-assessment.md) — the synthesis (the verdict, the five frontiers, the recommended attack order). Then the registers. The cards are the evidence base.
|
||||
|
||||
| File / dir | Phase | What it is |
|
||||
|---|---|---|
|
||||
| [`00-scope-and-method.md`](00-scope-and-method.md) | — | Charter, method, rules of engagement, phase/executor table |
|
||||
| [`01-phase0-ground-truth.md`](01-phase0-ground-truth.md) | 0 | Corpus triage, the five unreconciled articulations, system-map recovery |
|
||||
| [`02-layer-taxonomy.md`](02-layer-taxonomy.md) | 1 | The two-axis taxonomy: 7 macro layers + 2 cross-cuts over 33 zones; the Candidate Register (CR-1…4); completeness criteria |
|
||||
| [`03-card-schema.md`](03-card-schema.md) | 1 | The card metadata schema: liveness ⊥ fitness, design ⊥ build, sabotage-tested evidence, `verified_at` stamps |
|
||||
| [`10-layer-cards/`](10-layer-cards/) | 2 | Nine layer cards (M0–M6, MG, MV), every liveness label re-verified against code |
|
||||
| [`04-phase2-findings.md`](04-phase2-findings.md) | 2 | Stage-coverage audit; corrections to Phase 0; findings F-1…F-5 |
|
||||
| [`20-component-cards/`](20-component-cards/) | 3 | Eight component cards: the four zero-subsystem zones + always-on, derivation organs, surface selection, attention |
|
||||
| [`05-phase3-findings.md`](05-phase3-findings.md) | 3 | Corrections C-1…C-5; findings F-6…F-10; the consolidated Phase-4 seed list |
|
||||
| [`30-gap-register.md`](30-gap-register.md) | 4 | **The live gap register** — 20 entries, 4 tiers, each with evidence + deciding authority (proposes superseding `docs/gaps.md`) |
|
||||
| [`31-hindrance-audit.md`](31-hindrance-audit.md) | 4 | Eleven hindrances with fitness verdicts and better homes; five candidates examined and cleared |
|
||||
| [`40-assessment.md`](40-assessment.md) | 5 | The synthesis |
|
||||
|
||||
**Maintenance contract** (from §8 of the synthesis): a card whose `verified_at` falls behind a load-bearing arc is testimony, not evidence. Update cards when their subsystems move, or this directory becomes the next dead instrument it was built to replace.
|
||||
|
||||
**Standing note:** the PR #138 fabrication findings appear throughout as *measured & pinned, fix held for ADR + ratification* — recorded, never re-discovered, never fixed here, per explicit instruction.
|
||||
67
docs/conceptualizing_engineering_mastery.md
Normal file
67
docs/conceptualizing_engineering_mastery.md
Normal file
|
|
@ -0,0 +1,67 @@
|
|||
To abstract this methodology into a universal framework applicable to any discipline—whether you are writing a compiler, designing a skyscraper, or architecting a cognitive AI substrate—you must shift from *incremental optimization* to *first-principles reconfiguration*.
|
||||
|
||||
A truly rigorous engineering approach does not manage complexity; it eradicates it. This requires operating under hard constraints, governed by three core engineering pillars: **Semantic Rigor**, **Mechanical Sympathy**, and **The Third Door**. When integrated with a ruthless execution algorithm, this framework guarantees maximum yield for the time and capital spent.
|
||||
|
||||
Here is the generalized architecture of radical engineering improvement.
|
||||
|
||||
---
|
||||
|
||||
### Pillar I: Semantic Rigor (Defining the True Boundary)
|
||||
|
||||
Before a single line of code is written or a piece of metal is cut, the problem space must be scrubbed of all assumptions. Semantic Rigor demands that every requirement, term, and constraint is mathematically or physically defined. There are no thresholds tuned for “good enough”.
|
||||
|
||||
* **Assign Extreme Ownership:** Requirements are treated as inherently flawed. They must never come from an abstract "department" (e.g., "Safety requires..." or "Legal says..."). Every constraint must be attached to a specific, named human who can explain the fundamental law—physics, math, or strict logic—dictating it.
|
||||
* **Define the Absolute Limits:** Calculate the theoretical maximum efficiency or minimum mass allowed by the universe. If you are building software, what is the absolute minimum memory allocation required by the Turing machine? If you are building a bridge, what is the theoretical limit of the material's tensile strength? This mathematical limit becomes the baseline—anything less is a margin you must actively justify.
|
||||
|
||||
### Pillar II: Mechanical Sympathy (Aligning with the Medium)
|
||||
|
||||
A system that fights its underlying substrate is inherently inefficient. Mechanical Sympathy dictates that software must intimately understand the hardware it runs on, and hardware must be designed harmoniously with the physics of its environment.
|
||||
|
||||
* **Exploit the Environment:** Instead of building complex mechanisms to resist environmental challenges, use the environment to solve the problem. (e.g., SpaceX using its own cryogenic fuel to cool the engine, or software engineers utilizing zero-allocation architectures to bypass garbage collection entirely).
|
||||
* **Global Over Local Optimization:** You cannot optimize a component in a vacuum. A beautifully optimized microservice that requires massive network serialization overhead degrades the whole system. The boundary of the product is the *entire system*—co-optimize the macro structure.
|
||||
|
||||
### Pillar III: The Third Door (Bypassing Trade-offs)
|
||||
|
||||
Traditional engineering is obsessed with compromises: *"Do you want it fast, cheap, or reliable? Pick two."* Rigorous engineering rejects the dichotomy.
|
||||
|
||||
* **Orthogonal Problem Solving:** When facing a brutal design decision between two suboptimal paths, do not split the difference. If Path A increases mass and Path B reduces safety, the correct engineering answer is **Path C (The Third Door)**—a fundamental structural redesign that renders the trade-off irrelevant.
|
||||
* **Example in Practice:** Instead of choosing between a heavy flanged joint (Path A) or an expensive, slow-to-assemble gasket (Path B) for a high-pressure system, the Third Door is to eliminate the joint entirely via a single monolithic 3D print.
|
||||
|
||||
---
|
||||
|
||||
## The Universal Execution Algorithm
|
||||
|
||||
With the pillars established, the actual execution of the project must follow a strict, sequential algorithm. This sequence is a law of physics for project velocity; attempting step three before step two guarantees wasted capital and engineering hours.
|
||||
|
||||
1. **Scrub and Validate:** If you can't trace it to a physical/logical law, it's just a suggestion..
|
||||
Break the project down to its irreducible axioms. Force every team member to defend their requirements. If a constraint cannot be mathematically or logically proven, discard it.
|
||||
|
||||
|
||||
2. **Eradicate the Part or Process:** The best component is no component..
|
||||
Remove every feature, subsystem, or line of code that is not fundamentally critical to the core function. If you are not occasionally forced to add a part back in because the system broke, you are not deleting aggressively enough. Deletion eliminates manufacturing time, testing time, and failure modes simultaneously.
|
||||
|
||||
|
||||
3. **Simplify and Optimize:** Never optimize what shouldn't exist..
|
||||
Only after the system has been stripped to its absolute bare minimum do you begin optimizing. Streamline the remaining components to maximize their throughput, structural efficiency, or computational speed.
|
||||
|
||||
|
||||
4. **Accelerate Cycle Time:** Maximize the derivative of progress..
|
||||
Shorten the loop between design, testing, and failure. Do not spend months building a flawless prototype. Build a minimal viable component, test it to destruction to find its actual limits (not its theoretical ones), and feed that data immediately back into Step 1.
|
||||
|
||||
|
||||
5. **Automate:** Lock in the efficiency..
|
||||
Automation is the final step, never the first. Automating a complex, unoptimized process just ensures you manufacture garbage at high speed. Once the design is stripped down, proven, and stable, apply automation to scale production or execution infinitely.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## The Metrics of Success
|
||||
|
||||
To ensure a team is actually operating under this framework and not just paying lip service to it, measure the project's engineering culture against these binary checks:
|
||||
|
||||
| Metric | Traditional Engineering | First-Principles Engineering |
|
||||
| --- | --- | --- |
|
||||
| **Constraint Origin** | "Industry standard" or "Best practice" | Derived directly from fundamental math/physics limits |
|
||||
| **Response to Failure** | Add a new part (sensor, brace, patch) | Redesign the system to delete the failure mode entirely |
|
||||
| **System Boundary** | Team A builds X, Team B builds Y | The entire stack (hardware + software + factory) is one system |
|
||||
| **Optimization Goal** | Meeting the baseline requirements safely | Approaching the theoretical physical/logical limit |
|
||||
Loading…
Reference in a new issue