core/docs/assessment/40-assessment.md
Shay ee481c1c92 docs(assessment): Phase 5 — the synthesis, and the directory index
40-assessment.md — the verdict in one paragraph (skeleton real, discipline
exceptional, the two deepest commitments unproven in opposite directions);
the cognitive cycle stage-by-stage with the honest sentence for each; what
is excellent, stated as findings; the macro picture (built-and-dark,
capability outruns proof, the record decays faster than the code); THE
FIVE FRONTIERS (the reading, the verdict, the chooser, the proof of life,
the throughput — everything else on the registers is hygiene, enforcement,
or ceremony); the recommended attack order in five waves sequenced by the
mastery algorithm (scrub/rule -> delete -> simplify/enforce -> accelerate
evidence loops -> automate last, with the warning that an always-on
process consolidating an empty set is "garbage at high speed" and CORE
came within one flag of it); the charter's four questions answered
directly; and the maintenance contract that keeps this directory from
becoming the next dead instrument.

README.md — the directory index and reading order.

The assessment is complete: 21 documents, four phases, three
self-corrections, every claim SHA-stamped at 8927c563.
2026-07-27 16:08:21 -07:00

12 KiB
Raw Blame History

The Assessment

Phase 5 synthesis · Fable 5 · 2026-07-27 · verified at forgejo/main @ 8927c563 Method: docs/conceptualizing_engineering_mastery.md, applied per 00-scope-and-method.md. Everything below is traceable to a card or register entry; nothing below is new evidence.


1. The verdict, in one paragraph

CORE is an organism whose skeleton is real, whose discipline is exceptional, and whose two deepest commitments are unproven in opposite directions. The deterministic substrate, the typed learning boundary, the earned-license machinery, and the serving-truth discipline are built, enforced, and — where lanes exist — measured at wrong=0 across every ratified surface. Against that: the comprehension the telos requires is measurably narrow (a reader 19 constructions wide feeding a writer 1739 wide, fabricating on 22), and the continuity the telos names is built but unproven (a real always-on process whose falsifiable soak has never produced an artifact and whose pins run in no suite). The system's five self-descriptions did not agree until this assessment reconciled them, and its fastest-moving month outran every instrument that was supposed to describe it. The distance to the telos is not mysterious, and it is not large in kind: it is five named frontiers (§5), most of which are blocked on rulings and proof-runs rather than on invention.

2. The cognitive cycle, stage by stage

From the evidence-bearing stage-coverage audit (Phase 2, corrected by Phase 3):

Stage State The honest sentence
listen covered (text) / uncovered (non-text) The gate is sound and closure-checked; 59 sensorium modules wait disconnected with no entry criterion.
comprehend covered, narrow Deduction decides real arguments at wrong=0; the general reader is 19 constructions wide, fabricates on 22, and its expert replacement is gated on an experiment that has never returned a verdict.
recall covered Exact, verifiable, typed by standing; the tier that compounds and the tier that resets are different sets.
think covered ROBDD entailment 716/716 against an independent oracle; proof-gated idle consolidation climbs to closure — when anything feeds it.
articulate covered Selection-not-rewrite, disclosed estimates, typed refusal — served as an empty string (H-3).
learn covered, throttled The single reviewed path is proven by pinned lanes; volume is 24×73× under the floor and every curriculum band is capped at 16 entailed cases by one engineering item.
replay covered Eleven SHA-pinned lanes failing CI on drift; the strongest sustained discipline in the repository.
the runner of the cycle built, unproven The continuous life exists as code and has never been observed living longer than a test. Its learning loop is half-gated (F-6). Its guardian tests are orphaned (G-5/G-7).

3. What is excellent — the standard the rest should be held to

Named deliberately, because a system this self-critical earns the right to have its strengths stated as findings (F-5, extended):

  1. The typed learning boundary (M5) — durable-reviewed vs provisional-typed dissolves the autonomy-versus-safety trade-off; INV-21…30 make it law rather than intention. This is the Third Door executed, and it is CORE's most distinctive idea.
  2. Selection, never rewrite (M4) — the honest artifact survives every override; three surfaces serve three invariants and refuse to be conflated.
  3. The non-hardening invariant (M1) — no axiom flag can exist; the only closure in the architecture is mathematical. Most systems acquire an epistemic seal eventually; CORE structurally cannot.
  4. Fail-closed evidence machinery (MV) — unknown lane shapes refuse, degraded runs stamp themselves NON-CANONICAL, broken registries grant nothing, claims are machine-derived.
  5. The falsification bench (M2) — closed verdicts, first-sentence non-goals, checksummed evidence: what a v1 should look like.
  6. In-code prospective sabotage testssurface_resolution.py documents the regression that would silently pass and names the contract that catches it. Doctrine written where it executes.

4. Where it stands — the macro picture

The layer table (Phase 2, with Phase 3 corrections applied): no layer is wrong-solution. M0 and M1 are fit. MG and M5 are fit with strained enforcement/throughput. M2, M3, M4, M6, MV are strained — and every strain decomposes into register entries with named authorities. At subsystem depth the organism is far more built than zone labels imply (137/205 live at the last full sweep); the genuinely unbuilt mass is concentrated where the map said — except that the map's centerpiece claim was stale, and the process it called unbuilt has existed since June 14.

Three structural facts dominate the macro picture:

  • Built-and-dark. Seventeen capability flags default off; one is ratified on; three are daemon-forced. The gap between what CORE is and what CORE does by default is the widest gap in the system, and it is a governance artifact, not an engineering one (G-8).
  • Capability outruns proof. The daemon, Shape B+ persistence, the soak harness, the SME scaffolding — all built; none carried to verdict. The pattern is consistent enough to be cultural: this project finishes machinery and defers ceremonies. The mastery framework's step 4 (accelerate cycle time) applies to evidence loops, not just build loops.
  • The record decays faster than the code. Five articulations, three record/code contradictions, two dead registers, one stale map, and an assessment (this one) that had to correct itself twice using the only method that works. The cure is not more documents — it is verified_at stamps, failing pins for laws, and instruments that supersede rather than accumulate (H-8, H-9, G-7, G-9).

5. The five frontiers

Everything separating CORE-as-built from CORE-as-intended reduces to five named items. Nothing else on the registers is frontier; it is hygiene, enforcement, or ceremony.

  1. The reading — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier.
  2. The verdict — run ADR-0252 §5 (G-1). One experiment, already authorized, already scaffolded, NO-GO defined as full credit. It decides the shape of frontier 1 and retires or redeems the 18 condemned organs. Highest leverage in the project.
  3. The chooser — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention.
  4. The proof of life — run the soak, record the artifact, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness, awaiting execution.
  5. The throughput — curriculum query-scoping, the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). The learning engine is sound and starved; this frontier is volume with integrity.

Sequenced by the mastery algorithm — scrub, delete, simplify, accelerate, automate last. Waves, not dates. Each item names its register entry; nothing here is new.

Wave 0 — Scrub & rule (rulings, not builds; every later wave gets cheaper after it) The one-line and one-page rulings: CR-3 efferent (G-12), CR-4 temporal stance (G-13), CR-1 attention ADR (G-14), the daemon's owning ADR (G-15), the F-6 accrual ruling (G-6), the three record/code amendments (H-8), the register supersessions (H-9). Plus the two ADR-track items that unblock frontiers: run §5 to verdict (G-1) and the fabrication ADR (G-2) — both yours to ratify, both fully staged.

Wave 1 — Delete (the best part is no part) DriveGradientMap, InhibitionMask (H-2); docs/gaps.md and the ratchet marked historical (H-9); the map's phantom L12; the stale blueprint banner. Small, but it removes false testimony — after Wave 1, the code stops telling readers things that aren't true.

Wave 2 — Simplify & enforce (make the guarantees mechanical) The orphaned-pin meta-check (G-7 — likely the single highest-leverage mechanical change); failing pins for the three unpinned laws (G-9); the flag-default register with named profiles (G-8/H-6); the M2 trust table (H-7); refusal materialisation (H-3/G-20); the accrual-swallow counter (H-11); composer-precedence extension (H-4) before the next serving arm lands, not after.

Wave 3 — Accelerate the evidence loops (carry built machinery to verdict) The L10 soak to a recorded artifact (G-5); the Wilson re-count with honest demotions (H-1/G-19); curriculum query-scoping and the earning ledger (G-10); the fabrication fixes landing under their ratified ADR, then the widening program (G-3) in whatever shape §5's verdict dictates.

Wave 4 — Automate, last (only what Waves 03 proved) Soak cadence under a ruled schedule; flag profiles flipped per their registered evidence bars; the contemplation/proposal machinery lit only once the loop it feeds is whole (F-6 resolved) and the chooser (G-4) exists to steer it. Automating before this point manufactures the mastery framework's "garbage at high speed" — an always-on process consolidating an empty set is precisely that, and CORE came within one flag of it.

7. The charter's four questions, answered

Where does CORE stand on its cognitive cycle? §2's table, evidence-bearing per stage. Seven of nine stages covered; comprehension covered-but-narrow; the runner built-but-unproven. The full decomposition: 9 layer cards, 8 component cards, every claim SHA-stamped.

Is the layer model itself complete? It is now reconciled — the five articulations were answering five different questions and are dissolved into the two-axis taxonomy (D1). Four candidate functions the telos implies and no document names are registered with their ruling questions (CR-1 turned out to be live-ungoverned; CR-2 is the real absence; CR-3/CR-4 are one-line rulings). One phantom stratum (L12) is flagged for deletion. Nothing else missing at the layer level survived the completeness criteria.

What is the metadata? The card schema (03-card-schema.md) — liveness ⊥ fitness, design ⊥ build, evidence with the would-fail-if-absent bit, capacity with ceilings, verified_at stamps — plus 17 filled cards and two registers. This directory is the instrument you asked for: each layer and component now has a philosophical intent, a functional contract, an implementation status with evidence, and a fitness judgment, in one greppable place that travels with the repository.

What is hindering us? Eleven audited entries (H-1…H-11), each with evidence, better home, and authority — headlined by the license-counting basis, decoration-as-testimony, and record/code divergence — plus five candidates examined and cleared, so the audit's negative space is as deliberate as its findings. No ratified ADR was found to be a wrong decision; three were found to have wrong records.

8. Method, and what it earned

Four phases, three self-corrections, one direction: Phase 0 trusted a map and was wrong; Phase 2 read code and corrected it, then overstated twice; Phase 3 read deeper and corrected Phase 2; nothing in the chain was ever caught by re-reading documents. The assessment's authority rests on exactly this: every liveness claim traces to an import, a call site, a flag default, or a pinned lane, at a named SHA — and where verification stopped short (suite membership of individual pins, Shape B+ exact coverage, the curriculum-formation bypass), the cards say so instead of rounding up.

That is also the maintenance contract for this directory: a card whose verified_at falls behind a load-bearing arc is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint. The registers supersede the dead instruments only for as long as they are kept live. The cheapest way to keep them live is Wave 2's mechanical enforcement; the most expensive way is another assessment like this one.

— End of assessment. All deliverable sets complete: scope/method, ground truth, taxonomy, schema, 9 layer cards, 8 component cards, both registers, this synthesis.