Track A of docs/assessment/50-execution-plan.md §6, executed under its own protocol. Criterion pre-registered at dfc394d2 (that commit carries the thresholds as importable constants and NO results); this commit carries the run. VERDICT: NO-GO. Full credit by ADR-0252 §5.4's own terms. docs/research/sme-experiment-verdict-797ebad5.md evals/structure_mapping/adr0252_s5/results/report-797ebad5.json deterministic_digest b3d9d27592e213104c51bcace7415fd819d722a1b06f7d0a5ca65bd5a234e4ec 53 cases, 1378 pairs, 0.0% non-convergence — no exception contributed to any number. RS-A (structure-only) separates perfectly, AUC 1.0000, margin +0.0867, and fails structure-sensitivity. RS-B (attribute-bearing) fails all three. The mechanism, measured rather than argued: the similarity quotient that would deliver attribute-invariance is the same quotient that annihilates structural contrast. An add-vs-subtract minimal pair — one entity, identical numbers, one relation kind changed — aligns at residual exactly 0.0, its two configurations being related by a PROPER rotation about e1. Sweeping the attribute weight (committed as diagnostic_sweep), structure-sensitivity fails at every setting, and every regime where the SME property survives is a regime where attributes contribute nothing: AUC 1.00 -> 0.97 -> 0.83 -> 0.69 as they begin to matter. Scope stated narrowly: this refutes H1 for embeddings that encode role-structure as point positions aligned by conformal Procrustes under similarity. The argument is about the quotient, so it generalises across that class; it does not refute every Cl(4,1) representation, and it does not touch the symbolic structure-mapping lane already in evals/structure_mapping/. Two findings that were on nobody's list: N-8 — the experiment was NOT unrun. rnd/sme-experiment-v2 @96e5f468is titled "Verdict: GO" and rnd/structure-mapping-experiment @fc9d0c14carries an earlier one. Neither survives inspection: attempt 1 leaked the S1-S4 label into the embedding; attempt 2 was blind and then took its separability from `except ValueError: res = 1000.0` — a solver exception counted as a distance — on a corpus 46/51 of which is outside holdout_dev/v1 with nothing marked, with duplicate graphs, colliding ids, and an extractor that was never committed. The register, the plan and the ADR all read the same absence and inherited the same error for nine days. A branch tip is not a record. G-21 — the math reader returns a selected graph for 5 of 500 holdout_dev/v1 cases (1.0%), all one skeleton; 24/150 on the public lane. That is why §5.1's four-structure corpus is not extractable, and it is a sharper measurement of the comprehension frontier than G-3's construction count. Registers updated: G-1 carries the verdict, H-10 confirmed and discharged, G-21 added, plan §0 gains N-8, Track A closed, synthesis frontier 2 rewritten. Pins: tests/test_adr_0252_s5_blindness.py, each observed red before green (label leak and duplicate-graph sabotage both caught). Off-serving — evals/structure_mapping/adr0252_s5/ is imported by no serving path, emits no answers, and changes no flag. [Verification]:797ebad5+ this branch — `uv run core test --suite smoke -q` 641 passed in 199.76s; `uv run ruff check` clean; report regenerated after lint with an unchanged digest.
15 KiB
The Assessment
Phase 5 synthesis · Fable 5 · 2026-07-27 · verified at forgejo/main @ 8927c563
Method: docs/conceptualizing_engineering_mastery.md, applied per 00-scope-and-method.md. Everything below is traceable to a card or register entry; nothing below is new evidence.
Amended: 2026-07-27 at ed06dd64 (Opus 5, Phase 6) · re-verified 2026-07-28 at 797ebad5 — sizing the execution plan produced seven corrections, four of which touch statements made here. They are applied inline and marked [AMENDED N-x]; the derivations are in 50-execution-plan.md §0. The chain in §8 continues: nothing in it was ever caught by re-reading documents.
1. The verdict, in one paragraph
CORE is an organism whose skeleton is real, whose discipline is exceptional, and whose two deepest commitments are unproven in opposite directions. The deterministic substrate, the typed learning boundary, the earned-license machinery, and the serving-truth discipline are built, enforced, and — where lanes exist — measured at wrong=0 across every ratified surface. Against that: the comprehension the telos requires is measurably narrow (a reader 19 constructions wide feeding a writer 1739 wide, fabricating on 22), and the continuity the telos names is built but unproven (a real always-on process whose falsifiable soak has never produced an artifact and whose pins run in no suite). The system's five self-descriptions did not agree until this assessment reconciled them, and its fastest-moving month outran every instrument that was supposed to describe it. The distance to the telos is not mysterious, and it is not large in kind: it is five named frontiers (§5), most of which are blocked on rulings and proof-runs rather than on invention.
2. The cognitive cycle, stage by stage
From the evidence-bearing stage-coverage audit (Phase 2, corrected by Phase 3):
| Stage | State | The honest sentence |
|---|---|---|
| listen | covered (text) / uncovered (non-text) | The gate is sound and closure-checked; 59 sensorium modules wait disconnected with no entry criterion. |
| comprehend | covered, narrow | Deduction decides real arguments at wrong=0; the general reader is 19 constructions wide, fabricates on 22, and its expert replacement is gated on an experiment that has never returned a verdict. |
| recall | covered | Exact, verifiable, typed by standing; the tier that compounds and the tier that resets are different sets. |
| think | covered | ROBDD entailment 716/716 against an independent oracle; proof-gated idle consolidation climbs to closure — when anything feeds it. |
| articulate | covered | Selection-not-rewrite, disclosed estimates, typed refusal — served as an empty string (H-3). |
| learn | covered, throttled | The single reviewed path is proven by pinned lanes; volume is 24×–73× under the floor and every curriculum band is capped at 16 entailed cases by one engineering item. |
| replay | covered | Eleven SHA-pinned lanes failing CI on drift; the strongest sustained discipline in the repository. |
| the runner of the cycle | built, evidenced as prose | The continuous life exists as code and has been observed over 5000 idle beats with a reboot at 2500, all four falsifiable gates passing — recorded as prose in a contract file, with no committed artifact and no pinned digest. Its learning loop is half-gated (F-6), and no pre-merge gate runs its pins. [AMENDED N-4] |
3. What is excellent — the standard the rest should be held to
Named deliberately, because a system this self-critical earns the right to have its strengths stated as findings (F-5, extended):
- The typed learning boundary (M5) — durable-reviewed vs provisional-typed dissolves the autonomy-versus-safety trade-off; INV-21…30 make it law rather than intention. This is the Third Door executed, and it is CORE's most distinctive idea.
- Selection, never rewrite (M4) — the honest artifact survives every override; three surfaces serve three invariants and refuse to be conflated.
- The non-hardening invariant (M1) — no axiom flag can exist; the only closure in the architecture is mathematical. Most systems acquire an epistemic seal eventually; CORE structurally cannot.
- Fail-closed evidence machinery (MV) — unknown lane shapes refuse, degraded runs stamp themselves NON-CANONICAL, broken registries grant nothing, claims are machine-derived.
- The falsification bench (M2) — closed verdicts, first-sentence non-goals, checksummed evidence: what a v1 should look like.
- In-code prospective sabotage tests —
surface_resolution.pydocuments the regression that would silently pass and names the contract that catches it. Doctrine written where it executes.
4. Where it stands — the macro picture
The layer table (Phase 2, with Phase 3 corrections applied): no layer is wrong-solution. M0 and M1 are fit. MG and M5 are fit with strained enforcement/throughput. M2, M3, M4, M6, MV are strained — and every strain decomposes into register entries with named authorities. At subsystem depth the organism is far more built than zone labels imply (137/205 live at the last full sweep); the genuinely unbuilt mass is concentrated where the map said — except that the map's centerpiece claim was stale, and the process it called unbuilt has existed since June 14.
Three structural facts dominate the macro picture:
- Built-and-dark. Twenty-eight of
RuntimeConfig's 32 boolean flags default off; one of the four on is ratified on; three of the off ones are daemon-forced. The gap between what CORE is and what CORE does by default is the widest gap in the system, and it is a governance artifact, not an engineering one (G-8). (This section originally said seventeen; two counts of one set differing by eleven is the finding — nothing distinguishes a capability flag from a policy flag.) [AMENDED N-7] - Capability outruns proof — and the ceremony, not the run, is what is missing. The daemon, Shape B+ persistence, the soak harness, the SME scaffolding — all built. The soak was even run, at 5000 beats, and passed; what it never got was a committed artifact and a pinned digest. The SME experiment was never run at all. The pattern is consistent enough to be cultural: this project finishes machinery and defers ceremonies. The mastery framework's step 4 (accelerate cycle time) applies to evidence loops, not just build loops. [AMENDED N-4]
- The record decays faster than the code. Five articulations, three record/code contradictions, two dead registers, one stale map, and an assessment (this one) that had to correct itself twice using the only method that works. The cure is not more documents — it is
verified_atstamps, failing pins for laws, and instruments that supersede rather than accumulate (H-8, H-9, G-7, G-9).
5. The five frontiers
Everything separating CORE-as-built from CORE-as-intended reduces to five named items. Nothing else on the registers is frontier; it is hygiene, enforcement, or ceremony.
- The reading — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier.
- The verdict — run, 2026-07-28: NO-GO (
docs/research/sme-experiment-verdict-797ebad5.md, criterion pre-registered before the run, artifact and digest committed). Geometric structure-mapping does not carry relational structure in the embedding class tested: the similarity quotient that would give attribute-invariance is the same one that annihilates structural contrast. So frontier 1's shape is settled — widening proceeds by more constructions, not by geometric SME — and the §6 build-authorization question is answerable either way. Two side findings: the experiment had already returned GO twice on unmerged branches, both unsound (N-8); and the math reader decides 1.0% ofholdout_dev/v1(G-21). Awaiting ratification. [AMENDED N-8] - The chooser — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention.
- The proof of life — the soak passed at 5000 beats; commit the artifact, pin the digest, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness and answered in prose that nothing can regress against. [AMENDED N-4]
- The throughput — the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). Curriculum query-scoping already landed (ADR-0264 R5, discharged 2026-07-26), so this frontier has no engineering prerequisite left: four bands earn SERVE the moment one content-policy ruling is made. The learning engine is sound and starved; this frontier is volume with integrity. [AMENDED N-5]
6. The recommended attack order
Sequenced by the mastery algorithm — scrub, delete, simplify, accelerate, automate last. Waves, not dates. Each item names its register entry; nothing here is new.
Wave 0 — Scrub & rule (rulings, not builds; every later wave gets cheaper after it) The one-line and one-page rulings: CR-3 efferent (G-12), CR-4 temporal stance (G-13), CR-1 attention ADR (G-14), the daemon's owning ADR (G-15), the F-6 accrual ruling (G-6), the three record/code amendments (H-8), the register supersessions (H-9). Plus the two ADR-track items that unblock frontiers: run §5 to verdict (G-1) and the fabrication ADR (G-2) — both yours to ratify, both fully staged.
Wave 1 — Delete (the best part is no part)
DriveGradientMap, InhibitionMask (H-2); docs/gaps.md and the ratchet marked historical (H-9); the map's phantom L12; the stale blueprint banner. Small, but it removes false testimony — after Wave 1, the code stops telling readers things that aren't true.
Wave 2 — Simplify & enforce (make the guarantees mechanical) The orphaned-pin meta-check (G-7 — likely the single highest-leverage mechanical change); failing pins for the three unpinned laws (G-9); the flag-default register with named profiles (G-8/H-6); the M2 trust table (H-7); refusal materialisation (H-3/G-20); the accrual-swallow counter (H-11); composer-precedence extension (H-4) before the next serving arm lands, not after.
Wave 3 — Accelerate the evidence loops (carry built machinery to verdict) The L10 soak to a recorded artifact (G-5); the Wilson re-count with honest demotions (H-1/G-19); curriculum query-scoping and the earning ledger (G-10); the fabrication fixes landing under their ratified ADR, then the widening program (G-3) in whatever shape §5's verdict dictates.
Wave 4 — Automate, last (only what Waves 0–3 proved) Soak cadence under a ruled schedule; flag profiles flipped per their registered evidence bars; the contemplation/proposal machinery lit only once the loop it feeds is whole (F-6 resolved) and the chooser (G-4) exists to steer it. Automating before this point manufactures the mastery framework's "garbage at high speed" — an always-on process consolidating an empty set is precisely that, and CORE came within one flag of it.
7. The charter's four questions, answered
Where does CORE stand on its cognitive cycle? §2's table, evidence-bearing per stage. Seven of nine stages covered; comprehension covered-but-narrow; the runner built-but-unproven. The full decomposition: 9 layer cards, 8 component cards, every claim SHA-stamped.
Is the layer model itself complete? It is now reconciled — the five articulations were answering five different questions and are dissolved into the two-axis taxonomy (D1). Four candidate functions the telos implies and no document names are registered with their ruling questions (CR-1 turned out to be live-ungoverned; CR-2 is the real absence; CR-3/CR-4 are one-line rulings). One phantom stratum (L12) is flagged for deletion. Nothing else missing at the layer level survived the completeness criteria.
What is the metadata? The card schema (03-card-schema.md) — liveness ⊥ fitness, design ⊥ build, evidence with the would-fail-if-absent bit, capacity with ceilings, verified_at stamps — plus 17 filled cards and two registers. This directory is the instrument you asked for: each layer and component now has a philosophical intent, a functional contract, an implementation status with evidence, and a fitness judgment, in one greppable place that travels with the repository.
What is hindering us? Twelve audited entries (H-1…H-12), each with evidence, better home, and authority — headlined by the license-counting basis, decoration-as-testimony, and record/code divergence — plus five candidates examined and cleared, so the audit's negative space is as deliberate as its findings. No ratified ADR was found to be a wrong decision; three were found to have wrong records, and a fourth divergence was later found inside the code (H-8d). [AMENDED N-3/N-6]
8. Method, and what it earned
Four phases, three self-corrections, one direction: Phase 0 trusted a map and was wrong; Phase 2 read code and corrected it, then overstated twice; Phase 3 read deeper and corrected Phase 2; nothing in the chain was ever caught by re-reading documents. The assessment's authority rests on exactly this: every liveness claim traces to an import, a call site, a flag default, or a pinned lane, at a named SHA — and where verification stopped short (suite membership of individual pins, Shape B+ exact coverage, the curriculum-formation bypass), the cards say so instead of rounding up.
That is also the maintenance contract for this directory: a card whose verified_at falls behind a load-bearing arc is testimony, not evidence — the same failure mode as the map, the ratchet, and the blueprint. The registers supersede the dead instruments only for as long as they are kept live. The cheapest way to keep them live is Wave 2's mechanical enforcement; the most expensive way is another assessment like this one.
Phase 6 continued the chain, and it is worth naming what caught what. Seven further corrections (50-execution-plan.md §0) came from reading core/cli_test.py, three workflow files, core/config.py, evals/l10_always_on/contract.md, and one ADR's own supersession banner. The banner case is the instructive one: ADR-0264 had already corrected itself, in place, directly under the superseded heading — and Phase 4 quoted the heading anyway. A document that corrects itself only helps a reader who reads past the heading. That is an argument for the failing pin over the amended paragraph, everywhere it is available.
— End of assessment. All deliverable sets complete: scope/method, ground truth, taxonomy, schema, 9 layer cards, 8 component cards, both registers, this synthesis.