core/docs/assessment/README.md
Claude 199dc72d90
docs(specs): the four posture statements — and R-11's measurement dissolves its own question
The last docket item. Three decisions that cost nothing to make and something
real to leave silent, plus one proposal withdrawn because instrumentation showed
it could not be built as specified.

docs/specs/postures.md. Each posture names the CRITERION that would change it,
because an unstated boundary reads as an unexamined one and a future reader
cannot tell a deliberate limit from an accident.

P-1 (R-1 A, G-12) — efferent action deferred, in scope eventually. No efferent
surface beyond typed trace-folded tool operators until (i) the chooser exists
and is governed AND (ii) an efferent falsification bench exists. Both halves are
load-bearing: without (i) CORE acts with no account of what it should do next;
without (ii) it acts with no way to be shown wrong. ADR-0211's bench-level
prohibition is narrower than this and remains in force.

P-2 (R-5 A, G-11) — identity enforcement stays scoring-only until a named
held-out benign/adversarial corpus shows separation on the certified metric at a
floor named BEFORE the run. Same pre-registration discipline that made the §5
NO-GO full credit, for the same reason: a refusal gate authorized on a floor
chosen after seeing the numbers is not evidence, and identity refusal is
expensive to get wrong in both directions.

P-3 (R-6 A, G-17) — non-text ingest deferred, with the falsification bench as
the standard. A modality enters serving on the SAME terms text did: named
held-out corpus, holds/bites predicates, wrong=0-or-refuse. No modality is
admitted because the substrate can represent it. That makes the 59 sensorium
modules a capability awaiting evidence rather than an unexplained absence.

P-4 (R-11 B -> second ruling) — THE INTERIM FABRICATION GATE IS WITHDRAWN.

R-11 ruled "measure first, then re-ask". The measurement ran over 11,199
distinct serving-path inputs and 23,562 clauses and returned ZERO outside the
verified inventory — which is not a clean bill of health. It has a cause:

  atom_fact's template is "{p}" — one slot, no literal anchor. It matches ANY
  string, and it is one of the 19 verified constructions.

So "outside the verified inventory" is NOT A WELL-DEFINED PROPERTY. Nothing is
outside it. The proposed gate would have refused nothing while presenting as a
safety mechanism — a mechanism whose failure state is indistinguishable from its
success state, which is this repository's dominant defect class.

This reframes G-2, and the reframing is the finding. "every dog is a mammal" ->
member(every_dog, mammal) is not the reader ADMITTING an out-of-inventory
construction. The surface is admissible; the reader assigns it the WRONG
RELATION. The defect is in the mapping, not the admissibility set, and no gate
over the admissibility set can catch it.

Second ruling, delegated: option A is WITHDRAWN as unimplementable as specified,
NOT deferred — a deferred option is one that could be built later, and this one
cannot be built at all against the inventory as it stands. Option C is
operative. The fabrication ADR inherits the boundary question: either atom_fact
is narrowed so admissibility is decidable, or the guarantee moves from
admissibility to MAPPING CORRECTNESS. The evidence points at the second.

METHOD NOTE, recorded because the number nearly shipped wrong. Two earlier
passes were both invalid and both looked fine. The first matched templates
against whole multi-sentence inputs and reported 100% out-of-inventory when
every clause was in-inventory. The second misread the sentence-splitter's tuple
and measured "." 23,562 times. The catch-all was found only by a NON-VACUITY
CHECK — asserting the matcher could still say no to "most birds can fly" — and
it could not. A measurement that cannot fail is not a measurement, and that is
the same standard this arc applied to every pin it shipped.

Closes G-11, G-12, G-17. The entire adopted docket is now executed:
R-7 -> R-12 -> R-3+R-4 -> R-9+R-2 -> R-13 -> R-8 -> R-1/R-5/R-6/R-11.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wcw2pnMBwyvmNyQg4uPEt4
2026-07-28 08:26:48 +00:00

6.3 KiB
Raw Blame History

docs/assessment/ — The Holistic Macro→Micro Assessment (2026-07-27)

A read-only, evidence-bearing assessment of CORE's cognitive-cycle design versus implementation fulfillment, conducted under docs/conceptualizing_engineering_mastery.md at forgejo/main @ 8927c563. Phases 05 change no runtime behavior, fix no defect, and decide nothing — they produce evidence and judgments for ruling. Phase 6 adds the execution plan built on them and the ruling packet they require.

Start here: 40-assessment.md — the synthesis (the verdict, the five frontiers, the recommended attack order). Then 50-execution-plan.md for what follows from it, and 50-rulings.md for what is waiting on a decision. The registers rank the work; the cards are the evidence base.

File / dir Phase What it is
00-scope-and-method.md Charter, method, rules of engagement, phase/executor table
01-phase0-ground-truth.md 0 Corpus triage, the five unreconciled articulations, system-map recovery
02-layer-taxonomy.md 1 The two-axis taxonomy: 7 macro layers + 2 cross-cuts over 33 zones; the Candidate Register (CR-1…4); completeness criteria
03-card-schema.md 1 The card metadata schema: liveness ⊥ fitness, design ⊥ build, sabotage-tested evidence, verified_at stamps
10-layer-cards/ 2 Nine layer cards (M0M6, MG, MV), every liveness label re-verified against code
04-phase2-findings.md 2 Stage-coverage audit; corrections to Phase 0; findings F-1…F-5
20-component-cards/ 3 Eight component cards: the four zero-subsystem zones + always-on, derivation organs, surface selection, attention
05-phase3-findings.md 3 Corrections C-1…C-5; findings F-6…F-10; the consolidated Phase-4 seed list
30-gap-register.md 4 The live gap register — 23 entries (G-21, G-22, G-23 added during execution), 4 tiers, each with evidence + deciding authority; supersedes docs/gaps.md and the substrate-liveness ratchet per R-7
31-hindrance-audit.md 4 Fourteen hindrances with fitness verdicts and better homes; seventeen candidates examined and cleared — five in Phase 4, six in the first 2026-07-28 external-assessment triage, six in the second (of which one, the pack/domain seam, survived as G-23)
40-assessment.md 5 The synthesis
50-execution-plan.md 6 The execution plan — five waves + five frontier tracks over every G/H entry, with the dependency gates and the risks. §0 carries nine corrections to the assessment found while sizing and executing it; §2.1 carries the adopted ruling docket and its execution order
50-rulings.md 6 The ruling packet — R-1…R-14, each with evidence, options, a recommendation, and the exact diff that follows from each choice. RULED 2026-07-28: twelve adopted, R-10 and R-14 stricken, §5 NO-GO ratified. Carries the standing delegation under which residual sub-questions are decided, and marks each ruling EXECUTED as it lands
../specs/postures.md The posture statements (R-1/R-5/R-6/R-11) — efferent-action deferral, the identity-enforcement discrimination bar, non-text-ingest deferral, each with the criterion that would change it; plus P-4, the interim fabrication gate withdrawn because measurement dissolved its premise
../specs/flag_register.md The flag register (PR-5, R-3 + R-4) — all 32 RuntimeConfig booleans by class, governing ADR, and what evidence flips it; profiles as the unit of decision; §5 indexes every declared table in the repository and the pin that makes each true

Maintenance contract (from §8 of the synthesis): a card whose verified_at falls behind a load-bearing arc is testimony, not evidence. Update cards when their subsystems move, or this directory becomes the next dead instrument it was built to replace.

Execution status (2026-07-28). Phases 06 produced the evidence; the arc that followed executed everything needing no ruling, Wave 0 closed (twelve rulings adopted, two stricken, the §5 NO-GO ratified), and the entire adopted docket is now executed in its corrected order — R-7 → R-12 → R-3+R-4 → R-9+R-2 → R-13 → R-8 → R-1/R-5/R-6/R-11. Landed: Track A's §5 verdict (pre-registered), PR-1, PR-3+PR-3b, PR-4 (+pin 3), PR-4b, PR-5, PR-6, PR-7, PR-9, PR-11, PR-12, PR-14, R-12's ADR amendments, the posture statements, and five defect fixes (H-13, H-8e, G-22, G-23, plus pin 3's own union blind spot).

Closed: G-5…G-12, G-15, G-17, G-19, G-22, H-1, H-6, H-7, H-8 (all five instances), H-11, H-13. Added: N-8, N-9, G-21, G-22, G-23, H-13, H-14. Wave 4's F-6 half-gate is lifted. The gate is now four steps (smoke + warmed_session + deductive + teaching).

Three results worth reading before the registers. (1) ADR-0252 §5 returned NO-GO against a pre-registered criterion — the geometric structure-mapper is refuted for its embedding class, and §6's organ-retirement condition was amended because a NO-GO made it unsatisfiable as written. (2) Served capability shrank on purpose: the Wilson re-count took deduction from 25 licensed shape-bands to 4, and the curriculum outcome-mix rule dissolved the four licences N-5 predicted — in both cases wrong stayed 0, so no answer changed, only the claim attached to it. (3) R-11's instrumentation dissolved its own question: nothing is syntactically outside the verified inventory, because one of the 19 constructions matches any string — so the fabrications are mis-reads of admissible surfaces, and the fabrication ADR inherits the boundary question.

What remains is the next arc, not this one — see 50-execution-plan.md §Status and §2.1.

Standing note: the PR #138 fabrication findings appear throughout as measured & pinned, fix held for ADR + ratification — recorded, never re-discovered, never fixed here, per explicit instruction.