Squashes the arc's work into one commit; the workflow-file edit it originally carried is excluded (see the end of this message). ## Lane 1 — Workbench recorded a proved answer as ungrounded With deduction_serving_enabled ratified ON (ADR-0256), workbench/api.py's live chat route builds a bare ChatRuntime(), so the deduction composer decides Workbench turns and stamps grounding_source="deduction" — but _coerce_grounding_source carried a hand-copied whitelist of the six pre-arc labels and silently rewrote anything else to "none". The runtime comment reasoned this was inert because "REPL turns do not flow through Workbench's CognitivePipelineRecord path". True, and irrelevant: the traffic flows the other way. Stale since 2026-07-24. Scope is one field. workbench/api.py:818 prefers TurnEvent.epistemic_state, which read epistemic_state_needed — honest. So the UNregistered path degraded honestly while the hand-copied whitelist asserted a falsehood; a second copy of a closed enum was worse than no copy. Hence registration AND derivation: GROUNDING_SOURCES exposes the Literal's members, and the coercion reads it. workbench-ui badges/tokens/snapshot follow; enumCoverage.test.ts forces atomicity. ## Lane 2 — the ratification ceremony The discovery loop was instrumented but not closed. teaching/ratification.py turns a reviewed decision into a chain record, a corpus commit, and a receipt. Its design turns on one observation: _ratified_rows DROPS unadmissible rows silently — correct when serving, a trap when ratifying, because the file grows, the commit lands, and the band count does not move. So the ceremony refuses to call an append a ratification until it has re-read the curriculum through the real loader and seen the chain arrive; a non-admitted append is rolled back. Validation is a pre-flight courtesy, admission is the proof. Arena queue entry and ledger reseal are deliberately NOT performed (bridge rule 1); the receipt names them. Front door: `core proposal-queue ratify`, a sibling of `review` rather than a flag on it. ## Lane 3 — structural closures - ADR-0263 gains rule 5: absence policy is DECLARED in CAPABILITY_LEDGERS, not passed at the call site. An AST-matched test fails if a serving path passes missing_ok again. - Deductive suite added WHOLE to the pre-push gate: 285 tests in 29s against smoke's 216 in 62s, so no coverage trade was needed. - Smoke/CI parity assertion made bidirectional. It was one-directional, and had drifted. - test_prior_surface_deduction_binding.py pins correction binding on the deduction path. The review's diagnosis did NOT reproduce — hash_surface moves in lockstep — so it pins what is there. Mutation-checked. - Domain-keyed ADR index over 312 flat-numbered files, explicitly partial. - Arc-close brief template, plus this arc's own brief filled in against it. ## Lanes 4 and 5 — two premises falsified by measurement, one of them mine Math 4.2: baseline reproduced (correct=5 wrong=0 refused=495); all four named cases traced to one seam with each gap isolated by one-variable probes. Then the number that changes the recommendation: the gap blocking case 0000 affects 1 case in 500, the 'than' gap blocking 0001 affects 2. ADR-0251's prohibition on per-case growth now rests on a count. No reader change made. CGA: versor_condition is 0.22% of a turn, not the "~10x proof latency" I claimed — that multiplied an isolated microbenchmark by a call count and compared it to a single verdict's latency. The real cost is geometric_product at 33,986 calls/turn (~73%) via cga_inner in search paths. The obvious closed form is NOT bit-exact (954/4000 in f32); backend.vault_recall's serial fold IS (3000/3000, worst-rel 0) and is the correct target. cargo test could not run — static.crates.io is denied by the sandbox network policy — so the Rust parity question stays open and the typestate lane is carried forward, not shipped uncompiled. ## Not landed: three lines owed to .github/workflows/smoke.yml The CI smoke gate is narrower than the local one — test_pack_draft_serve_boundary.py (ADR-0253 INV-33) has been local-only, unseen because the parity pin checked one direction. The edit was authored and rejected at push for lacking the `workflow` OAuth scope, so it is recorded as a named, dated PENDING_IN_CI exception rather than dropped: the assertion still fires on any new divergence, and a second guard fires once the three land. [Verification]: pre-push gates all green — smoke 236 passed, warmed_session 10 passed, deductive 285 passed. Ratification 14, ADR index 5, CLI suites 10. Grounding/epistemic sweep 741 passed 1 skipped. workbench-ui 598 passed across 73 files, tsc -b clean. capability index 11 passed, digest unchanged. Math holdout correct=5 wrong=0 refused=495. Committed chain corpora byte-unchanged after the tests that write to them. Environment caveat: the repo pins requires-python ==3.12.13, which uv cannot fetch for linux-x86_64, so `uv sync --locked` fails. All Python runs used a scratch venv on 3.12.11 with declared deps — not the locked universe, not the full ~12k suite. The pin was left untouched. Re-run on a 3.12.13 host before treating this as merge evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FduW6Krm3PPQv3P5iwBYtx
15 KiB
Independent verification of the Tier-S arc assessment (2026-07-25)
Input: an external architectural assessment of the deduction-serve generalization arc
(18 numbered observations) plus a synthesized five-pillar execution blueprint built on it.
Method: every load-bearing claim re-checked against the tree at 6a06cb2 by reading the
implicated code — and, where possible, by executing it. Result: roughly a third of the
blueprint's work items are falsified by the codebase; one flagship item is diagnosed at the wrong
layer; and the most urgent live defect appears in neither document.
The forward-discipline praise is accurate and was confirmed mechanically:
chat/data/deduction_serve_ledger.json holds 25 sealed bands / 18,000 cases / wrong=0, and
evals/capability_index/baseline.json reports breadth 13, wrong_total=0,
assert_mode_valid=true. What follows concerns only what the assessment proposed to do next.
1. The live defect neither document names
Workbench records a proved deduction answer as ungrounded. Verified by execution, not by reading:
_run_chat_turn('If it rains then the ground is wet. It rains. Therefore the ground is wet.')
surface = 'Given: if it rains then the ground is wet; it rains.
Your premises entail: the ground is wet.'
grounding_source = 'none' <- TurnEvent carried 'deduction'
epistemic_state = 'epistemic_state_needed'
The chain:
workbench/api.py:646— the live HTTP chat route calls_run_chat_turn(prompt).workbench/api.py:773— constructs a bareChatRuntime(), i.e.DEFAULT_CONFIG.core/config.py:397—deduction_serving_enabled: bool = True(ratified ON, 2026-07-24).chat/runtime.py:1748-1757— the composer commits and stampsTurnEvent.grounding_source = "deduction".workbench/api.py:714-719—_coerce_grounding_sourcewhitelists{pack, teaching, vault, partial, oov, none}and returns"none"for anything else.
workbench/api.py contains no occurrence of "deduction" or "curriculum".
Scope of the falsification — one field, not two. workbench/api.py:818 prefers
TurnEvent.epistemic_state over the grounding-source fallback, and the runtime stamps
epistemic_state_needed for the unregistered label. That is honest. So grounding_source is
falsified; epistemic_state is not.
The asymmetry is the useful part: the unregistered path degrades honestly
(epistemic_state_needed — "this label needs determination"), while the hand-copied whitelist
asserts a falsehood (none = ungrounded). The coercion is strictly worse than doing nothing.
This inverts the assessment's reading. chat/runtime.py:477-481 argues the unregistered
literal is "inert today" because "core chat REPL turns do not flow through Workbench's
CognitivePipelineRecord path." True, and irrelevant — the traffic flows the other way. Workbench
turns flow through the deduction path. That comment has been stale since the flag was ratified ON.
2. Falsified items — no session should be spent on these
| Claim | What the code says |
|---|---|
parents[n] audit is "a 30-minute grep with significant payoff" |
Audit run. 168 Path(__file__).parents[n] sites repo-wide; 0 resolve outside the repo root. The derived_close_proposals.py bug was isolated. |
| Realizer-guard exemption is undocumented at the guard site | False. Documented at both sites — chat/runtime.py:2371-2375, :3007-3010 — with the invariant and a worked example (pack open=VERB rejecting "the door is not open"). |
| T12 weekly-ledger entry is "completely missing" | False. docs/analysis/weekly-audit-2026-07-22-stragglers-todo.md:170 — an open pattern-watch item. The claim came from grepping commit messages only. |
assert_corpus_sound() has one impl; a second subject must reimplement it |
False. evals/curriculum_serve/runner.py:71 is already assert_corpus_sound(domain, cases), and build_report() calls it plus assert_provenance plus assert_anti_recall_coverage unconditionally (:105-108). Structural, not discipline-based. |
| Manual promotion sweeps need an automated Promotion Oracle | Already exists. docs/research/tier-s-housekeeping-2026-07-24.md §3: "not a one-time manual check… The sweep's proof is the passing gate, not a separate pass." ds-mem-0020 surfaced as a lane failure. |
accrue_realized_knowledge is "dark," a closing temporal-consistency window |
Misread. core/config.py:316-321 gates session memory, is False for eval/one-shot runtimes, and is enabled by the production L10 process. A deployment profile, not an epistemic debt clock. |
| Cross-subject proof could "emerge accidentally" | Structurally impossible. curriculum_surface.resolve_domain requires both terms in one subject's vocabulary and refuses ambiguous_reading rather than picking. Deduction has no subject notion at all. |
3. The flagship item is diagnosed at the wrong layer
The blueprint states that v5-VP and v6-EX "produce intermediate proof chains" that
render_entailment_verb "drops to print a terminal verdict."
generate/proof_chain/entail.py:125-190 decides entailment by canonicalizing
(P1 & … & Pn) -> Q and asking is_tautology — a BDD equivalence check, not a derivation.
EntailmentTrace (:79-99) carries outcome, reason, and five canonical BDD node keys —
opaque hashes. There are no intermediate steps in the object. The renderers (render.py:62+)
already emit every premise plus the verdict.
Nothing is dropped because nothing exists to drop. Multi-step articulation is engine work — proof-term extraction, or a derivation calculus over the ROBDD — with a real soundness surface. Scoping it as "design the trace object, then write the renderer" would produce a renderer with nothing to render.
4. The "second subject arena" premise
The blueprint orders its work as "harden lateral contracts before spinning up the second subject
arena." docs/research/curriculum-volume-quantification-2026-07-24.md §2 measures otherwise:
11 bands across 4 subjects already route questions today, every one is 24×–73× short of the
entailed-bucket floor, and the document concludes in bold that no served subject has a family
close enough to flag as "next."
There is no subject to spin up; the wiring is done. What physics has and the others lack is ratified content volume. That makes the Ratification Ceremony not a third-tier item but the binding one.
5. Confirmed, but materially rescoped
missing_ok— real and narrow. Exactly one production call site passesTrue(chat/curriculum_serve_license.py:46), correctly, becausechat/data/curriculum_serve_ledger.jsondoes not exist yet. Default isFalse. A small manifest closes it.- Telemetry spine — the proposed fix is unimplementable as written.
pipeline.py:799hasheshash_surface;runtime.py:1212stamps the served surface;TurnEventnever carrieshash_surface, sotrace_hashis not merely divergent from telemetry but not reconstructable from it. Both documents also miss the larger sibling flagged in-code atruntime.py:1235-1241: the external telemetry sink still receives the pre-override event (audit-ledger R7). - Discovery-yield stratification — ~4× the implied cost.
teaching/discovery_yield.py:67-70reads a single monotonicturn_count; there is no per-source counter to stratify on. Requires new persisted manifest fields. The denominator has always included pack/vault/none turns — deduction adds to a pre-existing dilution rather than creating it. - Band-level capability index —
index.pyscores viacoverage_geomean. Adding 25 near-1.0 band entries would mechanically inflate the headlinecapability_score, against the index's own anti-gaming property. Needs a scoring decision before a schema change. engine_refusedas a standing gate —tests/test_vocab_trigger_instrument.pyis already in thedeductivesuite. The instrument is defined as a before/after admissions counter; promoting a measurement to a gate means pinning an expected value, which is precisely the stale-exact-pin failure mode the assessment's own test-freshness item exists to prevent. The two requests collide._prior_surface— mostly false as diagnosed.pipeline.py:719sets_prior_surface = hash_surface, which is updated in lockstep with the logos-morph override (:526) and the speculative marker (:676); the realizer guard runs upstream. The register exclusion is deliberate and documented (:715-717). A pin is still cheap insurance.- Test freshness — the cited instance was already fixed in
a4d6868with a docstring explaining the failure mode. Remaining== (assertions intests/test_cli_test_suites.pypin argv prefixes, which is appropriate. A repo-wide de-pin sweep is low yield. - Math Phase 4.2 case-first — fully accurate, and the one strategic item both documents get
right. Adopt
docs/research/math-reader-phase-4-1-status-2026-07-24.mdverbatim.
6. The Apple Silicon / MLX substrate blueprint
A companion document proposes five bare-metal optimizations. Both of its foundational premises are contradicted by this repo's own measured report.
- "MLX executes tensor operations on the Neural Engine" is
README.md:82— a design aspiration.evals/reports/apple_uma_mechanical_sympathy_latest.md§10 explicitly disclaims it: "No MLX semantic-backend claim. No Neural Engine acceleration claim. No CoreML acceleration claim." MLX appears in exactly two files, both underbenchmarks/. No MLX in any runtime path. - "Rust remains the incumbent native algebra backend" —
algebra/backend.py:6-8: "Pure Python is the deterministic default. Rust is an explicit opt-in backend viaCORE_BACKEND=rust." The benchmark run recordscore_rs import: False,using_rust(): False, statuspython_fallback.
The algebra is running in pure Python/NumPy. The document proposes GPU shader work for a layer
that has not switched on the CPU kernel it already wrote — core-rs/src/cl41.rs already hardcodes
the Cl(4,1) Cayley table at compile time via bitmask anticommutation-parity and metric contraction
([u8;1024] index + [i8;1024] sign), which is its own recommendation 1 minus the GPU.
Corrected 2026-07-25 by measurement — see
cga-hot-path-measurement-2026-07-25.md. This section originally argued thatversor_condition(3 calls/turn × the benchmark'sp50 = 0.536 ms) was "~10× the entire proof latency". That was wrong: it multiplied an isolated microbenchmark by a call count and compared the product againstFrameVerdictTTFV, which is a single verdict's latency rather than a turn's. Measured directly,versor_conditionis 0.448 ms/turn against a 200 ms turn — 0.22%, so switching the Rust backend on recovers essentially nothing through that path.The profile does show CGA dominating (~73% of turn time), but through
cga_inner→geometric_product(33,986 calls/turn) in nearest-neighbour and salience search, not through the versor invariant. Both this assessment and the hardware blueprint aimed at the wrong function. The correct target, and why the obvious shortcut is not bit-safe, are in the measurement document.
| Proposal | Verdict |
|---|---|
Rust typestate (UnverifiedClaim → VersorClaim) |
Adopt. No PhantomData/typestate anywhere in the tree — greenfield, not "recently explored." The only proposal with zero determinism cost, and it needs no GPU. |
| SoA memory layout | Resequence. AoS confirmed (vault.rs:111,149). But the copy truth-table names the marshalling copy as the cost (extract_f32_slice, lib.rs:348), while batch paths are already zero-copy (lib.rs:282). SoA pays off only after zero-copy scalar bindings. |
Custom sparse .metal Cl(4,1) kernel |
Premature. The table exists. Rust is off by default for determinism; a GPU kernel makes determinism harder. Per AGENTS.md:276 the ubuntu-latest:host label is native macOS on one machine, so a shader would have a single validation host against a local-first merge bar. |
| MLX lazy kernel fusion of the versor invariant | Blocked on doctrine. JIT fusion reassociates float ops; versor_condition < 1e-6 is a non-negotiable (algebra/versor.py:100) and core-rs/tests/test_crdt_hash_parity.rs pins bit-exact parity. Needs a ratifying ADR, not a ticket. |
| bfloat16 asymmetric precision routing | Reject as specified. bf16 ε ≈ 7.8e-3 — four orders coarser than the 1e-6 gate; upcasting inside the shader cannot recover bits lost at storage. cga.rs:embed_point_raw computes 0.5*(r²∓1), cancellation-prone at bf16 for r ≈ 1 — exactly the unit-normalized regime. |
The actionable question the document missed: why is CORE_BACKEND=rust not the default?
Establish whether core_rs still holds bit-exact parity. (Attempted 2026-07-25 and blocked:
cargo cannot reach static.crates.io under the sandbox network policy. The question stays open
— the measurement above just removes the urgency the performance argument was supplying.)
None of this is on the critical path of the arc that just closed — FrameVerdict TTFV is 0.151 ms
with producer=proof_chain.entail, and the deduction/curriculum serving path is pure-Python ROBDD
that never touches CGA.
7. Resulting order of work
- Stop the provenance falsification. Red test; register
deduction/curriculumincore/epistemic_state.py+ the TS contract in one commit (enumCoverage.test.tsalready forces atomicity); derive_coerce_grounding_source's whitelist from the enum rather than a hand-copied set; delete the stale comment atchat/runtime.py:471-481. - Build the Ratification Ceremony —
proposal → chain record → corpus commit → arena queue entry → ledger reseal, reusingcore/ratified_ledger.py's sealers and thecore proposal-queueCLI, with the vocab-trigger admissions delta as the acceptance metric. - Batch the cheap closures —
missing_okmanifest; telemetry contract documented honestly plus audit-ledger R7 raised; the_prior_surfacepin; deductive suite into the pre-push gate; ADR domain index and an arc-close brief template. - Math Phase 4.2 case-first, in parallel — cases
0000/0001/0148/0082, noNEUTRAL_COUNT_VERBSgrowth (ADR-0251). - Substrate track, behind the above — Rust-default parity question, then typestate.
Deferred with stated reasons: multi-step articulation (re-scope as engine work), yield
stratification (blocked on manifest counters), band-level index (blocked on a scoring decision),
engine_refused gate (bucket empty), cross-subject ADR (resolve_domain fails closed).
Relates to project-generalization-arc.