feat(rnd): ADR-0252 §5 run to a verdict — NO-GO, and the two prior GOs are void

Track A of docs/assessment/50-execution-plan.md §6, executed under its own
protocol. Criterion pre-registered at dfc394d2 (that commit carries the
thresholds as importable constants and NO results); this commit carries the run.

VERDICT: NO-GO. Full credit by ADR-0252 §5.4's own terms.
  docs/research/sme-experiment-verdict-797ebad5.md
  evals/structure_mapping/adr0252_s5/results/report-797ebad5.json
  deterministic_digest b3d9d27592e213104c51bcace7415fd819d722a1b06f7d0a5ca65bd5a234e4ec

53 cases, 1378 pairs, 0.0% non-convergence — no exception contributed to any
number. RS-A (structure-only) separates perfectly, AUC 1.0000, margin +0.0867,
and fails structure-sensitivity. RS-B (attribute-bearing) fails all three.

The mechanism, measured rather than argued: the similarity quotient that would
deliver attribute-invariance is the same quotient that annihilates structural
contrast. An add-vs-subtract minimal pair — one entity, identical numbers, one
relation kind changed — aligns at residual exactly 0.0, its two configurations
being related by a PROPER rotation about e1. Sweeping the attribute weight
(committed as diagnostic_sweep), structure-sensitivity fails at every setting,
and every regime where the SME property survives is a regime where attributes
contribute nothing: AUC 1.00 -> 0.97 -> 0.83 -> 0.69 as they begin to matter.

Scope stated narrowly: this refutes H1 for embeddings that encode role-structure
as point positions aligned by conformal Procrustes under similarity. The
argument is about the quotient, so it generalises across that class; it does not
refute every Cl(4,1) representation, and it does not touch the symbolic
structure-mapping lane already in evals/structure_mapping/.

Two findings that were on nobody's list:

N-8 — the experiment was NOT unrun. rnd/sme-experiment-v2 @ 96e5f468 is titled
"Verdict: GO" and rnd/structure-mapping-experiment @ fc9d0c14 carries an earlier
one. Neither survives inspection: attempt 1 leaked the S1-S4 label into the
embedding; attempt 2 was blind and then took its separability from
`except ValueError: res = 1000.0` — a solver exception counted as a distance —
on a corpus 46/51 of which is outside holdout_dev/v1 with nothing marked, with
duplicate graphs, colliding ids, and an extractor that was never committed. The
register, the plan and the ADR all read the same absence and inherited the same
error for nine days. A branch tip is not a record.

G-21 — the math reader returns a selected graph for 5 of 500 holdout_dev/v1
cases (1.0%), all one skeleton; 24/150 on the public lane. That is why §5.1's
four-structure corpus is not extractable, and it is a sharper measurement of the
comprehension frontier than G-3's construction count.

Registers updated: G-1 carries the verdict, H-10 confirmed and discharged, G-21
added, plan §0 gains N-8, Track A closed, synthesis frontier 2 rewritten.

Pins: tests/test_adr_0252_s5_blindness.py, each observed red before green
(label leak and duplicate-graph sabotage both caught).

Off-serving — evals/structure_mapping/adr0252_s5/ is imported by no serving
path, emits no answers, and changes no flag.

[Verification]: 797ebad5 + this branch — `uv run core test --suite smoke -q`
641 passed in 199.76s; `uv run ruff check` clean; report regenerated after lint
with an unchanged digest.
This commit is contained in:
Claude 2026-07-28 02:34:42 +00:00
parent 299c92be68
commit 7728a25dee
No known key found for this signature in database
9 changed files with 862 additions and 12 deletions

View file

@ -9,10 +9,12 @@
## Tier A — Frontier-blocking (each blocks a ratified commitment or the telos itself) ## Tier A — Frontier-blocking (each blocks a ratified commitment or the telos itself)
### G-1 · The ADR-0252 §5 experiment has never returned a verdict ### G-1 · The ADR-0252 §5 experiment **run to a NO-GO on 2026-07-28; awaiting ratification** **[AMENDED N-8]**
**Layer:** M3 · **Leverage: 1 (highest in the assessment)** **Layer:** M3 · **Leverage: 1 (highest in the assessment)**
The ratified governing paradigm's single load-bearing empirical claim — can Cl(4,1) geometry carry relational structure the SME way — sits authorized (§8.4), scaffolded (two unmerged `rnd/` worktrees, tip `bed29a09` "formalize §5 experiment scaffolding"), and unrun. Until it returns GO or NO-GO, the §6 comprehension correction cannot be authorized, and the 18 condemned organs serve indefinitely with no successor path. A well-controlled NO-GO is *defined by the ADR as full credit* — the experiment is cheap to finish and expensive to leave open. GSM8K's demotion to diagnostic mispriced this: the paradigm governs **all** future comprehension, not math. The ratified governing paradigm's single load-bearing empirical claim — can Cl(4,1) geometry carry relational structure the SME way — was authorized (§8.4), scaffolded, and is now **answered**: `docs/research/sme-experiment-verdict-797ebad5.md` records **NO-GO** against a criterion pre-registered at `dfc394d2` before the run, with a committed report artifact and a pinned `deterministic_digest`. A well-controlled NO-GO is *defined by the ADR as full credit*. The refutation is scoped: it holds for embeddings that encode role-structure as point positions and align them with conformal Procrustes under similarity, because the similarity quotient that would deliver attribute-invariance is the same quotient that annihilates structural contrast — measured, not argued (an `add`-vs-`subtract` minimal pair aligns at residual exactly `0.0`, the two configurations being related by a proper rotation). Across a sweep of attribute weights, every regime where the SME property survives is a regime where attributes contribute nothing.
**Evidence:** ADR-0252 §5/§8; worktree log; `M3` card. · **Authority:** execution (already authorized) + Shay's verdict ruling.
**This entry's original claim — "has never returned a verdict" — was wrong (N-8).** The experiment had been run **twice**, on `rnd/structure-mapping-experiment` @ `fc9d0c14` and `rnd/sme-experiment-v2` @ `96e5f468`, returning GO both times. Attempt 1 leaked the label into the embedding (disowned by its own successor); attempt 2 was blind and then harvested its separability from `except ValueError: res = 1000.0` — a solver exception counted as a distance — on a corpus that was 46/51 outside `holdout_dev/v1` with nothing marked, carried duplicate graphs and colliding ids, and whose extractor was never committed. Two GO verdicts sat unmerged and unratified for nine days while the register recorded the item as unrun.
**Evidence:** `docs/research/sme-experiment-verdict-797ebad5.md`; `docs/research/sme-experiment-preregistration-2026-07-28.md`; `evals/structure_mapping/adr0252_s5/`; `tests/test_adr_0252_s5_blindness.py`. · **Authority:** Shay's ratification of the verdict, and the §6 build-authorization ruling it unblocks.
### G-2 · The #138 fabrications — *measured & pinned, fix held for ADR + ratification* ### G-2 · The #138 fabrications — *measured & pinned, fix held for ADR + ratification*
**Layer:** M3 (locus: `generate/meaning_graph/reader.py`) → blast radius M4 · **Leverage: 2** **Layer:** M3 (locus: `generate/meaning_graph/reader.py`) → blast radius M4 · **Leverage: 2**
@ -113,6 +115,7 @@ Own `use_salience`, the two underived constants, the self-narrowing budget feedb
- **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling. - **G-17 · Non-text ingest** — 59 sensorium modules, no serving path, no entry criterion; projection heads do not exist. Position paper is honest about this. Needs either an entry criterion or an explicit deferral ruling (the falsification bench is the standard the track should be held to when it moves). **Authority:** ruling.
- **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR. - **G-18 · Identity-divergence curriculum may still bypass formation's gates** — known gap since 2026-05-17 (`teaching_order.md`); unverified at this SHA. **Authority:** Phase-3-style verification pass, then a routing PR.
- **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Note (Phase 6):** the exposure is **already pinned in-repo**`tests/test_volume_honesty.py` (ADR-0264 R9, in the local smoke gate) pins the 21-of-25 shortfall *"in BOTH directions"* and calls its inventory *"an EXPOSURE INVENTORY, not an approved baseline."* So the open work is **applying** the demotions, not discovering them, and that pin moves in the same PR. **Authority:** ADR amendment + re-count, authorized by R-13. - **G-19 · Wilson/replay evidence shortfall** — 21/25 ratified bands short if replays were counted as independent trials (see H-1 for the mechanism). Recorded here as *evidence debt on existing licenses*; the counting fix is the hindrance entry. **Note (Phase 6):** the exposure is **already pinned in-repo**`tests/test_volume_honesty.py` (ADR-0264 R9, in the local smoke gate) pins the 21-of-25 shortfall *"in BOTH directions"* and calls its inventory *"an EXPOSURE INVENTORY, not an approved baseline."* So the open work is **applying** the demotions, not discovering them, and that pin moves in the same PR. **Authority:** ADR amendment + re-count, authorized by R-13.
- **G-21 · The math reader decides 1.0% of `holdout_dev/v1`** *(new, 2026-07-28)* — measured while building the §5 corpus: `parse_and_solve` returns a selected graph for **5 of 500** held-out cases, all carrying the same relational skeleton (`compare_multiplicative`); on the public lane, 24/150. This is a sharper measurement of the comprehension frontier than G-3's construction count, on the corpus the project already treats as its held-out standard, and it is why ADR-0252 §5.1's four-structure corpus is not extractable today. Distinct from G-3 (which counts *constructions* the general reader admits); this counts *cases decided* on a standing eval corpus. **Authority:** the widening program (G-3) + a ruling on whether holdout decision-rate becomes a tracked lane metric.
- **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain). - **G-20 · The `refusal_reason` materialisation** — typed refusal evidence exists and is discarded at the public `str` boundary; the plumbing for materialisation already landed. Cross-listed as H-3. **Authority:** small ADR (anticipated by the ADR-0024 chain).
--- ---

View file

@ -78,6 +78,7 @@ Ranked by leverage (cognitive/structural load removed ÷ effort), per the AGENTS
**Evidence:** GSM8K was demoted to diagnostic (correct — the flags-and-benchmarks reasoning stands). The §5 SME experiment lives in GSM8K's neighborhood (`holdout_dev/v1`, math structures), so it inherited the demotion's priority — but its verdict governs the *comprehension paradigm for everything*, per ADR-0252's own §4 conformance bar. **Evidence:** GSM8K was demoted to diagnostic (correct — the flags-and-benchmarks reasoning stands). The §5 SME experiment lives in GSM8K's neighborhood (`holdout_dev/v1`, math structures), so it inherited the demotion's priority — but its verdict governs the *comprehension paradigm for everything*, per ADR-0252's own §4 conformance bar.
**Why it hinders:** the highest-leverage open item in the project (G-1) has been priced as math-lane housekeeping. **Why it hinders:** the highest-leverage open item in the project (G-1) has been priced as math-lane housekeeping.
**Better home:** none needed — G-1's execution *is* the fix; this entry exists so the mispricing mechanism is named and not repeated. **Better home:** none needed — G-1's execution *is* the fix; this entry exists so the mispricing mechanism is named and not repeated.
**Status (2026-07-28): confirmed and discharged.** G-1 ran to a recorded NO-GO in hours, not an arc (`docs/research/sme-experiment-verdict-797ebad5.md`) — which is the mispricing demonstrated rather than asserted. The entry also understated itself: the experiment had *already been run twice* and returned GO both times, and both results sat unmerged on `rnd/` branches while the highest-leverage item in the project was recorded as unstarted. The demotion's shadow fell not only on the priority but on the *retrieval* of work already done.
**Authority:** already covered by G-1's ruling. **Authority:** already covered by G-1's ruling.
## H-11 · A silent-failure pinhole inside a typed layer ## H-11 · A silent-failure pinhole inside a typed layer

View file

@ -51,7 +51,7 @@ Three structural facts dominate the macro picture:
Everything separating CORE-as-built from CORE-as-intended reduces to five named items. Nothing else on the registers is frontier; it is hygiene, enforcement, or ceremony. Everything separating CORE-as-built from CORE-as-intended reduces to five named items. Nothing else on the registers is frontier; it is hygiene, enforcement, or ceremony.
1. **The reading** — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier. 1. **The reading** — close the fabrications (G-2, held for your ADR), then widen from 19 under whatever paradigm G-1's verdict selects. This is the intelligence frontier.
2. **The verdict**run ADR-0252 §5 (G-1). One experiment, already authorized, already scaffolded, NO-GO defined as full credit. It decides the *shape* of frontier 1 and retires or redeems the 18 condemned organs. Highest leverage in the project. 2. **The verdict****run, 2026-07-28: NO-GO** (`docs/research/sme-experiment-verdict-797ebad5.md`, criterion pre-registered before the run, artifact and digest committed). Geometric structure-mapping does not carry relational structure in the embedding class tested: the similarity quotient that would give attribute-invariance is the same one that annihilates structural contrast. So frontier 1's *shape* is settled — widening proceeds by more constructions, not by geometric SME — and the §6 build-authorization question is answerable either way. Two side findings: the experiment had already returned GO twice on unmerged branches, both unsound (N-8); and the math reader decides 1.0% of `holdout_dev/v1` (G-21). Awaiting ratification. **[AMENDED N-8]**
3. **The chooser** — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention. 3. **The chooser** — CR-2 (G-4). The continuous life needs something to want; today every goal is operator-supplied and the drive machinery is decoration. This is the only frontier requiring genuine design invention.
4. **The proof of life** — the soak *passed* at 5000 beats; commit the artifact, pin the digest, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness and answered in prose that nothing can regress against. **[AMENDED N-4]** 4. **The proof of life** — the soak *passed* at 5000 beats; commit the artifact, pin the digest, schedule the pins, rule on the half-gated loop (G-5, G-6). The telos's own claim, made falsifiable by CORE's own harness and answered in prose that nothing can regress against. **[AMENDED N-4]**
5. **The throughput** — the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). Curriculum query-scoping **already landed** (ADR-0264 R5, discharged 2026-07-26), so this frontier has no engineering prerequisite left: four bands earn SERVE the moment one content-policy ruling is made. The learning engine is sound and starved; this frontier is volume with integrity. **[AMENDED N-5]** 5. **The throughput** — the ledger, the outcome-mix ruling, the Wilson re-count (G-10, H-1, G-19). Curriculum query-scoping **already landed** (ADR-0264 R5, discharged 2026-07-26), so this frontier has no engineering prerequisite left: four bands earn SERVE the moment one content-policy ruling is made. The learning engine is sound and starved; this frontier is volume with integrity. **[AMENDED N-5]**

View file

@ -3,7 +3,7 @@
**Planner:** Opus 5 · 2026-07-27 · verified at `forgejo/main` @ `ed06dd64` **Planner:** Opus 5 · 2026-07-27 · verified at `forgejo/main` @ `ed06dd64`
**Governs:** everything in `30-gap-register.md` (G-1…G-20) and `31-hindrance-audit.md` (H-1…H-12), sequenced by `40-assessment.md` §6. **Governs:** everything in `30-gap-register.md` (G-1…G-20) and `31-hindrance-audit.md` (H-1…H-12), sequenced by `40-assessment.md` §6.
**Method:** `docs/conceptualizing_engineering_mastery.md` — scrub → **delete** → simplify/enforce → accelerate → automate last. Nothing is automated that Waves 03 have not proven. **Method:** `docs/conceptualizing_engineering_mastery.md` — scrub → **delete** → simplify/enforce → accelerate → automate last. Nothing is automated that Waves 03 have not proven.
**Status:** Wave 0 in progress. Waves 14 and Tracks AE are PROPOSED and unstarted. **Status (2026-07-28):** Wave 0's builds are **done** — PR-0 merged (#140), PR-1 landed, R-10 discharged by merging #138. **Track A is run and returned NO-GO** (`docs/research/sme-experiment-verdict-797ebad5.md`), awaiting ratification. R-1…R-14 remain PENDING and gate Waves 14; Tracks BE are unstarted behind them.
--- ---
@ -20,6 +20,7 @@ The assessment's authority rests on its self-correction chain — *"nothing in t
| N-5 | G-10 "SERVE blocked by the 16-premise cap" | **the cap was removed 2026-07-26**; PR-13 withdrawn | | N-5 | G-10 "SERVE blocked by the 16-premise cap" | **the cap was removed 2026-07-26**; PR-13 withdrawn |
| N-6 | *(new)* H-8 has a fourth instance, in the code | strengthens R-3 toward "incomplete flag set" | | N-6 | *(new)* H-8 has a fourth instance, in the code | strengthens R-3 toward "incomplete flag set" |
| N-7 | G-8 "seventeen capability flags" | the real number is 28 of 32; G-8 grows | | N-7 | G-8 "seventeen capability flags" | the real number is 28 of 32; G-8 grows |
| N-8 | G-1 / this §6 "the §5 experiment is unrun" | **it had been run twice, returning GO twice, on unmerged branches** — both unsound; Track A's job was re-scoped from *run it* to *run it soundly* |
### N-1 · "No suite contains any `l10`/`always_on` test" is a scan artifact — the same artifact class caught twice before ### N-1 · "No suite contains any `l10`/`always_on` test" is a scan artifact — the same artifact class caught twice before
@ -94,6 +95,16 @@ G-10 says curriculum SERVE is "fully blocked by one engineering item," the 16-pr
Two flags' own documentation asserts a production configuration the production configuration does not have. This is H-8's failure mode located inside the source, one layer below the documents, and it moves R-3's evidence decisively: F-6 reads as an **incomplete flag set**, not intended dormancy. Dormancy remains a coherent ruling — but it would now require correcting two docstrings that say otherwise. Two flags' own documentation asserts a production configuration the production configuration does not have. This is H-8's failure mode located inside the source, one layer below the documents, and it moves R-3's evidence decisively: F-6 reads as an **incomplete flag set**, not intended dormancy. Dormancy remains a coherent ruling — but it would now require correcting two docstrings that say otherwise.
### N-8 · The §5 experiment was not unrun — it had returned GO twice, on branches nobody merged
*(Added 2026-07-28, while executing Track A. It is the eighth correction, and like the other seven it came from reading code — here, a branch tip's commit message and the script under it.)*
`rnd/sme-experiment-v2` @ `96e5f468` is titled **"Verdict: GO"** and carries a completed feasibility document dated 2026-07-19. `rnd/structure-mapping-experiment` @ `fc9d0c14` carries an earlier one. This plan described both branches as "scaffolding" and the item as "unrun," which is what G-1 said, which is what the ADR's §5 status implied. All three were wrong on the same point for nine days.
Neither verdict survives inspection — attempt 1 leaked the label into the embedding, and attempt 2 measured separability with `except ValueError: res = 1000.0`, a solver exception counted as a distance — so **the plan's conclusion was right and its premise was wrong**. Track A still needed doing; what it needed was not "run the experiment" but "run it under a binding criterion, and say what the existing verdicts are worth."
The mechanism is worth naming because it is H-9's, one level up: *work that finishes on an unmerged branch is invisible to every instrument that describes the project.* The register, the plan, and the ADR all read the same absence and inherited the same error. A branch tip is not a record.
### N-7 · The flag count is 28 of 32, not seventeen ### N-7 · The flag count is 28 of 32, not seventeen
`RuntimeConfig` is a single frozen dataclass with **32 boolean fields: 28 default `False`, 4 default `True`** (`allow_cross_language_recall`, `use_salience`, `discourse_planner`, `deduction_serving_enabled`). The assessment's "seventeen capability flags default off" understates the built-and-dark surface by eleven. `RuntimeConfig` is a single frozen dataclass with **32 boolean fields: 28 default `False`, 4 default `True`** (`allow_cross_language_recall`, `use_salience`, `discourse_planner`, `deduction_serving_enabled`). The assessment's "seventeen capability flags default off" understates the built-and-dark surface by eleven.
@ -221,12 +232,14 @@ G-2's fix lands only under its ratified ADR (R-11 may add an interim defensive g
Waves are hygiene and enforcement. The frontiers are the project. Track A runs **in parallel with Wave 0**, because it is execution-authorized already and everything else gets cheaper once it returns. Waves are hygiene and enforcement. The frontiers are the project. Track A runs **in parallel with Wave 0**, because it is execution-authorized already and everything else gets cheaper once it returns.
### Track A · **Run ADR-0252 §5 to a verdict** · G-1/H-10 · **L** · *leverage 1 in the whole assessment* ### Track A · **Run ADR-0252 §5 to a verdict** · G-1/H-10 · **DONE 2026-07-28 — NO-GO, awaiting ratification**
The scaffolding exists on two branches — `rnd/structure-mapping-experiment` @ `fc9d0c14` (feasibility doc + 214-line experiment script) and `rnd/sme-experiment-v2` @ `bed29a09` (corpus extractor with the `holdout_dev/v1` scope pin + single-pair probe). **Both worktrees are pruned from disk; both branches survive locally and on `forgejo`.** Nothing needs re-deriving. Protocol followed as written: criterion pre-registered at `dfc394d2` **before** the run (that commit carries the thresholds as importable constants and no results); corpus built with provenance on every case; full experiment run; `docs/research/sme-experiment-verdict-797ebad5.md` written with the criterion, the run, the numbers, and the verdict; report artifact committed with a pinned `deterministic_digest`.
Protocol: reconstruct one worktree from `rnd/sme-experiment-v2`; state the GO/NO-GO criterion **in writing before running** (the ADR's §5 conformance bar, not a criterion chosen after seeing results); extract the corpus; run the probe; run the full experiment; write `docs/research/sme-experiment-verdict-<sha>.md` with the criterion, the run, the numbers, and the verdict; bring it to Shay for ratification. **Verdict: NO-GO**, and it is the well-controlled kind the ADR pre-declares as full credit. Structure-sensitivity fails at every attribute weight tested — not a knife-edge on a constant. The mechanism is stated in one sentence and measured rather than argued: *the similarity quotient that would deliver attribute-invariance is the same quotient that annihilates structural contrast.* An `add`-vs-`subtract` minimal pair — one entity, identical numbers, one relation kind changed — aligns at residual exactly `0.0`, its two configurations being related by a proper rotation. Sweeping the attribute weight, every regime in which the SME property survives is a regime in which attributes contribute nothing (AUC 1.00 → 0.83 → 0.69 as they start to matter).
**A well-controlled NO-GO is full credit by the ADR's own terms** and is the cheapest possible outcome — it retires the §6 build authorization question permanently and redirects the widening program. This item has been mispriced as math-lane housekeeping (H-10); it governs the comprehension paradigm for everything. **Scope, deliberately narrow:** this refutes H1 for embeddings that encode role-structure as *point positions* aligned by conformal Procrustes *under similarity* — the argument is about the quotient, so it generalises across that class. It does not refute every Cl(4,1) representation, and it does not touch the symbolic structure-mapping lane already in `evals/structure_mapping/`.
**Two things the track found that were not on anyone's list.** (1) The experiment was **already run twice**, returning GO twice, on unmerged branches — see N-8; both verdicts are unsound and are voided by the verdict document. (2) The math reader decides **5 of 500** `holdout_dev/v1` cases (1.0%), all one skeleton, which is why §5.1's four-structure corpus was not extractable and is now registered as **G-21**.
### Track B · **The reading** · G-2 → G-3 · **XL** ### Track B · **The reading** · G-2 → G-3 · **XL**
Sequenced and gated: merge #138 (R-10) → fabrication ADR + ratification → the two known mutations land → **then** widen from 19, in whatever shape Track A's verdict dictates. G-16's latent defect class in `_inflect_predicate`'s aspect arms must be cleared *by* the widening program, not after it. This is the intelligence frontier and the largest capability gap to the telos. Sequenced and gated: merge #138 (R-10) → fabrication ADR + ratification → the two known mutations land → **then** widen from 19, in whatever shape Track A's verdict dictates. G-16's latent defect class in `_inflect_predicate`'s aspect arms must be cleared *by* the widening program, not after it. This is the intelligence frontier and the largest capability gap to the telos.

View file

@ -0,0 +1,210 @@
# ADR-0252 §5 — the structure-mapping experiment, run to a verdict
**Date:** 2026-07-28 · **Runner:** Opus 5 (Track A, `docs/assessment/50-execution-plan.md` §6)
**Base:** `forgejo/main` @ `797ebad5` · **Criterion:** pre-registered at `dfc394d2`, before the run
**Artifact:** `evals/structure_mapping/adr0252_s5/results/report-797ebad5.json`
**`deterministic_digest`:** `b3d9d27592e213104c51bcace7415fd819d722a1b06f7d0a5ca65bd5a234e4ec`
**Reproduce:** `uv run python -m evals.structure_mapping.adr0252_s5.experiment`
---
## VERDICT: **NO-GO**
Both embedding variants fail. Structure-sensitivity (§5.3c) fails at **every**
attribute weight tested — it is not a knife-edge on one constant. Per ADR-0252
§5.4, *"a well-controlled NO-GO is a full-credit result"*: it refutes geometric
SME for comprehension **in the embedding class tested** and redirects stage 2 to
a different representation.
For ratification. This document authorizes nothing on its own.
---
## 1. What was already there, and why it could not stand
The plan recorded §5 as *"authorized, scaffolded, and unrun."* It is scaffolded
and authorized. It is **not unrun** — it has been run twice, and both runs
returned GO on unmerged branches:
| Attempt | Branch @ SHA | Claimed | Why the claim does not hold |
|---|---|---|---|
| 1 | `rnd/structure-mapping-experiment` @ `fc9d0c14` | GO | Disowned by its own successor: the S1S4 label entered the embedding via hand-typed "canonical" matrices, so all cases of a structure were the same byte-array. Procrustes on identical matrices is trivially zero. |
| 2 | `rnd/sme-experiment-v2` @ `96e5f468` | GO | Genuinely blind — and then `except ValueError: res = 1000.0`. Its rule `cross_mean > same_mean*2 and same_mean < 1.0` is satisfied by that sentinel alone. **A solver exception is not a distance.** |
Attempt 2 carries four further defects, each of which independently voids it:
1. **Wrong branch of the instrument.** It passed mixed versor/point *sequences*
to `conformal_procrustes`, which dispatches those to field conjugacy
`W F W⁻¹` — the branch that raises `field conjugacy versor not closed`. Its
entire separability signal is that raise.
2. **Corpus provenance.** Its doc says "51 actual problem graphs from the
`holdout_dev/v1` corpus." **5 of 51** carry holdout ids; 46 carry bare
integers (`1`, `2`, `9`, … `2087`) matching nothing in that corpus, or
`gma-*` ids belonging to the **public** lane. §5.6's discipline is
*"`holdout_dev/v1` only"*, and §5.1 permits synthetic variants only *"marked
as such."* Nothing is marked.
3. **Duplicate data.** At least three exact-duplicate graphs under different ids
(`1`≡`51`, `2`≡`60`, `17`≡`46`) and three colliding ids (`gma-115`,
`gma-121`, `gma-135` each appear twice with *different* graphs, so its
`id → label` dict silently drops one of each pair). Duplicate graphs align at
residual 0 by construction — attempt 1's tautology re-imported through the
data instead of through the embedding.
4. **Unreproducible.** Its `scripts/extract_sme_corpus.py` is cited as an
artifact and was never committed. It also scored only the first 20 rows.
None of this was caught by re-reading the documents. It was caught by reading
the script, the corpus file, and the ids.
## 2. The coverage fact that shaped this run
**Measured at `797ebad5`: the serving reader returns a selected graph for 5 of
the 500 `holdout_dev/v1` cases — 1.0% — and all five carry the same relational
skeleton** (`compare_multiplicative`; they are exactly the organ cohort in
`evals/structure_mapping/scoring/labels.py`). On the public lane the same
pipeline yields 24/150.
So §5.1's corpus — "minimum four structures with clean labels … drawn from
`holdout_dev/v1`" — **is not extractable today**. That is almost certainly why
attempt 2 reached for other lanes; the mistake was doing it silently.
This run uses §5.1's own escape hatch, visibly: 5 real cases extracted through
`parse_and_solve`, 48 synthetic from four declared templates, every one carrying
`provenance="synthetic"` and pinned by `tests/test_adr_0252_s5_blindness.py`.
**Consequence, stated rather than buried:** this verdict speaks about **the
geometry**, not about coverage. §5.5's compression estimate ("how many holdout
families collapse to how few canonical structures") is **not answerable** at
1.0% reader coverage and was not attempted.
## 3. The instrument, characterised before use
`conformal_procrustes` has two branches and they are not interchangeable:
| Input | Branch | Behaviour |
|---|---|---|
| `(5,K)` grade-1 point cloud | Kabsch + Umeyama | never raises; invariant to rotation, translation **and uniform scale** |
| mixed versor/point sequence | field conjugacy | **raises** on unlike inputs |
Measured on dummy clouds before any corpus was embedded — identical `1.8e-16`,
rotated `7.8e-16`, translated `3.9e-15`, uniformly scaled ×2.5 `1.4e-14`,
non-similar z-stretch `5.9e-1`, unrelated `1.1e0`. This run uses the first
branch. **Non-convergence rate: 0.0%** across 1378 pairs in both variants — no
pair was excluded, and no exception contributed to any number below.
## 4. Results
53 cases, 1378 pairs, 334 same-structure / 1044 cross-structure.
| | RS-A (`ATTR_SCALE=0`) | RS-B (`ATTR_SCALE=0.02`) |
|---|---|---|
| **(a) Separability** | **PASS** — AUC **1.0000**, τ=0.0935, 100% / 100%, margin **+0.0867** | **FAIL** — AUC 0.8297, 75.1% / 75.7%, margin **0.4757** |
| **(b) Attribute-invariance** | PASS, but **analytic** — all 12 variants at exactly `0.000000` because attributes are not in the geometry | **FAIL** — median 0.0331 (ratio 0.107); the ×3 `rescale` pairs land at 0.0930.227 |
| **(c) Structure-sensitivity** | **FAIL**`add-vs-subtract` = **0.000000** | **FAIL**`S1-vs-S2` = 0.1764 < τ=0.2225 |
| (d) Systematicity *(reported, not gating)* | chain↔chain **0.0000**, chain↔flat **0.8703** | chain↔chain 0.0448, chain↔flat 0.8661 |
| **Verdict** | **NO-GO** | **NO-GO** |
### 4.1 Why RS-A's decisive control fails, exactly
The two graphs in the `add-vs-subtract` minimal pair differ in exactly one
field — one entity, identical numbers, `add` against `subtract`. Their
configurations are:
```
subtract : (1,0,0) (2,0,-1) (3,0,0)
add : (1,0,0) (2,0,+1) (3,0,0)
```
These are related by a **180° rotation about e1** — a *proper* rotation, det = +1.
Kabsch aligns them exactly, so the residual is not merely small, it is `0.0`.
**The similarity quotient that would give attribute-invariance is the same
quotient that annihilates the structural contrast.** A geometry blind to
rotation cannot distinguish "gain" from "loss" when their only encoded
difference is a signed offset on one axis.
This much *is* repairable by constants: with asymmetric kind heights
(`subtract → 7.0`) the three minimal pairs measure 0.162 / 1.940 / 0.217, all
above τ. **That repair is a diagnostic and does not change this verdict** — the
criterion was fixed in advance and a re-run under a repaired scheme requires its
own pre-registration. It is recorded here because concealing it would make the
NO-GO look deeper than it is.
### 4.2 What is *not* repairable by constants — the deeper result
Sweeping `ATTR_SCALE` under the same thresholds (committed as
`diagnostic_sweep` in the report):
| `ATTR_SCALE` | AUC | (a) | (b) | (c) | `rescale` median |
|---|---|---|---|---|---|
| 0.0 | 1.0000 | PASS | PASS | FAIL | 0.00000 |
| 0.001 | 1.0000 | PASS | PASS | FAIL | 0.00615 |
| 0.002 | 1.0000 | PASS | PASS | FAIL | 0.01245 |
| 0.005 | 1.0000 | PASS | PASS | FAIL | 0.03259 |
| 0.01 | 0.9742 | FAIL | PASS | FAIL | 0.06682 |
| 0.02 | 0.8297 | FAIL | FAIL | FAIL | 0.13505 |
| 0.05 | 0.6888 | FAIL | FAIL | FAIL | 0.40610 |
| 0.1 | 0.6694 | FAIL | FAIL | FAIL | 1.16296 |
Read down the table: **every regime in which the SME property survives is a
regime in which attributes contribute essentially nothing.** Where attributes
are geometrically negligible (≤0.005, `rescale` residual ≤0.033), separability
is perfect — but "invariance" there is a restatement of blindness. The moment
attributes carry real weight, separability collapses monotonically
(1.00 → 0.97 → 0.83 → 0.69) and invariance degrades with it.
There is no setting at which the geometry both *sees* surface attributes and
*factors them out*. That is the SME signature §5.3b was written to detect, and
it is absent — not narrowly missed.
The mechanism is structural rather than incidental: Procrustes measures a metric
distance between point configurations modulo similarity. Any attribute that
moves a point either moves it **within** the similarity orbit — in which case it
was never carrying information the aligner could use — or **outside** it, in
which case invariance fails by construction. There is no third option in this
class of embedding.
## 5. Scope of the refutation — stated narrowly on purpose
**What this refutes.** H1 for the class of embeddings that encode role-structure
as point positions in Cl(4,1) and align them with conformal Procrustes under
similarity. The §4.2 argument is about the similarity quotient itself, so it
generalises across position-encoding schemes in this class, not just the
particular constants used here.
**What this does not refute.** That *some* Cl(4,1) representation could carry
relational structure — e.g. one where relations are versors composed rather than
points positioned, or where alignment is a non-metric admissibility check rather
than a Procrustes residual. Those are different experiments and would need their
own pre-registrations.
**What it does not touch at all.** Whether the *symbolic* structure-mapping
already in `evals/structure_mapping/` works. That lane is untouched here.
## 6. What follows
Under ADR-0252 §5.4 a NO-GO "cleanly refutes geometric SME for comprehension and
redirects stage 2 to a different representation," and §5.4 pre-declares this
full credit. Three consequences for the registers, offered for ruling and not
taken:
1. **G-1 can close** with a recorded verdict, and the §6 build authorization
question is settled in the direction of "not on this representation."
2. **Track B's widening is unblocked in shape** (`50-execution-plan.md` §7's hard
gate: *"Track B widening ⟸ Track A verdict"*). The verdict says widening
proceeds by **more constructions**, not by geometric structure-mapping.
3. **H-10's mispricing is confirmed and discharged.** The experiment cost hours,
not an arc, and it governs the comprehension paradigm — exactly as H-10 said.
A fourth item is new and is the one worth your attention most: **the 1.0% reader
coverage of `holdout_dev/v1`** (§2) is a sharper measurement of the comprehension
frontier than G-3's construction count, on a corpus the project already treats as
its held-out standard. It belongs in the gap register on its own.
---
*Every number above is produced by `uv run python -m
evals.structure_mapping.adr0252_s5.experiment` at `797ebad5` and is reproduced by
the committed report artifact and its digest. The criterion constants in
`experiment.py` are unchanged since `dfc394d2`, which carried them before the run
and carried no results; the only post-run edit to that file is the additive
`diagnostic_sweep` block, and `git log -p` shows it.*

View file

@ -282,9 +282,9 @@ def sensitivity_pairs() -> list[tuple[str, Case, Case]]:
out.append( out.append(
( (
"add-vs-subtract", "add-vs-subtract",
case("sens-add", _s4(p[0], unit, 40.0, 3.0)), case("sens-subtract", _s4(p[0], unit, 40.0, 3.0)),
case( case(
"sens-sub", "sens-add",
graph_from_dict( graph_from_dict(
{ {
"entities": [p[0]], "entities": [p[0]],

View file

@ -16,7 +16,7 @@ import argparse
import hashlib import hashlib
import json import json
import statistics import statistics
from dataclasses import asdict, dataclass from dataclasses import dataclass
from pathlib import Path from pathlib import Path
from typing import Any, Final, Sequence from typing import Any, Final, Sequence
@ -189,6 +189,70 @@ def run_variant(name: str, attr_scale: float) -> dict[str, Any]:
} }
#: Diagnostic sweep. Added AFTER the run, and it touches no threshold above —
#: it exists to answer "is the verdict an artifact of ATTR_SCALE=0.02?" with a
#: curve instead of an opinion. The criterion constants are unchanged; `git log
#: -p` on this file shows the only post-run edit is this block.
SWEEP_SCALES: Final[tuple[float, ...]] = (0.0, 0.001, 0.002, 0.005, 0.01, 0.02, 0.05, 0.1)
def sweep() -> list[dict[str, Any]]:
"""Re-run (a), (b) and (c) across ATTR_SCALE, under the same thresholds."""
cases, labels = corpus_mod.build_corpus()
inv_pairs = corpus_mod.invariance_pairs()
sens = corpus_mod.sensitivity_pairs()
rows: list[dict[str, Any]] = []
for scale in SWEEP_SCALES:
clouds = {c.case_id: emb.embed(c.graph, attr_scale=scale) for c in cases}
ids = [c.case_id for c in cases]
same: list[float] = []
cross: list[float] = []
for i in range(len(ids)):
for j in range(i + 1, len(ids)):
r = residual(clouds[ids[i]], clouds[ids[j]])
if r is None:
continue
(same if labels[ids[i]] == labels[ids[j]] else cross).append(r)
auc = roc_auc(same, cross)
tau, below, above = best_threshold(same, cross)
inv_values = [
residual(emb.embed(left.graph, attr_scale=scale), emb.embed(right.graph, attr_scale=scale))
for _, left, right in inv_pairs
]
rescale_values = [
v for (kind, _, _), v in zip(inv_pairs, inv_values) if kind == "rescale" and v is not None
]
sens_values = {
kind: residual(emb.embed(left.graph, attr_scale=scale), emb.embed(right.graph, attr_scale=scale))
for kind, left, right in sens
}
clean_inv = [v for v in inv_values if v is not None]
cross_median = statistics.median(cross) if cross else float("nan")
rows.append(
{
"attr_scale": scale,
"auc": auc,
"threshold": tau,
"same_below_rate": below,
"cross_above_rate": above,
"invariance_median": statistics.median(clean_inv) if clean_inv else None,
"rescale_median": statistics.median(rescale_values) if rescale_values else None,
"cross_median": cross_median,
"minimal_pairs": sens_values,
"separability_pass": auc >= AUC_FLOOR
and below >= CLASSIFY_FLOOR
and above >= CLASSIFY_FLOOR,
"invariance_pass": len(clean_inv) == len(inv_values)
and all(v < tau for v in clean_inv)
and statistics.median(clean_inv) <= INVARIANCE_RATIO_CEILING * cross_median,
"sensitivity_pass": all(
v is not None and v > tau for v in sens_values.values()
),
}
)
return rows
def main(argv: Sequence[str] | None = None) -> int: def main(argv: Sequence[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__) parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--out", type=Path, default=None, help="write the JSON report here") parser.add_argument("--out", type=Path, default=None, help="write the JSON report here")
@ -204,6 +268,7 @@ def main(argv: Sequence[str] | None = None) -> int:
"verdict_rule": "GO iff separability AND attribute-invariance AND structure-sensitivity on the same variant", "verdict_rule": "GO iff separability AND attribute-invariance AND structure-sensitivity on the same variant",
}, },
"variants": [run_variant(name, scale) for name, scale in VARIANTS.items()], "variants": [run_variant(name, scale) for name, scale in VARIANTS.items()],
"diagnostic_sweep": sweep(),
} }
verdicts = {v["variant"]: v["verdict"] for v in report["variants"]} verdicts = {v["variant"]: v["verdict"] for v in report["variants"]}
report["verdict_by_variant"] = verdicts report["verdict_by_variant"] = verdicts

View file

@ -0,0 +1,440 @@
{
"criterion": {
"auc_floor": 0.9,
"classify_floor": 0.95,
"invariance_ratio_ceiling": 0.25,
"nonconvergence_ceiling": 0.05,
"verdict_rule": "GO iff separability AND attribute-invariance AND structure-sensitivity on the same variant"
},
"deterministic_digest": "b3d9d27592e213104c51bcace7415fd819d722a1b06f7d0a5ca65bd5a234e4ec",
"diagnostic_sweep": [
{
"attr_scale": 0.0,
"auc": 1.0,
"cross_above_rate": 1.0,
"cross_median": 0.19863087538407323,
"invariance_median": 4.18427621245356e-16,
"invariance_pass": true,
"minimal_pairs": {
"S1-vs-S2": 0.16161396145136686,
"S3-vs-S4": 0.19051290886337985,
"add-vs-subtract": 0.0
},
"rescale_median": 4.18427621245356e-16,
"same_below_rate": 1.0,
"sensitivity_pass": false,
"separability_pass": true,
"threshold": 0.09348293495747853
},
{
"attr_scale": 0.001,
"auc": 1.0,
"cross_above_rate": 1.0,
"cross_median": 0.19934413238921983,
"invariance_median": 0.0013282930860297809,
"invariance_pass": true,
"minimal_pairs": {
"S1-vs-S2": 0.1621836277726819,
"S3-vs-S4": 0.19730918637870587,
"add-vs-subtract": 0.007397449569965667
},
"rescale_median": 0.0061465391631068175,
"same_below_rate": 1.0,
"sensitivity_pass": false,
"separability_pass": true,
"threshold": 0.09288238132409174
},
{
"attr_scale": 0.002,
"auc": 1.0,
"cross_above_rate": 1.0,
"cross_median": 0.2068596013996259,
"invariance_median": 0.0026704005687431696,
"invariance_pass": true,
"minimal_pairs": {
"S1-vs-S2": 0.16276373558692778,
"S3-vs-S4": 0.20447303891189786,
"add-vs-subtract": 0.01477536901427399
},
"rescale_median": 0.012452138303227247,
"same_below_rate": 1.0,
"sensitivity_pass": false,
"separability_pass": true,
"threshold": 0.08991312513645279
},
{
"attr_scale": 0.005,
"auc": 1.0,
"cross_above_rate": 1.0,
"cross_median": 0.23288467474611535,
"invariance_median": 0.006826167438046735,
"invariance_pass": true,
"minimal_pairs": {
"S1-vs-S2": 0.16457580592410367,
"S3-vs-S4": 0.22799918174262307,
"add-vs-subtract": 0.03676078612038948
},
"rescale_median": 0.03258807626933125,
"same_below_rate": 1.0,
"sensitivity_pass": false,
"separability_pass": true,
"threshold": 0.10413040131434703
},
{
"attr_scale": 0.01,
"auc": 0.9742182302062542,
"cross_above_rate": 0.9042145593869731,
"cross_median": 0.26332954754005544,
"invariance_median": 0.014417905622498534,
"invariance_pass": true,
"minimal_pairs": {
"S1-vs-S2": 0.16790252476327816,
"S3-vs-S4": 0.27275294777749237,
"add-vs-subtract": 0.07278692111710225
},
"rescale_median": 0.06681904316690977,
"same_below_rate": 0.907185628742515,
"sensitivity_pass": false,
"separability_pass": false,
"threshold": 0.1533011869079436
},
{
"attr_scale": 0.02,
"auc": 0.8297399453965631,
"cross_above_rate": 0.7567049808429118,
"cross_median": 0.31040363592935377,
"invariance_median": 0.033068273551452323,
"invariance_pass": false,
"minimal_pairs": {
"S1-vs-S2": 0.17636627398017427,
"S3-vs-S4": 0.37012827734659925,
"add-vs-subtract": 0.1417226819224871
},
"rescale_median": 0.13504807553913323,
"same_below_rate": 0.7514970059880239,
"sensitivity_pass": false,
"separability_pass": false,
"threshold": 0.2224772412728761
},
{
"attr_scale": 0.05,
"auc": 0.6888206345928832,
"cross_above_rate": 0.6197318007662835,
"cross_median": 0.5444760496385956,
"invariance_median": 0.09178528893308188,
"invariance_pass": false,
"minimal_pairs": {
"S1-vs-S2": 0.22961316406869808,
"S3-vs-S4": 0.28848882111140384,
"add-vs-subtract": 0.2788978285534268
},
"rescale_median": 0.4061020972791902,
"same_below_rate": 0.6197604790419161,
"sensitivity_pass": false,
"separability_pass": false,
"threshold": 0.40330909855248576
},
{
"attr_scale": 0.1,
"auc": 0.6694427237479065,
"cross_above_rate": 0.6111111111111112,
"cross_median": 1.0001770231984946,
"invariance_median": 0.2075140383804508,
"invariance_pass": false,
"minimal_pairs": {
"S1-vs-S2": 0.3914077252277385,
"S3-vs-S4": 0.3016565786534484,
"add-vs-subtract": 0.13928060953384408
},
"rescale_median": 1.1629642847658526,
"same_below_rate": 0.6107784431137725,
"sensitivity_pass": false,
"separability_pass": false,
"threshold": 0.7387414312757857
}
],
"experiment": "ADR-0252 \u00a75 structure-mapping acceptance gate",
"overall_verdict": "NO-GO",
"variants": [
{
"attr_scale": 0.0,
"attribute_invariance": {
"median": 4.18427621245356e-16,
"pass": true,
"ratio_to_cross_median": 2.106558813861557e-15,
"rows": [
{
"left": "inv-S1-base",
"perturbation": "rename",
"residual": 5.709490266768425e-16,
"right": "inv-S1-rename"
},
{
"left": "inv-S1-base",
"perturbation": "rescale",
"residual": 5.709490266768425e-16,
"right": "inv-S1-rescale"
},
{
"left": "inv-S1-base",
"perturbation": "jitter",
"residual": 5.709490266768425e-16,
"right": "inv-S1-jitter"
},
{
"left": "inv-S2-base",
"perturbation": "rename",
"residual": 2.659062158138695e-16,
"right": "inv-S2-rename"
},
{
"left": "inv-S2-base",
"perturbation": "rescale",
"residual": 2.659062158138695e-16,
"right": "inv-S2-rescale"
},
{
"left": "inv-S2-base",
"perturbation": "jitter",
"residual": 2.659062158138695e-16,
"right": "inv-S2-jitter"
},
{
"left": "inv-S3-base",
"perturbation": "rename",
"residual": 8.5663214756370435e-16,
"right": "inv-S3-rename"
},
{
"left": "inv-S3-base",
"perturbation": "rescale",
"residual": 8.5663214756370435e-16,
"right": "inv-S3-rescale"
},
{
"left": "inv-S3-base",
"perturbation": "jitter",
"residual": 8.5663214756370435e-16,
"right": "inv-S3-jitter"
},
{
"left": "inv-S4-base",
"perturbation": "rename",
"residual": 0.0,
"right": "inv-S4-rename"
},
{
"left": "inv-S4-base",
"perturbation": "rescale",
"residual": 0.0,
"right": "inv-S4-rescale"
},
{
"left": "inv-S4-base",
"perturbation": "jitter",
"residual": 0.0,
"right": "inv-S4-jitter"
}
]
},
"n_cases": 53,
"n_pairs": 1378,
"nonconvergence_rate": 0.0,
"separability": {
"auc": 1.0,
"cross_above_rate": 1.0,
"cross_median": 0.19863087538407323,
"margin_min_cross_minus_max_same": 0.08672913410526156,
"n_cross_pairs": 1044,
"n_same_pairs": 334,
"pass": true,
"same_below_rate": 1.0,
"same_median": 5.709490266768425e-16,
"threshold": 0.09348293495747853
},
"structure_sensitivity": {
"pass": false,
"rows": [
{
"left": "sens-S1",
"minimal_pair": "S1-vs-S2",
"residual": 0.16161396145136686,
"right": "sens-S2"
},
{
"left": "sens-S3",
"minimal_pair": "S3-vs-S4",
"residual": 0.19051290886337985,
"right": "sens-S4b"
},
{
"left": "sens-subtract",
"minimal_pair": "add-vs-subtract",
"residual": 0.0,
"right": "sens-add"
}
]
},
"systematicity": {
"rows": [
{
"comparison": "chain-vs-chain",
"left": "syst-chain-a",
"residual": 9.519840216985692e-16,
"right": "syst-chain-b"
},
{
"comparison": "chain-vs-flat",
"left": "syst-chain-a",
"residual": 0.870344201932933,
"right": "syst-flat"
}
]
},
"variant": "RS-A",
"verdict": "NO-GO"
},
{
"attr_scale": 0.02,
"attribute_invariance": {
"median": 0.033068273551452323,
"pass": false,
"ratio_to_cross_median": 0.10653313854538253,
"rows": [
{
"left": "inv-S1-base",
"perturbation": "rename",
"residual": 2.339999357265776e-16,
"right": "inv-S1-rename"
},
{
"left": "inv-S1-base",
"perturbation": "rescale",
"residual": 0.09326453180940468,
"right": "inv-S1-rescale"
},
{
"left": "inv-S1-base",
"perturbation": "jitter",
"residual": 0.01564292374731886,
"right": "inv-S1-jitter"
},
{
"left": "inv-S2-base",
"perturbation": "rename",
"residual": 2.7776227440016997e-16,
"right": "inv-S2-rename"
},
{
"left": "inv-S2-base",
"perturbation": "rescale",
"residual": 0.22728592837727327,
"right": "inv-S2-rescale"
},
{
"left": "inv-S2-base",
"perturbation": "jitter",
"residual": 0.03068103638906839,
"right": "inv-S2-jitter"
},
{
"left": "inv-S3-base",
"perturbation": "rename",
"residual": 1.8389178249470913e-16,
"right": "inv-S3-rename"
},
{
"left": "inv-S3-base",
"perturbation": "rescale",
"residual": 0.12118122411314303,
"right": "inv-S3-rescale"
},
{
"left": "inv-S3-base",
"perturbation": "jitter",
"residual": 0.03545551071383625,
"right": "inv-S3-jitter"
},
{
"left": "inv-S4-base",
"perturbation": "rename",
"residual": 6.896644022921873e-16,
"right": "inv-S4-rename"
},
{
"left": "inv-S4-base",
"perturbation": "rescale",
"residual": 0.14891492696512346,
"right": "inv-S4-rescale"
},
{
"left": "inv-S4-base",
"perturbation": "jitter",
"residual": 0.05081216199265817,
"right": "inv-S4-jitter"
}
]
},
"n_cases": 53,
"n_pairs": 1378,
"nonconvergence_rate": 0.0,
"separability": {
"auc": 0.8297399453965631,
"cross_above_rate": 0.7567049808429118,
"cross_median": 0.31040363592935377,
"margin_min_cross_minus_max_same": -0.47567711686657266,
"n_cross_pairs": 1044,
"n_same_pairs": 334,
"pass": false,
"same_below_rate": 0.7514970059880239,
"same_median": 0.11177849968220438,
"threshold": 0.2224772412728761
},
"structure_sensitivity": {
"pass": false,
"rows": [
{
"left": "sens-S1",
"minimal_pair": "S1-vs-S2",
"residual": 0.17636627398017427,
"right": "sens-S2"
},
{
"left": "sens-S3",
"minimal_pair": "S3-vs-S4",
"residual": 0.37012827734659925,
"right": "sens-S4b"
},
{
"left": "sens-subtract",
"minimal_pair": "add-vs-subtract",
"residual": 0.1417226819224871,
"right": "sens-add"
}
]
},
"systematicity": {
"rows": [
{
"comparison": "chain-vs-chain",
"left": "syst-chain-a",
"residual": 0.04475134339414685,
"right": "syst-chain-b"
},
{
"comparison": "chain-vs-flat",
"left": "syst-chain-a",
"residual": 0.8661270451535805,
"right": "syst-flat"
}
]
},
"variant": "RS-B",
"verdict": "NO-GO"
}
],
"verdict_by_variant": {
"RS-A": "NO-GO",
"RS-B": "NO-GO"
}
}

View file

@ -0,0 +1,118 @@
"""ADR-0252 §5 — the blindness invariant, pinned.
The §5 experiment has already been run twice and returned GO twice on rules
that were not binding. Attempt 1's specific defect was **label leakage**: the
structural label reached the embedding coordinates, so Procrustes aligned
identical byte-arrays and the test was a tautology.
These pins make that defect a failing test rather than a thing a reader has to
notice. They are cheap, they are in the smoke suite, and they would each have
caught a real historical error.
"""
from __future__ import annotations
import inspect
import numpy as np
import pytest
from evals.structure_mapping.adr0252_s5 import corpus, embedding
from evals.structure_mapping.adr0252_s5.embedding import EmbeddingRefusal
def test_embedding_module_never_reads_labels() -> None:
"""The mapper may not import, name, or receive a structure label.
Attempt 1 failed exactly here. A grep-level pin is enough: the embedding is
a pure function of the graph, so any mention of the label vocabulary in its
source is either leakage or a comment that will become leakage.
"""
source = inspect.getsource(embedding)
body = "\n".join(
line for line in source.splitlines() if not line.strip().startswith("#")
)
# The docstring legitimately discusses the scheme; strip it before scanning.
body = body.replace(embedding.__doc__ or "", "")
for label in corpus.STRUCTURES:
assert f'"{label}"' not in body and f"'{label}'" not in body, (
f"embedding.py references the structure label {label!r} — the "
"mapper must be blind to labels (ADR-0252 §5.2)"
)
assert "corpus" not in body.replace("corpus.py", ""), (
"embedding.py must not import the corpus module: labels live there"
)
def test_embedding_is_a_pure_function_of_the_graph() -> None:
"""Same graph twice → byte-identical cloud. No hidden state, no RNG."""
cases, _ = corpus.build_corpus()
graph = cases[0].graph
first = embedding.embed(graph, attr_scale=0.0)
second = embedding.embed(graph, attr_scale=0.0)
assert np.array_equal(first, second)
def test_entity_names_never_reach_the_geometry() -> None:
"""Renaming every entity must not move a single coordinate.
This is what makes §5.3b's ``rename`` perturbation analytic rather than
measured, and the verdict document says so; the claim is pinned here so it
stays true if the scheme changes.
"""
for kind, left, right in corpus.invariance_pairs():
if kind != "rename":
continue
for scale in (0.0, 0.02):
assert np.array_equal(
embedding.embed(left.graph, attr_scale=scale),
embedding.embed(right.graph, attr_scale=scale),
), f"rename changed the geometry for {left.case_id} at attr_scale={scale}"
def test_undeclared_relation_kind_refuses_rather_than_guessing() -> None:
"""No default position for an unknown relation (fail closed, INV-34 shape)."""
cases, _ = corpus.build_corpus()
graph = cases[0].graph
saved = dict(embedding.KIND_HEIGHT)
try:
embedding.KIND_HEIGHT = {} # type: ignore[assignment]
with pytest.raises(EmbeddingRefusal):
embedding.embed(graph, attr_scale=0.0)
finally:
embedding.KIND_HEIGHT = saved # type: ignore[assignment]
def test_corpus_marks_every_synthetic_case() -> None:
"""§5.1 permits synthetic variants only 'marked as such'.
Attempt 2's corpus carried 46 of 51 cases from outside ``holdout_dev/v1``
with nothing marking them. This pin makes that unrepeatable.
"""
cases, labels = corpus.build_corpus()
assert cases, "corpus must not be empty"
for case in cases:
assert case.provenance in {"holdout_dev/v1", "synthetic"}
assert case.case_id in labels
if case.provenance == "holdout_dev/v1":
assert case.case_id.startswith("gsm8k-holdout-dev-v1-"), (
f"{case.case_id} claims holdout provenance without a holdout id"
)
def test_corpus_has_no_duplicate_graphs() -> None:
"""Identical inputs align at residual 0 by construction.
Attempt 2's corpus contained at least three exact-duplicate graphs under
different ids (and three colliding ids), which re-imported attempt 1's
tautology through the data instead of through the embedding.
"""
cases, _ = corpus.build_corpus()
seen: dict[str, str] = {}
for case in cases:
key = repr(case.graph.as_json())
assert key not in seen, (
f"{case.case_id} duplicates {seen[key]} — duplicate graphs make "
"same-structure alignment trivially zero"
)
seen[key] = case.case_id