core/evals/logos/repaired_ground.py
Shay cbfc8ccbf7 feat(evals,tests,docs,adr): FA-1 ruled — cross-language holonomy does not discriminate meaning (NO-GO)
The CORE-Logos "Crown Proof" — ADR-0015's claim that cross-language holonomy
closure is THE validation gate of meaning — was measured against a criterion
registered before the run, and refused.

  G1  separation, all negatives      AUC 0.557   (chance 0.500, bar 0.80)   FAIL
  G2  hardest class, cross-pairing   AUC 0.664   (bar 0.75)                 FAIL
  G3  word-order sensitivity         0.644       (bar 0.90)                 FAIL
  G4  no collapse re-entry           0 lost      (bar: 0)                   PASS

G4 is the half that succeeded: the repairs work. The claim is what failed --
which is the outcome the registration called "worth as much as a GO", because
it retires the largest piece of unearned design in the system.

Both preconditions had to be BUILT before the question could be asked at all,
which is why the 2026-06-14 negative constrained nothing: the shipped compiler
collapsed 37-53 coordinates, and holonomy_encode never closed. Measured on the
repaired ground with a genuine loop, over 1,016 aligned pairs and 58,375
mechanically-generated negatives in three classes, nothing hand-picked.

Diagnosis, not just a number: the encoding reacts more strongly to REORDERING a
clause (median deviation 41.2) than to CHANGING WHAT IT IS ABOUT (32.2). It
measures path shape, not content -- and that follows from the algebra, since
permutation acts on an ordered product through non-commutativity, a first-order
effect, while substituting one of three factors perturbs it only through that
factor. The negative classes do order correctly by meaning-distance
(aligned 27.2 < lexical 32.2 < order 41.2 < cross-pair 53.2), so the geometry is
weak signal rather than noise -- a correlation where the design asked for a gate.

Consequences, all landed here:

  * ADR-0005 and ADR-0015 AMENDED to record the refusal. The three-language
    architecture, the pack contract, morphology-as-structure and Hebrew root
    folding into geometry all survive untouched; only the claim that alignment
    or closure VALIDATES meaning is retired.
  * tests/test_fa1_gate_verdict.py is the tripwire: if a future encoding makes
    the gate real, it fails, and the amendment is withdrawn in the same commit
    that replaces it with a proof. Two sabotages observed red.
  * allow_cross_language_recall RULED by measurement (work-order item 3): it
    gates vault-recall depth over field states, has no language argument and no
    cross-language path, and reaches only walk_surface, which runtime_contracts
    declares telemetry. Stays ON, reclassified DEPLOYMENT, flip condition
    recorded. G-25's own claim that it was "the pillar's own switch" is
    CORRECTED -- that read the flag's name instead of its call graph, which is
    the error the audit exists to find.
  * A sixth defect found while building the harness: en_collapse_anchors_v1
    declares role "collapse_anchor", which is not a LanguageRole member, so
    load_pack() raises. It is registered in the resolver, consumed by opening
    lexicon.jsonl on a raw path that bypasses the loader, and is the target of
    all 24 alignment edges that still resolve to nothing. The enum's own comment
    records the identical failure once before (ADR-0097, domain_seed).

The repairs cross to the keel on their own merit regardless of the verdict:
geodesic blending removes every coordinate collision, and mount-wide edge
resolution connects 39 of the 63 dead edges. The collapse was damaging English's
own vocabulary; seeing that required no theory of meaning.

FA-1 is CLOSED: all four work-order items discharged, L2 verdict recorded as
defective-but-repairable with its central design claim retired.

[Verification]: uv run core test --suite smoke -q -> 786 passed in 239.45s, EXIT=0
                (+4 tests, +15s vs the prior 782/224s -- the FA-1 pin's stated cost).
                Doc-parsing pins re-run after the ADR and register amendments:
                test_adr_status_governance + test_adr_index -> 321 passed;
                test_flag_register -> 6 passed.
2026-07-28 14:37:07 -07:00

154 lines
7.4 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""The repaired semantic ground — R1/R2/R3 of the FA-1 pre-registration.
`docs/analysis/fa1-holonomy-gate-preregistration.md` registers three repairs and the
criterion that judges them. This module implements the repairs; it takes no verdict.
Nothing here modifies the serving compiler. R1 and R3 are applied by *substituting two
module-level functions* inside a context manager, against the real compile path, so the
experiment measures the shipped pipeline with two operations corrected rather than a
re-implementation that could drift from it. The caches are cleared on entry and on exit,
so a repaired ground never leaks into anything else in the process.
R1 blend on the group instead of overwriting — `algebra/rotor.py`'s geodesic
R2 close the loop around the *pair* — `loop_deviation` below
R3 resolve alignment edges against the packs actually mounted, not an id-prefix table
"""
from __future__ import annotations
from contextlib import contextmanager
from typing import Iterator
import numpy as np
from algebra.cl41 import geometric_product, reverse as cl_reverse
from algebra.rotor import rotor_power, word_transition_rotor
from algebra.versor import unitize_versor
from packs import compiler as _compiler
#: The six packs that carry logos data. The serving path grounds Hebrew and Greek against
#: the `*_cognition_v1` pair (`chat/pack_grounding.py:56-57`); the micro packs are the
#: ADR-0015 seed. R3 exists so both participate — 63 of the 83 authored edges live in the
#: cognition packs and currently resolve to nothing.
LOGOS_MOUNT: tuple[str, ...] = (
"en_minimal_v1",
"en_core_cognition_v1",
"he_logos_micro_v1",
"he_core_cognition_v1",
"grc_logos_micro_v1",
"grc_logos_cognition_v1",
)
#: `en_collapse_anchors_v1` is DELIBERATELY ABSENT and cannot be added: its manifest declares
#: `role: "collapse_anchor"`, which is not a member of `packs.schema.LanguageRole`, so
#: `load_pack()` raises before the lexicon is read. `chat/pack_grounding.py` consumes it by
#: opening `lexicon.jsonl` on a raw path — bypassing the loader that would reject it — and 24
#: alignment edges point into it, all of which therefore resolve to nothing. The enum carries
#: a comment recording this exact failure once before (ADR-0097, `domain_seed`); a registry
#: widened by hand falls behind by hand. Reported by the gate as `edges_unresolved_under_R3`
#: rather than silently excluded.
#: Identity of Cl(4,1): the scalar 1. A closed loop returns here.
IDENTITY = np.zeros(32, dtype=np.float64)
IDENTITY[0] = 1.0
# ---------------------------------------------------------------------------
# R1 — interpolation that stays on the versor group
# ---------------------------------------------------------------------------
def geodesic_blend(source: np.ndarray, target: np.ndarray, strength: float) -> np.ndarray:
"""Move `source` a fraction `strength` along the group geodesic toward `target`.
The operation the compiler's `_blend_feature_versors` was named and parametrised for.
A linear combination of two versors is not a versor — the group is not closed under
addition — which is why `VocabManifold.update()` refuses one, and why returning the
target verbatim became the escape. `R^t · source`, with `R` the transition rotor, is
lawful at every `t`: it stays on the group by construction.
Endpoints are exact: `t=0` returns the source, `t=1` the target.
"""
strength = max(0.0, min(1.0, float(strength)))
src = np.asarray(source, dtype=np.float64)
if strength <= 0.0:
return np.asarray(source, dtype=np.float32).copy()
tgt = np.asarray(target, dtype=np.float64)
if strength >= 1.0:
return np.asarray(target, dtype=np.float32).copy()
moved = geometric_product(rotor_power(word_transition_rotor(src, tgt), strength), src)
return np.asarray(unitize_versor(moved), dtype=np.float32)
def _clear_pack_caches() -> None:
_compiler._load_pack_cached.cache_clear()
_compiler._load_mounted_packs_cached.cache_clear()
_compiler._load_pack_entries_cached.cache_clear()
@contextmanager
def repaired_compiler(pack_ids: tuple[str, ...] = LOGOS_MOUNT) -> Iterator[None]:
"""Compile with R1 (geodesic blending) and R3 (mount-wide edge resolution) in force.
R3 replaces the hardcoded `{"he": …, "grc": …, "en": …}` prefix table with the packs
actually being mounted. Entry ids are pack-local, so a global prefix table can only
ever address the three packs it names; every edge belonging to any other pack of the
same language resolves to a pack that has never heard of it and is dropped in silence.
"""
original_blend = _compiler._blend_feature_versors
original_infer = _compiler._infer_foreign_pack_ids
def _infer_from_mount(home_pack_id: str, _graph) -> list[str]: # noqa: ANN001 - private type
# The graph is ignored on purpose: membership of the mount, not the shape of the
# edge ids, decides which packs an edge may address.
return [pack_id for pack_id in pack_ids if pack_id != home_pack_id]
_compiler._blend_feature_versors = geodesic_blend
_compiler._infer_foreign_pack_ids = _infer_from_mount
_clear_pack_caches()
try:
yield
finally:
_compiler._blend_feature_versors = original_blend
_compiler._infer_foreign_pack_ids = original_infer
_clear_pack_caches()
# ---------------------------------------------------------------------------
# R2 — the loop, closed around a pair of representations
# ---------------------------------------------------------------------------
def forward_walk(versors: list[np.ndarray]) -> np.ndarray:
"""`F(X) = v₁ · v₂ · … · vₙ` — the ordered geometric product of a clause's versors.
No per-token position rotor. The shipped encoder conjugates each token by its absolute
index, which was introduced so path order would survive degenerate scalar/vector test
fixtures; across languages, clauses differ in length, so that schedule injects a length
artefact into every cross-language comparison (RQ2 of the 2026-06-14 finding). Order
sensitivity does not depend on it: the geometric product is non-commutative.
"""
if not versors:
raise ValueError("cannot walk an empty clause")
walk = unitize_versor(np.asarray(versors[0], dtype=np.float64))
for versor in versors[1:]:
walk = geometric_product(walk, unitize_versor(np.asarray(versor, dtype=np.float64)))
walk = unitize_versor(walk)
return walk
def closed_holonomy(clause_a: list[np.ndarray], clause_b: list[np.ndarray]) -> np.ndarray:
"""`H(A,B) = unitize(F(A) · reverse(F(B)))` — transport out along A, return along B."""
return unitize_versor(geometric_product(forward_walk(clause_a), cl_reverse(forward_walk(clause_b))))
def loop_deviation(clause_a: list[np.ndarray], clause_b: list[np.ndarray]) -> float:
"""`d(A,B) = ‖H(A,B) 1‖` — the single declared metric of the FA-1 criterion.
Zero when the loop closes. The null hypothesis is the identity, so there is no
similarity threshold to calibrate and no metric to select after seeing the numbers —
the failure mode of the June measurement, forbidden here by construction.
Sign note: `H` and `H` describe the same transport (versors are defined up to sign),
so the deviation is taken to the nearer of `±1`.
"""
holonomy = closed_holonomy(clause_a, clause_b)
return float(min(np.linalg.norm(holonomy - IDENTITY), np.linalg.norm(holonomy + IDENTITY)))