Records the architectural floor for frontier-LLM performance on each
Phase 2 v1 lane.
The baseline is structural: every lane's scoring rubric measures a
property that frontier LLMs do not architecturally emit (Provenance
typed sources, pack_mutation_proposal, vault_hits, REJECTED_IDENTITY
outcome, deterministic trace_hash). The frontier score on each of
those sub-metrics is 0.0 by construction, not by failure — even a
live-API run would still record 0.0 on these typed-signal checks
because the evidence is absent regardless of prose quality.
Artifacts:
docs/frontier_baselines.md
Full per-lane analysis: what each sub-metric scores, why the
frontier value is 0, and where a live-API baseline would or
would not add information.
evals/<lane>/baselines/v1_structural_zero.json (× 5)
Per-lane baseline records in the same shape as lane reports.
Encodes 0.0 / None on each sub-metric with rationale.
evals/baseline_runner.py
Adds StructuralZeroBaseline adapter conforming to the
BaselineModel protocol — a real, non-stub adapter that returns
the deterministic floor. Live-API adapters (Anthropic, OpenAI)
can be wired alongside when API keys are configured; the
structural floor remains the comparison baseline.
Across 5 lanes / 14 typed-signal sub-metrics:
CORE v1: 1.0 (each)
frontier structural: 0.0 (each)
The gap is "CORE measures a property frontier output does not
expose", not "CORE outperforms on a shared benchmark". v2 lanes may
add content-level sub-metrics where direct comparison via live-API
runs becomes meaningful.
Adds the third Phase 2 lane: calibration measures whether CORE's runtime
emits distinguishable, typed evidence for three cognitive states:
no_grounding vault_hits == 0 (gate fired, no recall)
coherent vault_hits > 0 (vault recall fired)
correction_proposed pack_mutation_proposal is not None
Each case runs on its own fresh CognitiveTurnPipeline to avoid
cross-case field-state drift (the gate's geometric recall score is
sensitive to vault content drift across turns).
v1 results: dev 12/12, public/v1 24/24, holdouts/v1 18/18 — all classes
score 1.0 across all splits.
Architectural findings logged in evals/calibration/gaps.md:
1. The ingest gate fires on a *geometric* CGA-recall score, not on
semantic OOD. 6/42 hand-chosen OOD prompts fire the gate with a
warmed vault; the other 36 land geometrically near in-pack
versors after morphological grounding. v1 measures the reliable
recall/correction signals, not semantic OOD detection.
2. CognitiveTurnPipeline.run() unconditionally overrides the
runtime's gate-safety surface with the realizer surface. The OOD
marker survives in walk_surface but not in surface. v1 classifies
on vault_hits (preserved) rather than surface (overridden).
Both findings are filed as suggested follow-up work, not v1 blockers.