* fix(cognition): add explicit surface resolution policy
* test(cognition): cover explicit surface resolution policy
* fix(cognition): route pipeline surfaces through resolver
* fix(cognition): address PR #76 review comments
- hoist `_is_useful_surface` import from inside `run()` to module top
- call `_render_walk_surface` / `_render_compose_surface` via the class
name (both are @staticmethod) for consistency with the existing
`_fold_*_into_surface` helpers
- drop redundant `realized_surface` truthiness check in
`resolve_surface` — `realizer_useful` already excludes empty /
placeholder surfaces via `_is_useful_surface`
Tests: tests/test_surface_resolution.py + tests/test_cognitive_turn_pipeline.py
green (16 passed); cognition suite 120/1s, smoke suite 67/0.
The original "Why does light exist?" complaint that motivated ADR-0084
was specifically about CAUSE-intent surfaces. ADR-0084 (substrate) +
PR #65 (content) already moved DEFINITION/RECALL to gloss-grounded
surfaces ("Light is visible medium that reveal truth."). But CAUSE
still dispatched through the chain-walk path:
Before: light — teaching-grounded (cognition_chains_v1):
cognition.illumination; logos.core.
light reveals truth (cognition.truth).
No session evidence yet.
After: Light exists as visible medium that reveal truth.
pack-grounded (en_core_cognition_v1).
The chain-walk is structurally correct but the wrong SHAPE for a why-
question — it's a graph traversal, not an explanation. ADR-0085 fixes
the shape using the same gloss material that DEFINITION/RECALL already
consume, with no new content authoring.
Additive composer
chat/pack_grounding.py:gloss_aware_cause_surface()
- Resolves gloss via lexicon-residency-checked resolve_gloss().
- Frames POS-aware:
NOUN -> "{Lemma} exists as {gloss}."
VERB -> "To {lemma} is to {gloss}."
ADJ -> "To be {lemma} is to {gloss}."
* -> falls back to _frame_gloss (predicate-identity).
- Threads anchor lens via the existing helper (ADR-0073c parity).
- Returns None when no gloss exists — runtime falls through to the
existing chain-walk path. Additive: no CAUSE case loses its surface.
Runtime dispatch
chat/runtime.py — IntentTag.CAUSE tries gloss path FIRST under the
flag; falls through to teaching_grounded_surface* on None.
Unconditional fallback — never silent.
Opt-in flag
core/config.py — RuntimeConfig.gloss_aware_cause: bool = False
Default off preserves pre-ADR-0085 chain-walk surfaces byte-
identically (null-drop invariant, CI-pinned).
Prompt-diversity classifier update
evals/prompt_diversity/runner.py — _CAUSE_MARKERS widened with the
explanation-frame markers ("exists as", "is to", "to be", "is for",
"purpose of") plus bare-form predicates ("reveal" alongside
"reveals"). Neither composer path is penalised on shape_fit just on
inflection grounds.
v1/public lift (flag OFF vs ON, 26 cases)
intent_accuracy : 65.4% -> 65.4% ( — )
versor_closure_rate : 100.0% -> 100.0% ( — )
response_shape_fit : 57.7% -> 57.7% ( — , both frames recognized)
audit_in_surface_rate : 42.3% -> 42.3% ( — , envelope ADR's job)
gloss_quote_rate : 11.5% -> 23.1% (+11.5pp, structural lift)
Tests (15)
- 5 pure composer (NOUN/VERB frame, unknown/empty None, no chain-
walk artifacts in surface)
- 5 runtime dispatch (flag-off chain-walk, flag-on gloss, parametrized
across glossed subjects, VERIFICATION unchanged under flag, no-
gloss fallback engages)
- 5 cognition lane invariance (aggregate metrics byte-identical
under both flag states; surfaces deliberately shift on the 2 CAUSE
cases with glossed subjects — the structural-change-vs-metric-
invariance both-sides invariant)
Lanes
smoke 67/0, cognition 120/0/1 skipped, packs 6/0, teaching 17/0,
runtime 19/0. core eval cognition byte-identical 100/91.7/100/100
under both flag states.
Scope limits (per ADR §Scope limits)
- CAUSE only; VERIFICATION still chain-walks (different shape).
- English pilot only; Greek/Hebrew packs not opted into definitional
layer yet (ADR-0084 scope limit).
- Single-lemma subjects; compound/anaphoric fall through.
- Opt-in until cognition holdout confirms the lift transfers off-
fixture. Future PR flips default on.
Out of scope
- Surface-vs-envelope cleanup ("pack-grounded (...)" still leaks).
- Predicate licensing (ADR-0086).
- Content style pass (bare lemma forms in glosses — separate brief).
Strict superset of ADR-0062's depth-1 composer. `max_depth` is the
number of follow-up hops appended beyond the initial chain:
max_depth=0 → byte-identical to single-chain surface
max_depth=1 → byte-identical to ADR-0062 composed
max_depth=2 → byte-identical to ADR-0062 when no second hop
survives, strict superset when one does
The composer surfaces content the realizer was silently dropping
from chains already ratified in `cognition_chains_v1`. Example
live lift on `"Why does light exist?"`:
composed: "light reveals truth, which grounds knowledge."
transitive(2): "...which grounds knowledge, which requires evidence."
Cycle-safe at every depth via a single visited-set; single-corpus
traversal in v1 (cross-corpus transitive deferred to a follow-up
ADR alongside ADR-0064's cross-pack model).
Both flags default False — every existing surface is preserved
byte-identically. When both `composed_surface` and
`transitive_surface` are True, transitive wins.
Implementation:
- `core/config.py`: `transitive_surface: bool = False`,
`transitive_max_depth: int = 2`.
- `chat/teaching_grounding.py`: `_resolve_followup` shared helper
refactored out of the depth-1 composer (no behavioural change),
plus new `teaching_grounded_surface_transitive(subject,
intent_tag, *, max_depth)`.
- `chat/runtime.py`: dispatch order — transitive > composed > single.
Verification:
- tests/test_transitive_surface.py: 16 new tests covering pure-fn
contract, visited-set cycle guard at every depth, runtime
integration, and the cognition-lane null-drop invariant at
`max_depth=2` (public + holdout splits).
- tests/test_composed_surface.py: 11/11 pass after the helper
refactor (ADR-0062 behaviour preserved).
- `core test --suite smoke`: 67 pass.
- `core test --suite cognition`: 120 pass, 1 skipped.
- `core test --suite teaching`: 17 pass.
- `core eval cognition`: 100 / 91.7 / 100 / 100 (byte-identical).
* chore(evals, cli): contract standardization + bench --json stdout cleanliness
End-of-session shippability pass. Three concrete fixes:
1. core/cli.py — bench --json no longer pollutes stdout
Several bench paths call scripts.run_pulse.run_pulse which prints
verbose [pulse] traces unconditionally to stdout, breaking jq /
programmatic consumers of --json output.
New _bench_stdout_guard() redirects stdout → stderr for the
duration of the bench run when --json is set. Operator still sees
the pulse trace (on stderr), but --json consumers get a clean JSON
document on stdout. Applied to all four bench paths: cost,
articulation, default suite, and --suite all.
Verified: core bench --suite determinism --json now produces
parseable JSON; human path still shows 1140 [pulse] lines.
2. evals/{frontier_compare,realizer_guard}/contract.md (new)
core/contemplation/contract.md (new)
Each new contract follows the established pattern (37 contracts
already exist under evals/<lane>/contract.md):
- What it measures
- Why it matters (structural win)
- How to run
- How to read the output
- Pass criteria table
- When it has failed and why
- Runner / module layout
Coverage:
- frontier_compare: both Lane A (CORE-only suites) and Lane B
(cross-provider prompt_battery) with explicit guardrails
against mixing — operator asks for the wrong lane combination,
runner exits 2 with helpful error.
- realizer_guard: C1/C2 articulation safety boundary — synthetic
illegal candidates rejected directly by check_surface AND
former-bug runtime prompts now produce legal articulations.
- contemplation (ADR-0080): not under evals/ since it's runtime
infrastructure that consumes eval reports — contract lives at
core/contemplation/contract.md. Documents the read-only +
SPECULATIVE-only + deterministic-replay invariants and the
shared DiscoveryCandidateSink plumbing convergence (ADR-0080).
3. evals/CLAIMS.md — Tier 2 rows added
- frontier_compare Lane A: determinism.primary_score, max_versor_condition
- frontier_compare Lane B: prompt_battery.primary_score (CORE adapter),
cross-provider artifact persistence
- realizer_guard: all_claims_supported
- contemplation: SPECULATIVE-only invariant, deterministic replay,
additive sink path, no pack mutation (all CI-pinned by tests)
Verification
------------
$ core test --suite smoke -q
67 passed in 27.22s (no regression)
$ uv run pytest -q tests/test_contemplation_loop.py \
tests/test_contemplation_pipeline_convergence.py \
tests/test_frontier_compare_cross_provider.py
27 passed in 4.87s
$ core bench --suite determinism --json 2>/dev/null | jq .results[0].passed
true (was: JSONDecodeError on prior [pulse] pollution)
* feat(evals/ui): report viewer renders Lane B cross-provider + pass-rate chart
Stop-hook caught that #62 only covered contracts — the 929-line
report_viewer.html was never audited against the new cross-provider
report shape from #61. Two real gaps:
1. Lane-aware observation drawer
The drawer hardcoded Lane A (CORE-native) fields: surface,
grounding_source, anchor_lens_mode_label, versor_condition.
Lane B (cross-provider) observations carry different fields:
provider, model, elapsed_ms, error_type, error_message.
Loading a cross-provider report rendered only the surface row
with empty `grounding` — the provider + model + timing data
was unreachable without expanding "Show raw JSON".
Fix: detect Lane B (presence of `obs.provider`) and render the
appropriate field set. Lane A still renders identically (now
also surfaces trace_hash + register_id when present, which were
silently buried in the raw JSON before).
2. Pass-rate chart per suite
The summary strip showed one aggregate Primary % across all
suites, with no way to see WHICH suite is dragging the score.
Multi-suite runs (e.g. --suite all) had to expand each panel
individually to find the failing one.
Fix: new .passrate-chart element below the summary strip,
one horizontal bar per suite showing passed/total. All-pass =
solid green, all-fail = solid red, partial = green/red split
at the pass fraction. CSS only — no new dependencies.
3. SUITE_PREAMBLES gains the prompt_battery entry so the sidebar
shows the "side-by-side surface evidence across providers"
description when loading a Lane B report.
Verified
--------
- Brace/paren/div balance unchanged (308/308 / 380/380 / 54/54)
- One <script> tag pair preserved
- Generated a real Lane B report via
`python -m evals.frontier_compare --provider core --suite prompt_battery`
for visual confirmation
Out of scope (noted for future PR)
----------------------------------
Sampled 3 `core demo` targets:
- register-tour: clean schema (all_claims_supported, claims, grid)
- audit-tour: both scene_1_* keys AND an empty scenes:[] array — inconsistent
- anti-regression: no all_claims_supported key, uses all_gates_held instead
Demo schema standardization deserves its own PR — operator tooling
would benefit from a uniform top-level success field across demos.
* docs(evals) + chore(demos): systematic audit + uniform success field
Stop-hook caught two real gaps after the contract+UI PR:
- demos had divergent success-field names (all_gates_held vs
learning_loop_closed vs claim_supported vs nested claims_supported)
- no systematic look at the 48 eval directories had been done
Both addressed concretely; remaining work captured in audit doc
rather than vaguely deferred.
1. Demo schema standardization — uniform all_claims_supported field
----------------------------------------------------------------------
All 9 ``core demo`` targets now emit a top-level
``all_claims_supported: bool`` field. Existing per-demo fields
(``all_gates_held``, ``learning_loop_closed``, ``claim_supported``,
nested ``claims_supported``) are preserved for backwards compat —
the new field is an alias derived from the demo's existing success
signal, not a replacement.
Operator tooling and the CI gate can now target
``all_claims_supported`` without knowing each demo's idiomatic
field name.
Files touched:
- evals/anti_regression/run_demo.py — adds AND of all_gates_held +
active_corpus_byte_identical
- evals/learning_loop/run_demo.py — adds AND of learning_loop_closed +
active_corpus_byte_identical
- scripts/publish_pack_measurements.py — adds AND of the three
entries in the nested claims_supported dict
- evals/long_context_cost/comparison_runner.py — adds alias for
claim_supported (singular)
The 5 demos already using ``all_claims_supported`` (audit-tour,
register-tour, anchor-lens-tour, orthogonality-tour, articulation)
are unchanged.
Verified across all 9 demos:
audit-tour : True
register-tour : True
anchor-lens-tour : True
orthogonality-tour : True
pack-measurements : True ← new alias
anti-regression : True ← new alias
learning-loop : True ← new alias
articulation : True
long-context-comparison : True ← new alias
2. docs/EVAL_AUDIT_2026-05-20.md — systematic 48-lane audit
------------------------------------------------------------
Replaces the "future PR" deferral with a concrete document.
Contains:
- Method (what was inspected for each lane).
- Summary (40/48 have contract.md; 18/48 have saved results;
empty results/ ≠ broken — most lanes regenerate on demand).
- Cross-provider relevance triage:
* 9 lanes are cross-provider-relevant and could benefit
from the prompt_battery-style adapter pattern (cognition,
english_fluency_ood, hebrew_fluency, koine_greek_fluency,
grammatical_coverage, inference_closure, multi_step_reasoning,
discourse_paragraph, foundational_*_ood, etc.).
* 29 lanes are CORE-only by design (versor closure, anchor
lens, identity divergence, provenance, etc.) — wiring
providers would be category-erroneous.
- Demo schema standardization status (this PR closes that).
- UI/UX coverage matrix.
- 5 concrete follow-up items, each focused enough for a single
PR, none requiring architectural change.
Regenerated reports
-------------------
evals/long_context_cost/results/comparison_v1.json and
evals/results/phase2_pack_measurements.json now contain the new
all_claims_supported field (auto-regenerated when validating the
schema change).
evals/frontier_compare/results/sample_core_promptbattery.json
added as a reference Lane B report so the new viewer always has
something to load on first open.
Two follow-up fixes from end-of-session verification of recent merges:
1. core/cli.py — wire `core contemplation` subcommand
PR #55 + #58 added the contemplation CLI at python -m core.contemplation
but never registered it under the `core` umbrella command, so
`core --help` didn't show it. Adds a subparser mirroring the existing
pattern (chat/test/check/.../doctor) that delegates to the existing
core.contemplation.__main__:main() — no duplication of arg parsing.
Surface preserved verbatim: reports (positional, 1+), --lane
{frontier_compare, contradiction_detection}, --pack-id, --note,
--report, --sink-root.
2. tests/test_architectural_invariants.py — restore INV-02 allowlist
PR #57's evals/lab/phi_separation_probe.py imports normalize_to_versor
for construction-time experimental rotor + embedding work, which
triggered INV-02's AST-scan failure (the test enforces that
normalize_to_versor is only called from a small allowed file set).
evals/lab/ is research-only, never imported by runtime — adding the
probe to allowed_files doesn't weaken the runtime invariant the
test enforces.
Verification
------------
$ core test --suite smoke -q
67 passed in 26.63s (was 66 passed / 1 failed before)
$ core contemplation --help
... shows the new subcommand surface
$ core contemplation evals/contradiction_detection/results/v1_public_*.json \
--lane contradiction_detection \
--sink-root /tmp/sink \
--report /tmp/run.json
... 4 SPECULATIVE findings; sink writes to /tmp/sink/2026/2026-05.jsonl
Connects ADR-0080's read-only contemplation loop to the existing
teaching-pipeline plumbing without forcing a type collapse. The
SPECULATIVE-only invariant from #55 is preserved verbatim; what
changes is *where the findings flow*.
What was wrong with the prior shape
-----------------------------------
PR #55 shipped a parallel core/contemplation/ package whose findings
were written as one JSON blob per CLI invocation, with no consumer.
The SPECULATIVE-only invariant protected a write path that didn't
exist. My closed PR #56 (second miner) would have entrenched the
duplication.
What this PR changes
--------------------
1. Schema (core/contemplation/schema.py)
- Adds a BOUNDARY note documenting why EvidencePointer (teaching)
and ContemplationEvidenceRef (core) intentionally stay separate:
EvidencePointer.source is constrained to {corpus, pack,
vault_coherent} — pointers into reviewed in-process memory the
runtime trusts. ContemplationEvidenceRef points to external
report files that have NOT been reviewed. Converging them would
either widen the runtime-grounding enum (losing the "reviewed
memory only" guarantee) or force benchmark reports to masquerade
as vault_coherent. Both are worse than keeping them separate.
- Adds format_contemplation_finding_jsonl(finding) — the canonical
JSONL formatter mirroring teaching.discovery.format_candidate_jsonl.
2. Runner (core/contemplation/runner.py)
- Both runners gain an optional sink: DiscoveryCandidateSink | None
parameter. When supplied, each finding is emitted as one
canonical JSONL line via the SHARED protocol — same protocol
that backs DiscoveryBufferSink and DiscoveryMonthlyFileSink.
- Sink path is additive: the ContemplationRun blob is byte-identical
whether or not a sink is supplied (pinned by test).
- No sink supplied → existing in-memory behavior preserved exactly.
3. CLI (core/contemplation/__main__.py)
- Adds --lane {frontier_compare, contradiction_detection} flag.
Default unchanged.
- Adds --sink-root <path> flag. When set, instantiates a
DiscoveryMonthlyFileSink and findings land at
<root>/<YYYY>/<YYYY-MM>.jsonl — the SAME layout discovery
candidates use, so operators can grep one stream.
4. Miner (core/contemplation/miners/contradiction_detection.py)
- Restored from closed PR #56 under the unified pipeline.
- Failure-mode split preserved (missed_contradiction /
false_contradiction_flag) with asymmetric repair actions.
What this PR does NOT do
------------------------
- Does NOT unify ContemplationFinding with DiscoveryCandidate.
DiscoveryCandidate.trigger is Literal[would_have_grounded,
successful_comparison, hedge_acknowledged, oov_resolved_via_decomp]
— all turn-loop flavored. None describe "I parsed a benchmark
report." Forcing a 5th trigger that no turn-loop extractor
produces would pollute the turn-loop type for the schema's sake.
- Does NOT extend teaching/gaps.py. Gap aggregates DiscoveryCandidate
cells by (subject, intent) — domain nouns. ContemplationFinding
subjects are namespaced ("contradiction_detection/CON-PUB-002").
Different operator views. A sibling aggregator can come later
when an operator actually asks for it.
Why this is the right unification point
---------------------------------------
The honest convergence is at the *sink* (so all SPECULATIVE evidence
lives in one rooted append-only stream), not the *aggregator* (which
appropriately produces typed views per evidence family). The boundary
doctrine from #55 is preserved; it now connects to existing plumbing
instead of writing JSON to disk with no consumer.
Tests (tests/test_contemplation_pipeline_convergence.py, 10 cases)
------------------------------------------------------------------
- DiscoveryBufferSink satisfies DiscoveryCandidateSink (shared protocol)
- frontier runner emits findings to shared sink
- contradiction runner emits findings to shared sink
- sink is optional — no-op when absent
- emission is canonical JSONL (sorted keys, no newline, deterministic)
- DiscoveryMonthlyFileSink persists findings at <root>/<YYYY>/<YYYY-MM>.jsonl
- sink emission does not alter the ContemplationRun blob (additive)
- contradiction miner predicate split + repair-action asymmetry
- config_hash differs between lanes (replay can distinguish)
- BOUNDARY doc is present in schema.py (regression guard)
- ContemplationEvidenceRef field invariants
- format_contemplation_finding_jsonl is deterministic + canonical
All 18 tests pass (5 original ADR-0080 + 13 new convergence).
Live evidence
-------------
$ uv run python -m core.contemplation \
evals/contradiction_detection/results/v1_public_*.json \
--lane contradiction_detection \
--sink-root /tmp/sink_demo
/tmp/sink_demo/2026/2026-05.jsonl ← same layout as discovery candidates
predicate=missed_contradiction subject=contradiction_detection/CON-PUB-002
predicate=missed_contradiction subject=contradiction_detection/CON-PUB-004
predicate=false_contradiction_flag subject=contradiction_detection/CON-PUB-005
predicate=false_contradiction_flag subject=contradiction_detection/CON-PUB-006
Wires observational telemetry on the composer-vs-graph atom-set
relationship. Phase 1 is strictly observational: no enforcement,
no surface mutation, no grounding-source change, no trace-hash impact.
New telemetry fields on TurnEvent + ChatResponse:
composer_graph_atom_status ∈ {equivalent, divergent,
graph_unconstrained,
composer_no_atoms,
not_applicable, ""}
composer_atom_set_hash SHA-256 over sorted unique atoms
graph_atom_set_hash SHA-256 over sorted unique atoms
composer_graph_atom_overlap_count int
Composer atoms come from existing pack candidate metadata
(pack_semantic_domains channel through _maybe_pack_grounded_surface).
Graph atoms come from build_graph_from_input + resolve_lemma on
node.subject/predicate/obj — no prose parsing. When a grounded
composer path lacks explicit atom provenance, status is
'composer_no_atoms'.
New pure helper:
chat/atom_equivalence.py — normalize_atoms, hash_atoms,
atoms_for_graph_nodes, compare_atom_sets
Tests (tests/test_composer_graph_atom_equivalence.py):
- Pack DEFINITION path produces observable equivalence
- Divergent atom sets produce distinct hashes
- Register invariance: atom hashes + status identical across
{neutral, terse, convivial}; trace_hash also constant (R5 axis)
- Anchor lens engaged case still ASCII-only on surface
- No prose-parsing helper symbols introduced in runtime.py
(extract_candidate_surface_lemmas, surface_lemma,
parse_surface_atoms) — enforces Phase 1 boundary
Performance note: build_graph_from_input now runs on every warm
English turn (previously only when forward_graph_constraint=True).
Phase 1 accepts this cost to make the telemetry universally
available; Phase 2+ can introduce a feature flag if needed.
Validation:
- Cognition eval byte-identical: 100/100/91.7/100
- Full lane: 2864 passed, 3 skipped, 0 failed (+5 over baseline)
- Targeted lane: 72 passed in tests/test_{graph_constraint,
pack_grounding,register_tour_demo,anchor_lens_tour_demo,
orthogonality_tour_demo,realizer_guard_holdout,
composer_graph_atom_equivalence}.py
R5 (ADR-0072) shipped the register *machinery*; ADR-0074's orthogonality
tour proved the axis was decoratively orthogonal to anchor-lens but
inspection of the cognition-eval surfaces revealed two structural gaps:
* On pack-grounded DEFINITION/RECALL/COMPARISON composers, the only
realizer override any register consumed was `disclosure_domain_count`
— which only fires on the no-gloss disclosure path. Under terse_v1,
every gloss-DEFINITION cell was byte-identical to default_neutral_v1.
* The register-tour's `surfaces_vary_at_least_once` gate could be
satisfied by convivial's decorative wrapper alone, masking that
regression in CI.
R6 closes both:
Layering separation (the load-bearing fix):
* New TurnEvent/ChatResponse field `register_canonical_surface` carries
the composer output BEFORE any register transformation. The pipeline
hashes this field for `trace_hash`, preserving R5's invariant that
per-prompt trace_hash is CONSTANT across registers even while
substantive transforms produce visibly different surfaces.
Substantive transforms (`chat/register_substantive.py`):
* terse_v1 gains 3 bool knobs: `drop_provenance_tag`, `compress_gloss`,
`drop_articles` — all pure regex transforms on the canonical surface.
* convivial_v1 gains `append_semantic_domain_clause` — appends a single
bounded "Related: <atom>." clause using the lemma's pack atoms.
* default_neutral_v1 leaves overrides empty; substantive transform is
byte-identical no-op (preserves `byte_identity_null_lift`).
* C1 (ADR-0075) safety preserved: drop_articles refuses to drop
articles following `not` (avoids R3 violations); no knob combination
trips R2/R3.
Strengthened tour gate (`evals/register_tour/run_tour.py`):
* Replaces `surfaces_vary_at_least_once` with two falsifiable claims:
- `terse_substantively_differs_from_neutral_on_pack_grounded_definition`
- `convivial_substantively_differs_from_neutral_on_pack_grounded_definition`
Both restrict to DEFINITION+pack-grounded cells and require
difference beyond whitespace/punctuation.
* New claim `register_canonical_surfaces_identical` directly proves
the layering separation.
* Preserves R5's `all_grounding_sources_identical` +
`all_trace_hashes_identical`.
Pack ratification:
* Loader widened to accept `bool` for closed-set R6 keys
(drop_provenance_tag / compress_gloss / drop_articles /
append_semantic_domain_clause).
* `_KNOWN_OVERRIDE_KEYS` ratify gate extended with same.
* terse_v1 + convivial_v1 reratified with new knobs; companion
mastery reports re-sealed. default_neutral_v1 unchanged.
Invariants pinned:
* `invariant_register_canonical_surface_constant_across_registers` (new)
* `invariant_terse_substantively_distinct_from_neutral` (new)
* `invariant_convivial_substantively_distinct_from_neutral` (new)
* `invariant_realizer_no_illegal_articulation` (C1, preserved)
* `invariant_realizer_guard_byte_identity_on_currently_passing_cases`
(C1, preserved)
Verification:
* `core eval cognition`: 100.0% / 91.7% / 100.0% / 100.0% — byte-
identical under default_neutral_v1.
* `core demo register-tour`: all 5 claims green, exit 0.
* `core demo anchor-lens-tour`: green (no anchor-lens code touched).
* `core demo orthogonality-tour`: green (5/5 claims).
* Full lane: 2858 passed, 1 pre-existing failure
(test_all_preamble_explains_combined_run, carried forward
unchanged from main). 56 new R6 tests across three files.
C1 coherence floor: a deterministic verifier that runs on every
candidate surface produced by the truth path, before assignment to
ChatResponse.surface. Rejects illegal articulations and routes them
to a bounded disclosure string — admission control with a
deterministic fallback, not normalization.
Active rules (R1 deferred during ratification — see ADR):
R2_aux_neg_requires_verb — "<aux> not <wrong-POS>" rejected
R3_be_neg_requires_predicate — "<be> not <verb>" rejected
Fail-open on unknown POS, fail-closed on explicit wrong POS.
Cognition eval byte-identical (100/91.7/100/100).
Original bug class — "Light reveals truth, right?" → "Right does not
thought." — now routes to "I do not have a reviewed articulation for
that yet." with grounding_source=none, walk_surface preserving the
rejected candidate, and telemetry carrying R2_aux_neg_requires_verb.
Files:
generate/realizer_guard.py NEW — pure verifier
chat/runtime.py hook on stub + main paths
chat/telemetry.py serialize guard fields
core/physics/identity.py TurnEvent +2 fields
evals/realizer_guard/run_holdout.py NEW — 6-prompt cluster
tests/test_realizer_guard_*.py NEW — 46 tests (unit/seam/holdout)
docs/decisions/ADR-0075-*.md NEW — ratified
Invariants pinned:
invariant_realizer_no_illegal_articulation
invariant_realizer_guard_byte_identity_on_currently_passing_cases
Lanes (excluding 1 pre-existing TestDemoPreambles failure unrelated
to C1, already present at 4426f38):
smoke 67/67 cognition 120/120(+1s) teaching 17/17
packs 6/6 runtime 19/19 algebra 132/132 full 2792/2793
A single demo that walks the full 3 × 3 × 2 matrix (register × lens
× prompts, 18 cells) and pins five claims simultaneously, packaging
both single-axis invariants into one composition gate.
The single-axis tours assert opposite invariants:
register-tour : per (lens, prompt), trace_hash CONSTANT across
registers (R5 / ADR-0072).
anchor-lens-tour : per (register, prompt), engaged lens diverges
in trace_hash from the unanchored baseline
(L1.4 / ADR-0073d).
Orthogonality-tour packages both claims simultaneously across the
full matrix, plus three surface-level claims that pin the markers
operators actually see.
Composed claims (all five must hold)
A) inner_register_invariant_within_lens
For each (lens, prompt) cell, the three register runs share an
identical trace_hash. (R5 register-tour, applied 6 times:
3 lenses × 2 prompts.)
B) outer_lens_distinctness_within_register
For each (register, prompt) cell where any non-unanchored lens
engages, that engaged lens's trace_hash differs from the
unanchored baseline at the same (register, prompt).
(L1.4 anchor-lens-tour, applied 6 times: 3 registers × 2 prompts.)
C) surface_carries_register_marker_under_convivial
Every convivial cell with a non-empty surface has a non-empty
register_variant_id.
D) surface_carries_lens_annotation_when_engaged
Every engaged cell carries [lens(<id>):<mode>] in surface AND
a non-empty anchor_lens_mode_label.
E) no_substrate_glyph_leak_across_grid
No cell's surface contains Greek/Hebrew/Syriac/Arabic glyphs.
(ADR-0073c gate re-asserted across the full matrix.)
CLI wiring
core demo orthogonality-tour human-readable grid + claims
core demo orthogonality-tour --json structured report
Exit code 0 iff all five claims hold.
Files
evals/orthogonality_tour/__init__.py NEW
evals/orthogonality_tour/run_tour.py NEW
core/cli.py EDIT
- cmd_demo handler wires orthogonality-tour
- demo choices + EPILOG examples updated
tests/test_orthogonality_tour_demo.py NEW (9 tests)
docs/decisions/ADR-0074-orthogonality-tour.md NEW
Sanity check baked into tests
test_engaged_cells_appear_for_both_non_trivial_lenses pins that
grc_logos_v1 engages on knowledge in all 3 registers (3 cells)
and he_logos_v1 engages on truth in all 3 registers (3 cells).
Prevents the lift claims being vacuously satisfied by a future
engagement regression.
Lane evidence
- 9 new orthogonality-tour tests pass.
- core demo register-tour → all_claims_supported: True
- core demo anchor-lens-tour → all_claims_supported: True
- core demo orthogonality-tour → all_claims_supported: True
- python -m core.cli eval cognition → byte-identical 100/100/91.7/100.
- Full lane: 2745 passed / 4 skipped / 1 pre-existing failure
(+9 over L1.4's 2736; the one failure remains
test_all_preamble_explains_combined_run, unrelated).
No runtime / composer / loader / pack / schema changes. Pure demo
consumer of existing telemetry contracts.
A live walkthrough that shows CORE actually being used. Four scenes,
five turns, rendered as a chat transcript ('You: …' / 'CORE: …') with
plain-English captions between turns.
Streamed by default (per-character prompt, per-word response, brief
"thinking" pause) so the layperson sees the answer arriving live.
--no-stream disables delays for CI / tests / fast capture.
Scenes:
1. Pack lookup — "What is truth?"
Shows deterministic lexicon-grounded answer.
2. Teaching-chain — "Walk me through recall."
Shows CORE chaining reviewed facts.
3. Compound prompt — "What is truth, and why does it matter?"
Shows compound decomposition + composition.
4. Cold turn → learn — "Why does narrative exist?"
Shows CORE refusing to fabricate, an operator
teaching it one new chain (real propose →
replay-gate → accept), then re-asking the same
prompt and getting a grounded answer.
The learning-loop scene reuses the production learning_loop demo so
the underlying machinery is exactly what ships — active corpus is
byte-identical pre/post.
Test gate: tests/test_conversation_demo.py (9 tests — per-scene
grounding source + content checks, learning loop closes,
active-corpus byte-identical, stable JSON shape).
Usage:
core demo conversation # live streamed transcript
core demo conversation --no-stream # instant rendering
core demo conversation --json # structured report (no chat output)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Renames the original phase5+phase6 combo to its more honest name
'adr-0024-chain' and repurposes 'all' to mean what users expect: every
demo (eight in total) in one shot.
Demos covered:
1. phase5 — stratified mechanism isolation
2. phase6 — three-condition head-to-head
3. audit-tour — pack-layer story
4. pack-measurements — pack-layer claims → numbers
5. long-context-comparison — exact NIAH vs transformer baselines
6. anti-regression — three-gate defense
7. learning-loop — cold turn → grounded surface
8. articulation — discourse-planner spine
Per-demo runners retain their native preambles + reports. The
aggregator captures each demo's load-bearing boolean (already pinned
by that demo's test gate) and prints a consolidated PASS/FAIL table.
Exits non-zero if any demo fails.
Under --json, sub-runner stdout is suppressed and a single
consolidated JSON object is emitted with one key per demo plus
'passed' and 'all_demos_passed'.
'core demo adr-0024-chain' preserves the historical phase5+phase6
combined-summary semantics for callers who depended on it.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Four-scene investor/operator-facing walkthrough proving the discourse-
planner spine is load-bearing. Each scene runs the same prompt under
flag-off (BRIEF baseline) and flag-on (RuntimeConfig.discourse_planner)
and pins a falsifiable lift assertion.
S1. EXPLAIN — Explain truth.
Flag-on: pack→teaching upgrade + 2 chain
continuation sentences over baseline.
S2. COMPOUND — What is truth, and why does it matter?
Flag-on: 9 grounded sentences across two sub-
plans; flag-off routes to OOV.
S3. WALKTHROUGH — Walk me through recall.
Flag-on emits the CLOSURE chain hop
'Recall reveals memory.'; flag-off
does not.
S4. Determinism — N=3 reruns × 3 prompts, unique(surface)=1.
Read-only against live packs + active corpus. Demo is test-gated
(7 tests, all green) and ships a stable JSON contract for downstream
consumers.
Wired into CLI as `core demo articulation [--json]` alongside the
existing trilogy (audit-tour / anti-regression / learning-loop).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds an aggregate ``all`` choice to ``core bench --suite`` that
exercises every benchmark CORE ships:
[1/4] Core six — determinism / latency / speedup / versor /
convergence / realizer (via run_benchmarks)
[2/4] Teaching-loop determinism
[3/4] Articulation suite — breadth / determinism / footprint /
cross-topic / discourse-planner /
ollama
[4/4] Cost — measurement bench (no PASS/FAIL by design)
Behavior:
* Each section prints its native report shape (run_benchmarks rows,
articulation summary, cost summary). Final consolidated tally
prints ALL PASSED / FAILURES DETECTED across the three pass/fail
groups; cost is reported separately as a measurement section so
it can't false-positive the gate.
* JSON mode emits a single consolidated object with one key per
section so a downstream report consumer gets every artifact from
one command.
* psutil is treated as optional: when missing, the articulation
footprint sub-bench is skipped (new ``skip_footprint`` kwarg on
``run_articulation_suite``) instead of aborting the whole run.
The other three articulation sub-benches all run, so the spine's
determinism + planner-on capability evidence is preserved.
CLI surface:
core bench --suite all
core bench --suite all --runs 50
core bench --suite all --json --report bench_all.json
Defaults for ``core bench`` (no suite) are untouched — still runs
the six core benches exactly as before.
EPILOG examples updated; ``--suite`` ``choices`` extended with
``"all"``.
Validation:
* core bench --suite all --runs 3: 4 sections run end-to-end;
consolidated tally reports per-bench PASS/FAIL. Pre-existing
backend_speedup FAIL (0.9999x — Rust kernel not built locally)
surfaces correctly; every other bench PASS including
articulation_suite_overall.
* core bench --runs 3 (no --suite): unchanged behavior, same six
benches as before.
* tests/test_articulation_bench.py + test_cli*.py: 25 passed.
* smoke suite 67/67.
Step 5 of the discourse-planner sequencing. Closes the chain:
classify_intent + classify_response_mode
-> grounding_bundle_for(subject)
-> plan_discourse(intent, mode, bundle)
-> render_plan(plan)
-> response_surface
Adds RuntimeConfig.discourse_planner (default False). When True, the
runtime — after the warm pack/teaching-grounded surface is set —
classifies the response mode, assembles a GroundingBundle from the
ADR-style accessors, builds a DiscoursePlan, and replaces the warm
surface with the deterministic multi-clause rendering whenever the
plan has more than one move.
Gating discipline:
* Engages only on warm_grounding_source in {"pack", "teaching"} so
vault/none turns and the discovery-signal CAUSE/VERIFICATION
disclosure are preserved exactly.
* BRIEF mode always collapses to a single ANCHOR move, so flag-on
with BRIEF intent is byte-identical to flag-off.
* Empty bundles produce empty plans; the runtime falls through to
the existing warm surface untouched.
Adds render_plan(plan) to generate/discourse_planner.py — a pure,
deterministic multi-clause renderer with fixed canonical connectives:
ANCHOR : capitalized opening sentence
SUPPORT : "Furthermore, ..."
RELATION : "In turn, ..."
TRANSITION: "Consequently, ..."
CLOSURE : skipped when fact is None
Every visible token is a verbatim pack lexicon entry, gloss, or
reviewed teaching chain string — no synthesis.
13 new tests pin:
* render_plan empty/brief/paragraph shape
* canonical connectives present in paragraph rendering
* deterministic + verbatim-fact invariants
* RuntimeConfig.discourse_planner defaults False
* Flag-off surface has no planner connectives
* Flag-on lifts produce structurally well-formed multi-sentence
output on grounded substrate
Lift measurement (multi_sentence_response public/v1, 15 cases):
* flag off: multi=0.40, connective=0.50, grounded=0.40
* flag on : multi=0.40, connective=0.60, grounded=0.40
-> connective_present_rate +10pp; multi-sentence count flat
because the existing narrative composer's literal "." chars in
tags like "cognition.truth" already trigger sentence splits in
the lane regex. Real lift is form quality: e.g. "Tell me about
truth" now renders as "Truth is a claim or state grounded by
evidence and coherent judgment. Furthermore, truth belongs to
cognition.truth. In turn, truth grounds knowledge." instead of
the prior provenance-laden narrative surface.
Critical gates (all green):
* flag off: cognition eval byte-identical
- public 100/100/91.7/100, holdout 100/100/83.3/100
* smoke suite 67/67
* conversational_thread_coherence: 3 unwanted placeholders flag off
and flag on (no regression)
* planner JSON byte-stable across calls (contract tests)
* grounding source order preserved (sidecar tests)
The 2026-05-19 design review's P0 #1 finding:
> CognitiveTurnPipeline can replace a useful runtime surface with
> placeholder prose.
Evidence at core/cognition/pipeline.py:147-149 (pre-fix):
if realized_plan.surface and not gate_fired:
surface = realized_plan.surface
articulation_surface = realized_plan.surface
The override gate was JUST "non-empty + gate didn't fire". No
usefulness check. Result: a realizer output of
"Truth is defined as ..." (with <pending> rendered as ...) silently
overrode a perfectly-grounded runtime pack surface, and the runtime
audit log still held a third surface.
Fix: gate the override through ``_is_useful_surface`` from
generate/intent_bridge.py — the same predicate that already gates
the bridge's articulate_with_intent fallback path. An ungrounded
realizer surface cannot honestly override a grounded runtime
surface. When the realizer cannot produce a useful surface, we
keep the runtime answer the user sees.
Measured lift on the warmed_session_consistency lane (3 of its 4
metrics):
BEFORE AFTER
no_placeholder_rate 0.4444 → 1.0000
telemetry_consistency_rate 0.4444 → 1.0000
warm_grounding_stability 0.0000 → 0.0000 (separate bug — see below)
The two metrics that flipped to 1.00 are now CI-pinned in
tests/test_warmed_session_lane.py:
TestPipelineOverrideGateInvariants — any future weakening of the
override gate fails the suite immediately.
Cognition eval byte-identical:
public: 100 / 100 / 91.7 / 100
holdout: 100 / 100 / 83.3 / 100
KNOWN FOLLOW-UP — not in this commit:
warm_grounding_stability remains 0.0 because of a SEPARATE bug
the warmed lane surfaces:
Turn 1: "What is truth?" -> pack-grounded ("truth — pack-grounded
(en_core_cognition_v1): cognition.truth; ...")
Turn 2: "What is truth?" -> vault-grounded ("Truth infer.")
After turn 1 ingests pack content into the vault, turn 2's gate
source flips from ``empty_vault`` to ``vault``, so the runtime's
``_maybe_pack_grounded_surface`` dispatcher is bypassed entirely
and the field-walk path produces gibberish ("Truth infer.").
This is the SurfaceSelector-shaped problem from the design review:
pack-grounding should fire by intent shape and lemma residency, not
by vault gate state. Fix scope crosses runtime.py:chat() + the
vault gate logic; deferred to its own commit / design proposal
rather than absorbed here.
The warmed lane already records the metric (0.0 baseline) so when
the fix lands it shows up as a measurable lift.
Workstream 1 eighth pack. Closes the polarity-marker + frequency-
adverb gap. Common conversational markers (yes/no/maybe/always/never)
had zero coverage in any prior pack.
Pack composition (16 entries — 2 INTJ / 14 ADV):
polarity.affirm.* yes indeed surely definitely
polarity.negate.* no hardly
polarity.uncertain.* maybe perhaps
polarity.frequency.* always sometimes often rarely never
usually occasionally frequently
``certain``/``certainly``/``uncertain`` deliberately excluded — those
remain in en_core_attitude_v1 (epistemic.certainty/uncertainty).
Regression test pins the invariant.
tests/test_correction_topic_lemma.py:
Three fixtures swapped from "No that is wrong" to "Nope that is
wrong". ``no`` is now correctly pack-resident in en_core_polarity_v1
(polarity.negate.dissent), so the "no pack-resident lemma" contract
these tests pin needed a fixture where every content token is
genuinely OOV. ``nope`` is OOV across all 10 mounted packs; ``wrong``
remains OOV (collision with attitude's ``right`` blocked spatial-
direction ``right`` but did not add ``wrong``).
Authoring:
Three parallel subagents — affirm / negate+uncertain / frequency.
Workstream 1 sixth pack. Closes the spatial-vocabulary gap. Prior
packs had zero coverage of here/there, location nouns, or spatial
prepositions.
Pack composition (24 entries — 7 ADV / 8 ADP / 9 NOUN):
spatial.deictic.* here there (2 ADV)
spatial.direction.* forward backward left up down (5 ADV)
spatial.relation.* near far above below inside outside
between beyond (8 ADP)
spatial.noun.* place location area region space
end top bottom side (9 NOUN)
``right`` was deliberately omitted — en_core_attitude_v1 already owns
it as evaluative.positive, and first-match-wins resolution preserves
that claim. A regression test pins this invariant explicitly.
Files: lexicon.jsonl / manifest.json + 12 contract tests.
Verification: full lane 2204 passed / 2 skipped / 0 failed.
Cognition eval byte-identical both splits.
Workstream 1 fifth pack. Closes the quantifier + basic-numeric gap.
Prior packs had zero coverage of universal / existential / comparative
quantifiers — queries about *all*, *some*, *many*, *more*, *most* all
fell through to OOV.
Pack composition (24 entries — mixed POS, 18 DET / 3 NUM / 2 ADJ / 1 NOUN):
quantitative.universal.* (6 DET) all every each both none neither
quantitative.existential.* (6 DET) some any several few many much
quantitative.comparative.* (6 DET) more less fewer most least enough
quantitative.numeric.* (3 NUM) one two three
quantitative.unit.* (3 mix) single (ADJ) half (NOUN) whole (ADJ)
The composer is POS-agnostic; surface composition uses
``semantic_domains`` rather than POS, so DET/NUM/ADJ/NOUN entries all
surface identically.
Files:
language_packs/data/en_core_quantitative_v1/
lexicon.jsonl — 24 entries, SHA-256 checksum-sealed
manifest.json — operational_base / D0
chat/pack_resolver.py
Appended to DEFAULT_RESOLVABLE_PACK_IDS after action.
core/config.py
Added to RuntimeConfig.input_packs default mount.
tests/test_en_core_quantitative_v1_pack.py
11 contract tests (load / POS-dist / namespace / no-collision /
contiguous-ids / mount / resolver-order / routing / invariance).
Authoring:
Three parallel subagents — universal+existential / comparative /
numeric. Strict exemplar + forbidden-lemma list against all 7
prior packs.
Verification:
Full lane: 2192 passed, 2 skipped, 0 failed.
Cognition eval byte-identical on both splits.
Workstream 1 fourth pack. Closes the common-action verb gap. Prior
packs covered reasoning (cognition), speech/perception (meta), and
adjectives (attitude); this pack covers what an agent *does*.
Pack composition (26 VERB entries):
action.doing.perform do perform execute carry conduct
action.doing.make make
action.doing.achieve achieve accomplish
action.creating.originate create build form produce generate develop
action.changing.transform change transform
action.moving.translate move
action.moving.depart_arrive go come
action.moving.transfer send receive
action.possessing.acquire get take
action.possessing.transfer give
action.possessing.retain keep
action.possessing.deploy use
Files:
language_packs/data/en_core_action_v1/
lexicon.jsonl — 26 entries, SHA-256 checksum-sealed
manifest.json — operational_base / D0
chat/pack_resolver.py
Appended to DEFAULT_RESOLVABLE_PACK_IDS after temporal.
core/config.py
Added to RuntimeConfig.input_packs default mount.
tests/test_en_core_action_v1_pack.py
11 contract tests covering load / POS / namespace / no-collision /
contiguous-ids / mounted-by-default / resolver-order / routing /
prior-pack invariance.
tests/test_procedure_surface.py
Swapped two test fixtures from "do stuff" to "fix bugs". ``do``
is now correctly pack-resident in en_core_action_v1 (semantically
correct — "How do I do stuff?" should ground on ``do``), so the
"no pack lemma exists" contract needed a fixture where both verb
and noun are genuinely OOV. ``fix bugs`` satisfies this across
all 7 mounted packs.
Authoring:
Three parallel subagents — doing / creating / moving+possessing.
Strict exemplar + forbidden-lemma list against all 6 prior packs.
Verification:
Cognition eval byte-identical on both splits (100/100/91.7/100 and
100/100/83.3/100).
All 70 pack tests pass (cognition + meta + attitude + temporal +
action + quant tests run together).
Live composer probes confirm every action lemma surfaces
deterministically from en_core_action_v1.
Workstream 1 third pack. Closes the temporal-vocabulary gap — prior
to this pack zero time/sequence/aspect terms existed in any mounted
English pack, so queries about *when*, *before*, *after*, *now*,
*future*, *past* all fell through to OOV.
Pack composition (28 entries, mixed POS — 12 ADV / 9 NOUN / 5 ADP /
1 SCONJ / 1 ADJ):
temporal.deictic.* (10 ADV) now today tomorrow yesterday soon
later recently eventually currently
formerly
temporal.relative.* (9 mix) before after during while until since
ago prior henceforth
temporal.noun.* (9 NOUN) moment period duration instant era
future past present time
The pack composer is POS-agnostic — surface composition uses the
ratified ``semantic_domains`` list rather than the POS tag. Mixed-POS
entries surface identically to noun/verb entries.
Files:
language_packs/data/en_core_temporal_v1/
lexicon.jsonl — 28 entries, SHA-256 checksum-sealed
manifest.json — operational_base / D0 / checksum-verified
chat/pack_resolver.py
Appended to DEFAULT_RESOLVABLE_PACK_IDS after attitude.
core/config.py
Added to RuntimeConfig.input_packs default mount.
tests/test_en_core_temporal_v1_pack.py
11 contract tests: checksum, POS-distribution invariant, primary-
domain namespace, no-collision regression gate against all 5 prior
packs, contiguous entry_ids, mounted-by-default, resolver-order
invariant, routing correctness, and prior-pack resolution unchanged.
Authoring:
Three parallel subagents — deictic / relative / nouns. Strict
exemplar + forbidden-lemma list against all 5 prior packs.
Verification:
Full lane: 2170 passed, 2 skipped, 0 failed (+11 new tests).
Cognition eval byte-identical on both splits.
Live composer probes confirm every temporal lemma surfaces
deterministically from en_core_temporal_v1.
Workstream 1 second pack. Closes the ADJ POS gap — prior to this pack
zero adjectives existed in any mounted English content pack, so the
runtime could not emit grounded surfaces for predicative queries like
"What is true?" or "What is important?".
Pack composition (40 ADJ entries):
attitude.truth_value.* (8) true false valid invalid accurate
inaccurate factual sound
attitude.evaluative.* (6) good bad right better worse best
attitude.epistemic.* (10) certain uncertain possible impossible
likely unlikely probable clear obscure
evident
attitude.modal.* (4) necessary sufficient required optional
attitude.importance.* (6) important essential relevant central
primary useful
attitude.scope.* (6) general specific broad narrow universal
particular
Files:
language_packs/data/en_core_attitude_v1/
lexicon.jsonl — 40 entries, SHA-256 checksum-sealed
manifest.json — operational_base / D0 / checksum-verified
chat/pack_resolver.py
Appended to DEFAULT_RESOLVABLE_PACK_IDS after cognition + meta.
core/config.py
Added to RuntimeConfig.input_packs default mount.
tests/test_en_core_attitude_v1_pack.py
11 contract tests: checksum, POS=ADJ uniformity, primary-domain
namespace, no-collision regression gate against all 4 prior packs,
contiguous entry_ids, mounted-by-default, resolver-order invariant,
routing correctness, and cognition+meta resolution unchanged.
Authoring:
Three parallel subagents (1 per cluster) — truth/eval, epistemic/modal,
importance/scope. Strict exemplar + forbidden-lemma list against all
prior packs. Main pass assembled, validated, sealed.
Verification:
Full lane: 2159 passed, 2 skipped, 0 failed (+11 new tests over the
previous 2148 baseline).
Cognition eval byte-identical on both splits:
public 100 / 100 / 91.7 / 100
holdout 100 / 100 / 83.3 / 100
Live composer probes: every ADJ lemma emits a deterministic
pack-grounded surface from en_core_attitude_v1.
Workstream 1 (pack content scale-up) first load-bearing step.
Adds a new ratified content pack covering the conversational vocabulary
en_core_cognition_v1 deliberately omits — speech acts, mental states,
perception, self-reference, and discourse-object nouns. These are the
lemmas that show up in nearly every model response and that previously
fell through to the OOV invitation surface.
Pack composition (73 entries, 49 VERB + 24 NOUN):
meta.speech_act.* (20 verbs) say tell speak reply claim state
describe express name mention note
observe declare assert deny confirm
suggest propose articulate respond
meta.mental_state.* (18 verbs) know believe think suppose assume
expect hope want prefer doubt wonder
guess recognize realize consider intend
decide hold
meta.perception.* (11 verbs) see hear feel sense perceive watch
look listen find detect notice
meta.self_reference.* (10 nouns) self mind view perspective position
role agent model system speaker
meta.discourse.* (14 nouns) response reply statement fact idea
point argument proposal suggestion
case instance example kind type
Files:
language_packs/data/en_core_meta_v1/
lexicon.jsonl — 73 entries, SHA-256 checksum-sealed
manifest.json — operational_base / D0 / checksum-verified
chat/pack_resolver.py
Appended en_core_meta_v1 to DEFAULT_RESOLVABLE_PACK_IDS after
en_core_cognition_v1 so cognition lemma resolution stays first-
match-wins on any future collision (preserves cognition-lane
byte-identity invariant).
core/config.py
Added en_core_meta_v1 to RuntimeConfig.input_packs default mount.
tests/test_en_core_meta_v1_pack.py
11 contract tests: checksum-verified load, POS split, primary-
domain namespace, no-collision-with-cognition-v1 regression gate,
pack registration order, resolver routing, and cognition-lemma
resolution unchanged.
tests/test_procedure_surface.py
Swapped two test fixtures from "claim" to "hypothesis". ``claim``
is now correctly pack-resident (meta.speech_act.claim) so the
procedure composer's object-first selector picks it over the verb
— the new behavior is semantically correct. ``hypothesis`` is
genuinely OOV across all mounted packs and preserves the verb-
fallback contract these tests pin.
Authoring methodology:
Four parallel subagents authored one cluster each from a strict
exemplar + word list + forbidden-lemma list (every en_core_cognition_v1
lemma listed explicitly to prevent collision). Each subagent wrote
only its cluster JSONL; the main pass assembled, validated, computed
the SHA-256 over bytes-on-disk, and wrote the manifest.
Verification:
Full lane: 2148 passed, 2 skipped, 0 failed (+11 new tests).
Cognition eval byte-identical on both splits:
public 100 / 100 / 91.7 / 100
holdout 100 / 100 / 83.3 / 100
Live runtime probes: fresh ChatRuntime() for "What is X?" with
X ∈ {fact, doubt, statement, model, self} all emit a
pack-grounded sentence from en_core_meta_v1.
OOV path still honest for genuinely-unknown terms (e.g. hypothesis).
Scope note:
This is one pack of ~70 lemmas, not "the model now articulates
open-domain English." The architecturally-honest articulation
story still requires more pack and teaching-chain content; this
pack moves the conversational-substrate boundary forward by ~70
lemmas in one ratifiable, replay-stable step.
Phase 5 (ADR-0067 follow-up):
teaching/cross_pack_supersede.py — supersede_cross_pack_chain()
CLI: core teaching supersede ... --cross-pack
--subject-pack-id ... --object-pack-id ...
Strict per-chain residency, anti-leakage, byte-identical rollback
on any post-append re-load failure. 9 new tests.
Articulation benchmark suite (Phase 4 capability proof):
benchmarks/articulation.py — 5 sub-benches
[1] breadth — every intent shape (9 + OOV + cross-pack)
[2] determinism — N reruns / unique-surface count
[3] footprint — psutil RSS profile across T turns
[4] cross-topic — thread context across mixed subjects
[5] ollama-compare — opt-in side-by-side with local Ollama
CLI: core bench --suite articulation
--runs N (det rerun count)
--turns N (footprint sample window)
--ollama-model MODEL --ollama-reruns N
Full operator preamble + JSON report path.
10 new tests cover the bench shape (psutil import-skipped).
Documentation:
benchmarks/README.md — full operator manual: catalogue of every
bench suite, how to read good/neutral/bad results for each sub-
bench, why CORE vs Ollama comparisons are valid on the
determinism axis and not on linguistic quality, workflow guide.
README.md — articulation bench listed in the live-demo grid and
quick-start examples.
Reference run (llama3:8b, 100 turns, 5 reruns):
determinism_all_identical=True
per-turn ΔRSS ≈ 23 KiB
CORE byte_identical_on_every_prompt=True
Ollama unique_surfaces≥2 on every prompt
Verification:
18 new tests pass
Full lane: 2116 passed, 2 skipped, 0 failed in 2:38
ADR-0066 P3.1 + P3.2. Conversation now reads as a thread: turns
carry structured summaries of their predecessors and (optionally)
prefix new pack/teaching surfaces with deterministic backreferences.
P3.1 — chat/thread_context.py.
TurnSummary(turn_index, intent_tag_name, subject, grounding_source,
chain_id, corpus_id) — frozen, structured-fields-only.
ThreadContext — bounded FIFO (default MAX_THREAD_TURNS=8) with
snapshot(), recent_for_subject(), recent_subjects(), clear().
recent_for_subject() excludes ungrounded tiers (oov/partial/none)
by default — those are not strong-enough anchors.
ChatRuntime.thread_context is owned at construction.
_push_thread_summary runs at end-of-turn on BOTH stub and walk
paths. Teaching-grounded turns carry chain_id + corpus_id so
downstream composers (P3.2) can detect same-chain reference.
Cold-start intent classification now runs unconditionally (was:
gated on sink attachment) so thread context captures subject
regardless of sink state.
P3.2 — chat/anaphora.py.
thread_anaphora_prefix(ctx, subject, intent_name, source) returns
a deterministic prefix when:
- current turn is pack/teaching tier
- a prior pack/teaching turn on the same subject exists
- the prior intent differs from the current intent
Format (structural-fields-only — no prose):
"(Recalling turn N: chain <chain_id>.) " # prior was teaching
"(Recalling turn N: <subject> grounded pack.) " # prior was pack
Opt-in via RuntimeConfig.thread_anaphora=False. Default off keeps
every existing surface byte-identical.
Live verification (with thread_anaphora=True + seeded context):
> What is light? # following a "Why does light exist?" teaching turn
[pack] (Recalling turn 0: chain cause_light_reveals_truth.)
light — pack-grounded (en_core_cognition_v1): cognition.illumination;
logos.core; perception.clarity. No session evidence yet.
32 new tests passed. Curated lanes green. Cognition eval
byte-identical to pre-ADR baseline.
Mirrors the chain-gap pipeline (Phase 1.1+1.2) for vocabulary gaps:
the OOV invitation surface (P2.1) now emits structured signals that
operators can aggregate, rank, and auto-promote into reviewed
PackMutationProposal candidates — closing the OOV loop the same way
Phase 1 closed the chain loop.
Three new modules + two new CLI surfaces:
teaching/oov_sink.py.
OOVCandidate dataclass mirroring teaching.discovery.DiscoveryCandidate.
OOVBufferSink (in-memory) + OOVMonthlyFileSink (append-only JSONL
under <root>/<YYYY>/<YYYY-MM>.jsonl — same layout as discovery sink
so the aggregator reuses the file-walk machinery).
hash_oov_candidate_id(token, intent, trace_hash) — deterministic
32-char hex id matching DiscoveryCandidate's replay invariant.
format_oov_candidate_jsonl — sorted-keys compact JSONL line.
teaching/oov_gaps.py.
aggregate_oov_gaps(root, since, sample_limit) groups emitted
candidates by token, tracks intent-shape union (a token asked under
multiple intents is a stronger curriculum signal), splits
boundary_clean from boundary_tainted counts, supports --since
YYYY-MM filtering via the sink's file naming convention.
Pure reader; never mutates the sink. Deterministic ordering:
(count desc, token asc).
teaching/oov_promotion.py.
promote_oov_gaps(gaps, threshold, include_tainted, suggested_packs)
lifts threshold-crossing tokens to OOVPromotion records.
- boundary_clean_count gates promotion by default (tainted-only
tokens may indicate the prompt hit a safety axis rather than a
vocab gap).
- --include-tainted flag for operator override.
- threshold < 1 raises.
- queue_id deterministic: ``oov:<token>@<threshold>`` — diffable
across runs.
- suggested_packs lists mounted packs but does NOT recommend one
— domain inference is out of scope (would require a stochastic
classifier). Operator picks the destination.
Runtime wiring:
ChatRuntime.attach_oov_sink(sink) mirrors attach_discovery_sink.
Runtime emits one OOVCandidate JSONL line per turn whose
grounding_source == "oov", no-op when no sink is attached.
Intent classifier is now invoked when EITHER sink is attached
(was: only discovery sink) — both downstream paths need it.
CLI:
core teaching oov-gaps [--top N] [--since YYYY-MM] [--root PATH]
[--sample-limit N] [--json]
core teaching oov-queue [--threshold N] [--include-tainted]
[--root PATH] [--since YYYY-MM] [--json]
ADR-0065 documents the full design (five-tier honesty gradient,
P2.1-P2.4 architecture). README.md updated with the ADR-0065
index entry.
Verification:
tests/test_oov_pipeline.py 24 passed
Operator workflow round-trip verified live:
> rt.attach_oov_sink(sink); rt.chat("What is photosynthesis?")
→ sink receives:
{"boundary_clean":true,"candidate_id":"f51bf8...",
"intent":"definition","token":"photosynthesis","trigger":"unresolved_subject",
"source_turn_trace":"","review_state":"unreviewed"}
> core teaching oov-gaps --root /tmp/oov_demo
→ ranked table by count, intent-set per token
> core teaching oov-queue --root /tmp/oov_demo --threshold 2
→ promoted tokens + suggested mounted packs
Full lane: 2005 passed, 2 skipped, 0 failed in 2:34 (xdist).
Full lane wall-time: 6:35 → 2:25 (2.7× speedup). No behavioral
changes; same 1933 passed, 2 skipped.
Three wins, biggest first:
1. pytest-xdist as a project dependency.
``pyproject.toml`` gains ``pytest-xdist>=3.6``. ``cmd_test``
injects ``-n auto`` for ``--suite full`` when xdist is importable;
curated suites stay single-process because worker-spawn overhead
is net-negative on the smaller suites. Operator can override
via passing ``-n <N>`` or ``--dist`` explicitly.
Verified: ``core test --suite full -q`` prints ``bringing up
nodes...`` and parallelises across the runner's CPUs.
2. Module-scoped fixture for run_demo() in test_learning_loop_demo.py.
The 7 demo tests each previously called ``run_demo(emit_json=True)``
from scratch — and ``run_demo`` itself runs the cognition lane
twice via the replay-equivalence gate. ~15s/file → ~3s/file.
Module scope (not session) is intentional: pytest-xdist
distributes by test, so a session-scoped fixture would still be
re-evaluated per worker that picks up a test from this file.
Module scope keeps the cost paid once per worker per file, which
is the actual lower bound.
3. Module-scoped fixture for the teaching-loop bench.
``test_teaching_loop_bench.py``'s 5 tests previously each ran
``run_teaching_loop_determinism(runs=2 or 3)`` — 12 pipeline
invocations across the file. One ``runs=3`` invocation shared
across all 5 tests covers every assertion: ~25s → ~7s.
For local iteration, ``core test --suite cognition -q`` etc. remain
fast (no xdist overhead). The full-lane speedup is most visible
under CI / pre-merge runs.
Closes the corpus flywheel. ADR-0055 Phase B emits DiscoveryCandidate
JSONL to the discovery sink, but until now there was no operator-facing
view: candidates accumulated to disk, no one grepped them, the system's
"I would have grounded this if I had a chain" signal went into a void.
P1.1 — Discovery aggregator (teaching/gaps.py).
Pure reader over the discovery-sink monthly-rollover layout
(<root>/<YYYY>/<YYYY-MM>.jsonl). aggregate_gaps(root, since,
sample_limit) groups emitted candidates by (subject, intent) cell
and returns a deterministic ranked tuple of Gap records.
- count: total emissions
- boundary_clean_count: subset whose boundary_clean flag held
(refusal/hedge-tainted emissions split out so operators can filter)
- sample_candidate_ids: up to N retained ids per cell, sorted
- months_seen: every month token where the cell appeared
--since YYYY-MM filters by file naming convention (no timestamp
dependency). Malformed lines silently skipped. Default root:
teaching/discovery_log.
CLI: core teaching gaps [--root PATH] [--since YYYY-MM] [--top N]
[--sample-limit N] [--json]
P1.2 — Auto-promotion queue (teaching/promotion.py).
promote_gaps(gaps, threshold, include_tainted) lifts cells whose
effective count meets the threshold into GapPromotion records.
- Default mode: boundary_clean_count gates promotion. Tainted-only
cells (count > 0 but all emissions refusal/hedge-tainted) do not
auto-promote — those may indicate the prompt hit a safety axis,
not a curriculum gap.
- include_tainted=True counts every emission (operator override).
- Threshold must be >= 1 (zero threshold defeats the queue).
- queue_id is stable + deterministic (gap:<intent>:<subject>@<N>).
- No content synthesis — promotion never invents connective or
object; only an operator can author a complete chain via the
propose/replay/accept pipeline.
CLI: core teaching queue [--threshold N] [--include-tainted]
[--root PATH] [--since YYYY-MM] [--json]
Operator workflow (closed loop):
operator → core chat # asks question
← cold turn emits DiscoveryCandidate
operator → core teaching gaps --top 10 # ranked gaps
operator → core teaching queue --threshold 3 # auto-promoted
operator → authors candidate JSONL
operator → core teaching propose <path> # replay gate runs
operator → core teaching review <id> --accept # corpus mutates
24 new tests (13 gaps + 11 promotion), all pure / no I/O dependencies,
fast (<1s combined). Full lane: 1933 passed, 2 skipped.
The full lane carried 13 long-standing red tests whose premises were
invalidated by reviewed-corpus growth that landed in earlier commits.
None reflected runtime bugs — all four classes are corpus-state drift
where the test fixture became stale. Curated lanes were green, full
lane stayed quietly red. Closes that gap.
1. test_teaching_audit (2 tests).
* test_audit_real_corpus_runs_clean asserted dropped == () and
lines_on_disk == lines_loaded — premise written before any
supersession existed. Curriculum saturation v2 (commit a0edbb4)
ratified the wisdom_grounds_judgment → wisdom_requires_knowledge
supersession; the audit now correctly shows 1 dropped line.
Rewritten as the line-conservation invariant:
lines_loaded + len(dropped) == lines_on_disk
plus a typed-reason check on every dropped entry.
* test_default_superseded_by_is_null_in_loaded_entries asserted
ALL loaded entries have superseded_by == None. Wrong even by
ADR-0055 design: the replacement entry IS loaded and carries
the back-pointer to the retired chain. Rewritten as the
active-set invariant: any non-null superseded_by on a loaded
entry must reference a dropped (retired) chain id, never a live
one — no double-live state.
2. test_learning_loop_demo (7 tests).
The demo's headline prompt was "Why does thought exist?", and the
ADR-0057 demo trilogy (commit 82dac4b) chose (thought, cause) as
the cold cell. Cognition saturation v2 (commit a0edbb4) ratified
cause_thought_reveals_meaning into the active corpus — so the
cold turn now grounds, no discovery candidate is emitted, every
demo scene breaks. Rotated the cold subject to ``narrative``
(pack-resident, no chain, same thematic shape, same affirming
evidence pointer cause_creation_reveals_meaning). Demo headline,
evals/learning_loop/run_demo.py, core/cli.py preamble, and the
test assertions all updated together so the demo reads cleanly:
before: [none] I don't know — insufficient grounding...
after : [teaching] narrative — teaching-grounded ... narrative
reveals meaning ...
3. test_discovery_candidates (4 tests).
Test fixture used (judgment, CAUSE) as the still-cold pair.
Epistemology v1 (commit 2acf71f) ratified
cause_judgment_requires_wisdom — (judgment, cause) is no longer
cold. Rotated to ``principle`` (pack-resident, no chain on either
intent today). Added a pytest.skip self-guard so when a future
curriculum unit ratifies a (principle, *) chain the test rotates
cleanly instead of going red.
Full lane: 1892 passed, 2 skipped, 0 failed (was 4 failed pre-fix,
13 failed pre-ADR-0063). Cognition eval unchanged: public 100/100/
91.7/100, holdout 100/100/83.3/100.
ADR-0063 closes the ADR-0048/0050/0053/0061 hardcoded-cognition-pack
asymmetry. New chat/pack_resolver.py provides resolve_lemma(lemma,
pack_ids) → (resolving_pack_id, semantic_domains) across an ordered
tuple of mounted lexicon packs (first-match-wins, lru_cache per-pack).
Surface composers in chat/pack_grounding.py now consult the resolver
instead of a hardcoded en_core_cognition_v1. en_core_relations_v1
joins RuntimeConfig.input_packs defaults; kinship lemmas now ground
on the live path:
> What is a parent?
parent — pack-grounded (en_core_relations_v1):
kinship.ascendant.direct; kinship.parent; biology.progenitor.
No session evidence yet.
Cross-pack comparison (knowledge × parent) renders composite tag
(en_core_cognition_v1 × en_core_relations_v1). Cognition lane
remains byte-identical: cognition is resolved first and the surface
format for cognition lemmas is unchanged.
Cognition eval (byte-identical to pre-ADR baseline):
public → intent 100% / surface 100% / term 91.7% / closure 100%
holdout → intent 100% / surface 100% / term 83.3% / closure 100%
Curated lanes green: smoke 67 / cognition 121 / teaching 17 /
packs 6 / runtime 19 / algebra 132.
New tests: test_pack_resolver.py (28) + test_cross_pack_grounding.py
(17). test_en_core_relations_v1_pack.py: default-input-packs guard
inverted. test_pack_grounding.py: two stale ADR-0048 tests rewritten
(premises invalidated by ADR-0052/0061; now use fully-out-of-pack
prompts).
chat/teaching_grounding.py UNCHANGED — cognition_chains_v1 corpus
stays cognition-only. Cross-pack teaching corpora are the natural
ADR-0064.
Pre-ADR-0062, the teaching-grounded composer emitted exactly one
reviewed chain per surface — "light reveals truth" — even when the
corpus already contained an immediate follow-up "truth grounds
knowledge". With 21 active chains after curriculum saturation v2,
many grounded prompts had a corpus-ratified follow-up the composer
silently dropped.
ADR-0062 adds the composed composer + an opt-in config flag:
flag OFF (default):
light — teaching-grounded (cognition_chains_v1): cognition.illumination;
logos.core. light reveals truth (cognition.truth). No session evidence yet.
flag ON:
light — teaching-grounded (cognition_chains_v1): cognition.illumination;
logos.core. light reveals truth (cognition.truth), which grounds
knowledge (cognition.knowledge). No session evidence yet.
Follow-up resolution:
- prefer cause; fall back to verification (deterministic preference)
- cycle guard: 1-step cycles (A→B, B→A) blocked
- pack-residency guard: follow-up's object must be pack-resident
- bounded depth: v1 follows exactly one hop
- degrades to single-chain BYTE-IDENTICALLY when no follow-up
survives the guards (drop-in replacement)
Trust-boundary invariants preserved:
- Every visible non-template token is lemma / pack-domain /
humanize_predicate connective / template constant. Only added
template constant: ", which "
- Deterministic: same chains → same surface bytes
- Default-False flag pattern mirrors ADR-0047/0058
- `versor_condition < 1e-6` invariant untouched (surface composition only)
Cognition lane null-drop invariant CI-pinned:
Composed mode emits a strictly LONGER surface (extra follow-up
clause); every expected_term passing flag-OFF must still pass flag-ON.
Asserted in test_cognition_lane_metrics_unchanged_with_composed_flag
for both public and holdout splits. If a future change drops tokens,
the test fails as a deliberate regression.
public flag OFF: intent 100% / surface 100% / term 91.7% / versor 100%
public flag ON : intent 100% / surface 100% / term 91.7% / versor 100% (identical)
holdout flag OFF: intent 100% / surface 100% / term 83.3% / versor 100%
holdout flag ON : intent 100% / surface 100% / term 83.3% / versor 100% (identical)
Live-prompt lift visible on ~12 of 21 active chains; the rest hit
cycle or pack-residency guards. Saturation v2's clusters were
authored partly with composition in mind (thought→meaning→
understanding, inference→evidence→knowledge, etc.).
- core/config.py — `RuntimeConfig.composed_surface: bool = False`
- chat/teaching_grounding.py — `teaching_grounded_surface_composed`
sibling to `teaching_grounded_surface`
- chat/runtime.py — dispatch branch in `_maybe_pack_grounded_surface`
selects composed vs single-chain based on config flag
- tests/test_composed_surface.py — 11 tests pin: function-level
(None on no chain / degrades when no follow-up / two-clause when
follow-up exists / includes intermediate + final domains /
deterministic / cycle guard / trust label preserved); runtime
integration (default single-chain / flag-on composed / frozen
config); cognition-lane null-drop invariant.
Lanes (regression): smoke 67 / cognition 121 / teaching 17 /
composed-surface 11 — all green.
Three external-facing demos / benchmarks now match the existing
audit-tour / pack-measurements / long-context-comparison treatment:
preamble printed before the run, README index entries, claims table.
- core/cli.py — _ANTI_REGRESSION_PREAMBLE, _LEARNING_LOOP_PREAMBLE,
_TEACHING_LOOP_BENCH_PREAMBLE. Each lists reference ADRs, what to
expect, trust boundary, test gate, and machine-readable invocation.
Wired through _print_preamble in the demo dispatch + bench dispatch
(suppressed under --json).
- README.md — new "Inter-Session Memory — Reviewed Learning" section
between Teaching Order and Architecture: the three-gate trust
property table, the three live-demo table, and the operator-surface
command list. Quick-start block lists `core demo anti-regression`,
`core demo learning-loop`, and `core bench --suite teaching-loop
--runs 100` alongside the existing demos.
No code paths changed — preambles are stdout-only when not under JSON.
Tests unchanged; 17/17 green (5 anti-regression + 7 learning-loop + 5 bench).
`core bench --suite teaching-loop [--runs N]` runs the full reviewed-
corpus extension pipeline (propose → real replay-equivalence gate →
operator accept) N times against an identical input and asserts
byte-identical artifacts every run:
- proposal_id (SHA-256 of canonical-JSON payload)
- replay_baseline (cognition lane metrics on active corpus)
- replay_candidate (cognition lane metrics on transient corpus)
- regressed_metrics (sorted tuple)
- chain_id_written
Also reports per-iteration latency (mean / p50 / p95) and total wall.
100-run result against today's main:
unique(proposal_id)=1 unique(baseline)=1 unique(candidate)=1
unique(chain_id)=1 active_corpus_byte_eq=True
mean=1.849s p50=1.838s p95=1.851s
The full learning loop is replayable bit-identically across N
independent invocations. Pairs naturally with ADR-0045's 100% exact-
NIAH recall numbers — same epistemic class of guarantee, applied to
the *learning loop* itself rather than only to retrieval. No LLM
provider can publish equivalent numbers on a learning path.
- benchmarks/teaching_loop.py — `run_teaching_loop_determinism(runs)`
returns a typed `TeachingLoopBenchReport` with uniqueness counts,
determinism flag, byte-identical-active-corpus flag, and latency
distribution (mean / p50 / p95 / total). Pure-stdlib percentile —
no numpy dep on this path.
- benchmarks/run_benchmarks.py — `bench_teaching_loop_determinism`
shim + `_SUITES["teaching-loop"]` registration + runs= passthrough.
- core/cli.py — `--suite teaching-loop` choice added to bench parser.
- tests/test_teaching_loop_bench.py — 5 tests pin determinism at
small N, proposal_id SHA-256 shape, canonical chain_id layout,
latency stats well-formedness, JSON serialisation.
Trust boundary: every write is confined to a tempdir created inside
the bench loop; the active corpus is read once at start, once at end,
and any byte difference would fail the bench.
`core demo learning-loop` (+ `--json`) walks a single prompt through the
full ADR-0055..0057 inter-session-memory architecture:
S1. Cold turn → universal disclosure, grounding_source=none
S2. Discovery emission → DiscoveryCandidate to attached sink
S3. Operator proposal → real replay-equivalence gate, no regression
S4. Operator accept → TRANSIENT corpus only; active untouched
S5. Same prompt → teaching-grounded surface with the new chain
Before / after on the deterministic prompt "Why does thought exist?":
before: [none] I don't know — insufficient grounding for that yet.
after: [teaching] thought — teaching-grounded (cognition_chains_v1):
cognition.thought; logos.internal. thought reveals meaning
(cognition.meaning). No session evidence yet.
The active corpus on disk is byte-identical pre/post. The demo writes
only to a transient corpus, then swaps `_CORPUS_PATH` for the after
turn — the same pattern the replay-equivalence gate uses.
- evals/learning_loop/run_demo.py — `run_demo(emit_json=False)` returns
a structured `DemoReport` with both surfaces and per-scene detail.
- core/cli.py — `core demo learning-loop` target wired.
- tests/test_learning_loop_demo.py — 7 tests pin: full loop closes,
before is ungrounded, after contains new chain atoms (thought /
reveal / meaning), discovery emits ≥1, replay gate reports no
regression, S4 byte-identical active + 1 line on transient, same
prompt drives both surfaces.
Lane state: learning-loop-demo 7 new — green. Demo runs in ~15s
end-to-end (cognition lane runs twice via replay gate).
No LLM provider has a published equivalent of this loop: per-fact
provenance from operator accept to surface, replay-equivalence gate
proving non-regression, byte-identical active state regardless of
outcome, full audit trail back to the originating cold turn.
`core demo anti-regression` (+ `--json`) is a self-contained walkthrough of
the three independent gates that every reviewed-corpus extension must pass.
Designed for showcasing CORE's epistemic discipline to reviewers / industry
observers — no LLM provider has a published equivalent.
Scenes:
- S1. Eligibility predicate refuses an undetermined-polarity candidate
before any replay is invoked. ProposalError raised; no log row.
- S2. Replay-equivalence gate auto-rejects a regressing candidate with
the named regressed metrics in the operator note. Uses the documented
`run_replay=` kwarg of `propose_from_candidate` to inject a controlled
regression of the same `ReplayEvidence` shape the real gate produces.
- S3. Real `teaching.replay.run_replay_equivalence` runs the cognition
public lane. A replay-equivalent candidate reaches 'pending' — operator
`--accept` is still required to write.
Each scene asserts the active corpus is byte-identical pre/post.
- evals/anti_regression/run_demo.py — `run_demo(emit_json=False)` returns
a structured `DemoReport`; verbose human output by default, JSON on flag.
- core/cli.py — `core demo anti-regression` target wired alongside
audit-tour / pack-measurements / long-context-comparison.
- tests/test_anti_regression_demo.py — 5 tests pin each scene's
load-bearing claim + the corpus-byte-identical invariant.
Lane state: anti-regression-demo 5 new — green. Demo runs in ~10s end-to-end.
`core teaching supersessions` (+ `--json`) pairs each retired chain with its
active replacement. Derived view over `audit_corpus()`; pure, read-only.
- teaching/audit.py — `SupersessionRecord` + `supersession_history(report)`
returns retired→replacement pairs ordered by retired-line (disk order,
oldest first). Orphan supersessions (retired with no live entry carrying
the matching `superseded_by` — e.g. chained retirements where the middle
link itself was retired) surface as `replacement=None` so silent corpus
drift is inspectable.
- core/cli.py — `core teaching supersessions [--json]`. Exit 1 if any
orphan is detected (catches silent drift in CI); 0 otherwise.
- tests/test_supersession_history.py — 7 tests pin empty-history,
single-pair shape, chained-supersession surfaces both pairs, line-no
ordering, orphan detection, JSON round-trip, no corpus mutation.
Lane state: smoke 67 / cognition 121 / supersession-history 7 new / supersede 13 /
audit 23 — green. `core eval cognition`: unchanged (intent 100% / surface 100% /
term 91.7% / versor 100%). Real corpus today reports `(no supersessions)`.
`core teaching supersede <old_chain_id> --subject ... --intent ... --connective ...
--object ... --review-date YYYY-MM-DD` is the second corpus mutation surface
(alongside accept_proposal). No replay gate — it's a deliberate operator action
that replaces a hand-authored or previously discovery-promoted chain.
- teaching/supersede.py — `supersede_chain()` orchestrator with pre-checks
(review_date format, intent whitelist, pack-consistency via re-audit,
no double-supersede, no self-supersede, no new-chain-id collision) and
byte-identical rollback on post-audit failure.
- teaching/proposals.py — extended `append_chain_to_corpus` with optional
`superseded_by` kwarg; remains the only function in the codebase that
writes to the active teaching corpus.
- core/cli.py — `core teaching supersede` subcommand wired to the live
`_CORPUS_PATH`; EPILOG updated with example.
- tests/test_supersede.py — 13 tests pin every gate, byte-identical
rollback on rejection, append-only at disk level, audit-and-runtime
parity after supersession, hand_authored provenance with
`supersede(<old_chain_id>)` tag.
Lane state: smoke 67 / cognition 121 / teaching 17 / supersede 13 / audit 23 /
proposals 16 / contemplation 16 / contemplation-wiring 6 / discovery 24 — green.
`core eval cognition`: intent 100% / surface 100% / term 91.7% / versor 100% — unchanged.
The only path by which CORE extends its own active teaching corpus.
Closes ADR-0055 Phase C alongside ADR-0056's cognitive surface.
Three load-bearing calls (recorded in ADR-0057):
1. Replay-equivalence is a precondition, not a permission;
operator --accept remains required.
2. Eligibility = polarity in {affirms, falsifies} AND at least
one source='corpus' evidence pointer AND boundary_clean AND
claim_domain != evaluative (unless --allow-evaluative) AND
proposed_chain complete.
3. Append-only proposal log; corpus history append-only too.
Changes
- teaching/proposals.py — TeachingChainProposal, ReplayEvidence,
ProposalLog (event-sourced replay → current_state), eligibility
predicate, propose_from_candidate, accept/reject/withdraw,
append_chain_to_corpus (the sole corpus-write surface). Uses
TYPE_CHECKING guards to break the circular import with
chat.pack_grounding.
- teaching/replay.py — run_replay_equivalence; swaps _corpus_index
path to a tmp file, runs cognition lane on the active corpus
AND a transient copy with the proposed chain appended, returns
regressed-metrics list; trust-boundary assertion that the active
corpus bytes are byte-identical pre/post.
- teaching/discovery.py — moved chat.pack_grounding /
chat.teaching_grounding imports inside extract_discovery_candidates
to break the cycle (was masked when chat.runtime was the entry
point; surfaced by CLI entry).
- core/cli.py — three new subcommands:
core teaching propose <candidate-jsonl-path> [--allow-evaluative]
core teaching proposals [--state pending|accepted|rejected|withdrawn] [--json]
core teaching review <proposal_id> --accept --review-date YYYY-MM-DD
core teaching review <proposal_id> --reject [--note ...]
core teaching review <proposal_id> --withdraw [--note ...]
- tests/test_teaching_proposals.py — 16 tests covering: every
eligibility gate, proposal_id idempotency, append-only log,
replay-equivalent stays pending, regression auto-rejects with
named regressed metrics, --accept appends one line with typed
Provenance, --accept refused on non-equivalent, state-machine
blocks double-accept, real replay gate runs cognition lane
twice and asserts byte-clean active corpus pre/post.
Invariants preserved
- versor_condition(F) < 1e-6 — C2 touches no algebra path.
- Active corpus bytes byte-identical regardless of replay outcome.
- No clock-time reads, no LLM, no async.
- Proposal-only — accept_proposal is the sole corpus-write path.
Lanes: smoke 67 / cognition 121 / runtime 19 / teaching 17 /
new proposals 16. Cognition eval unchanged.
Open follow-ups (not in scope):
- supersession via operator review action
- cross-pack falsification arbitration (ADR-0056 Call 2 deferred)
- pack-data migration of frame-dependent connectives
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Lands the three load-bearing pieces of ADR-0055 Phase A so later
phases (DiscoveryCandidate, TeachingChainProposal) have a safe
substrate to write into.
- teaching/audit.py: pure, deterministic re-parse of the reviewed
corpus with same gates as the runtime loader but keeps drop
reasons (invalid_json, missing_required_field:*, unsupported_intent,
pack_missing_subject, pack_missing_object, superseded_by:*).
- teaching/provenance.py: typed Provenance(adr_id, source,
review_date, raw); legacy "reviewed" maps to "hand_authored" so
current corpus reports the canonical enum without a file rewrite.
- chat/teaching_grounding._corpus_index honors superseded_by —
active view drops superseded entries while disk preserves history.
- core teaching audit CLI subcommand (--json optional); exits 1 on
any drop so CI catches silent corpus shrinkage from pack swaps.
Observable behaviour unchanged: corpus is 10/10 loaded, all five
core lanes green (smoke 67, cognition 121, runtime 19, teaching 17,
packs 6), cognition eval metrics identical on dev / public /
holdout splits. versor_condition < 1e-6 invariant untouched.
Tests: tests/test_teaching_audit.py — 23 tests covering provenance
parser, real-corpus determinism, every drop-reason path,
supersession semantics, runtime/audit parity, read-only contract.