Final phase of the articulation arc. Consumes the per-turn
``PlanMetrics`` + ``ContemplationFinding`` streams produced by
Phases 3 + 4 and aggregates across many turns to emit
SPECULATIVE ``PACK_MUTATION_CANDIDATE`` findings that the operator
reviews via the existing proposal-review-ratify chain.
This is the doctrine-aligned answer to the user's question:
"Should we... realize a way to score whether it should use what
it produced towards memory confidence for future use?"
Yes — and it stays inside ADR-0080: read-only, SPECULATIVE-only,
deterministic, no parallel learning path, no autonomous memory
mutation.
What it adds
------------
* New module ``chat/articulation_telemetry.py``:
- ``ArticulationObservation`` frozen dataclass — per-turn
bundle of (turn_id, anchor_subject, prompt_hash,
plan_substrate_hash, metrics, findings).
- ``format_articulation_observation_jsonl(...)`` — deterministic
sort-keys JSONL line.
- ``load_articulation_observations(lines)`` — schema-tolerant
loader; malformed lines drop without aborting.
- ``ArticulationObservationSink`` protocol — structurally
identical to ``TurnEventSink`` but distinct named type so
consumers can subscribe to one stream without the other.
* New module ``core/contemplation/miners/articulation_quality.py``:
- ``mine_articulation_observations(observations, paths)`` —
pure deterministic aggregator with three v1 rules.
- **recurring_predicate_monotony** — when the same
(subject, predicate) pair is flagged WEAK_SURFACE in
>= _MIN_RECURRENCE (default 3) observations, propose
substrate diversification with non-dominant predicates.
- **recurring_planner_gap** — when the same subject is
flagged PLANNER_GAP >= _MIN_RECURRENCE times across modes,
propose substrate expansion.
- **low_average_predicate_diversity** — when mean
``predicate_diversity_ratio`` < 0.5 across >= _MIN_RECURRENCE
observations on the same anchor subject, propose
diversification.
* Runtime wiring (``chat/runtime.py``):
- New ``ChatRuntime.attach_articulation_sink(sink)`` method.
Mirrors ``attach_telemetry_sink`` pattern.
- Emission point at the end of
``_maybe_apply_discourse_planner``: when contemplation
enabled + sink attached + plan engaged, builds an
``ArticulationObservation`` and emits one JSONL line.
Sink errors propagate (fail-fast, no swallowing).
- Per-runtime ``_articulation_turn_counter`` increments on
every emission; gives downstream consumers a stable
sequence index.
Tests
-----
* ``tests/test_articulation_quality_miner.py`` (11 tests):
- Empty / sub-threshold cases yield no findings.
- Each of the three rules fires at threshold.
- Recurring_predicate_monotony separates by subject (no
cross-subject merging).
- Recurring_planner_gap collects distinct modes into a
sorted comma-joined string.
- Determinism — byte-equal finding IDs across two runs.
- SPECULATIVE doctrine pin.
- JSONL round-trip preserves observation identity.
* ``tests/test_articulation_quality_e2e.py`` (7 tests):
- Sink-detached + contemplation-on → no emission.
- Sink-attached + contemplation-off → no emission.
- Engaged turn emits exactly one observation line.
- BRIEF prompt emits nothing (fast-path).
- **Full loop** — run compound prompt 3x → 3 observations →
miner emits PACK_MUTATION_CANDIDATE with subject='truth',
predicate='recurring_predicate_monotony', object='belongs_to'.
- Full loop is deterministic (byte-equal finding IDs across
two complete runs).
- Every full-loop finding is SPECULATIVE.
Doctrine pins
-------------
| Claim | Pinned by |
|--------------------------------------|----------------------------------------------------------|
| SPECULATIVE-only | test_all_findings_remain_speculative |
| Deterministic across runs | test_miner_is_deterministic_across_runs |
| Full-loop determinism (e2e) | test_full_loop_is_deterministic_byte_equal_finding_ids |
| No autonomous mutation | Sink is append-only; miner outputs ContemplationFinding |
| | objects only; nothing writes to packs/vault/teaching. |
| Append-only stream | Sink protocol has emit(line: str) and nothing else. |
Live demo (3 identical compound-prompt turns)
---------------------------------------------
Runtime emits 3 observations. Offline miner aggregates and emits:
[pack_mutation_candidate] subject='truth'
predicate='recurring_predicate_monotony' object='belongs_to'
evidence_refs: 3 observations
proposed_action: "diversify substrate for 'truth': across 3
observations the plan repeatedly over-concentrated on
predicate 'belongs_to'. Candidates: add teaching chains
rooted on 'truth' with relations OTHER than 'belongs_to'
(grounds / requires / reveals / contrasts / precedes /
follows) so the planner's RELATION selector has more
variety to draw from."
epistemic_status: speculative
The system observed its own articulation patterns across many
turns, identified the corpus expansion priority, and emitted a
specific reviewable proposal — without mutating anything. The
operator decides whether to act on it via the existing review
chain.
Verification
------------
pytest test_articulation_quality_miner.py 11/11 pass
pytest test_articulation_quality_e2e.py 7/7 pass
pytest test_plan_metrics*.py 18/18 pass (Phase 4)
pytest test_plan_contemplation*.py 17/17 pass (Phase 3)
pytest test_discourse_planner_*.py 99/99 pass
pytest test_articulation_demo.py all claims supported
pytest test_narrative_example_intents.py pass
core test --suite smoke 67/67 pass
core test --suite runtime 19/19 pass
The articulation arc is complete. Future work documented in
``docs/sessions/SESSION-2026-05-21-articulation-arc.md`` §8:
connective rotation, generalised pronoun selection, doctrine-gated
plan revision, Phase 2.5 mid-sentence reflection. None blocking.
Quantitative companion to Phase 3 (commit 664e081). Where Phase 3
emits SPECULATIVE *findings* about plan quality, Phase 4 emits
typed *measurements* — pure-function projection of a
``DiscoursePlan`` into a ``PlanMetrics`` dataclass.
Why this matters
----------------
The discourse planner now produces multi-clause grounded
articulations (Phase 1), the renderer pronominalizes across
consecutive same-subject moves (Phase 2), and the contemplation
pre-flight emits qualitative concerns about plan shape (Phase 3).
What was missing was the *aggregable* layer: per-turn structured
numbers that downstream consumers can stream across many turns
to score quality patterns the per-turn observer cannot see.
Phase 4 lands that layer. Phase 5 (offline contemplation miner)
becomes possible because there's now structured signal to mine.
What it measures
----------------
Structure
* move_count — total moves in plan
* fact_bearing_count — moves with fact != None
Move-kind distribution
* anchor_count / support_count / relation_count
/ transition_count / closure_count
Diversity
* unique_predicates — distinct predicates across
fact-bearing moves
* unique_subjects — distinct subject lemmas
* unique_sources — distinct FactSources
Topic dynamics
* topic_shift_count — consecutive pairs where
subject changed
* pronominalization_opportunities — consecutive pairs where
subject held (= Phase 2's
anaphora trigger count)
Derived ratios
* predicate_diversity_ratio — unique_predicates /
fact_bearing_count
* subject_focus_ratio — pronominalizations /
(pronominalizations +
topic_shifts)
Every field is a deterministic pure function of the plan: same
plan in → byte-equal ``PlanMetrics.as_dict()`` out. This is the
load-bearing claim that lets Phase 5 aggregate across turns
without "is this the same metric?" ambiguity.
Doctrine alignment
------------------
Per ADR-0080 contemplation discipline:
* Read-only — metrics are pure projections of the plan; no
mutation of plan, runtime state, or memory tiers.
* No autonomous learning — metrics are observations, not
learned policy. Promotion to memory still flows through
the existing proposal-review-ratify chain.
* Deterministic replay — pinned by test_metrics_are_deterministic_
and_byte_equal_as_dict plus the runtime-level
test_metrics_byte_equal_across_runs.
Wiring
------
* New ``ChatRuntime.last_plan_metrics`` property — read-only
``PlanMetrics`` from the most recent turn where the planner
engaged (and ``discourse_contemplation`` was on); ``None``
otherwise. Reset between turns alongside ``last_plan_findings``
via the existing top-of-call reset block.
* Same opt-in flag as Phase 3 (``discourse_contemplation``).
When True, the runtime computes both findings AND metrics in
the same block; when False (default), both stay at empty/None.
Demo (config: discourse_contemplation=True)
-------------------------------------------
"What is knowledge?" → metrics: None (BRIEF fast-path)
"Tell me about memory." → moves=3 fact_bearing=3
kinds=A:1/S:1/R:1/T:0/C:0
unique_predicates=3 subjects=1
pronominalization_ops=2 shifts=0
predicate_diversity=1.000
subject_focus=1.000
"What is truth, and why does
it matter?" → moves=7 fact_bearing=6
kinds=A:2/S:2/R:2/T:1/C:0
unique_predicates=4 subjects=1
pronominalization_ops=4 shifts=1
predicate_diversity=0.667 ← Phase 3
WEAK_SURFACE
quantified
subject_focus=0.800
+ 1 finding (weak_surface)
The compound-prompt numbers are particularly informative:
``predicate_diversity=0.667`` is the algebraic expression of the
Phase 3 ``WEAK_SURFACE`` rule — the rule fires precisely because
6 fact-bearing moves used only 4 distinct predicates.
``subject_focus=0.800`` quantifies that 80% of consecutive pairs
held the same subject — high topic stickiness that Phase 2's
reflective renderer leveraged into 4 ``it`` substitutions.
Tests
-----
* ``tests/test_plan_metrics.py`` — 10 unit tests pinning each
field, derived ratios, bridge-move handling (``fact=None``
resets the focus channel), and determinism via ``as_dict()``
byte-equality.
* ``tests/test_plan_metrics_runtime.py`` — 8 end-to-end tests
proving the runtime wiring: disabled by default, populated
when enabled, BRIEF prompts yield None, no cross-turn leak,
byte-equal across runs, parametrized co-population check
alongside findings.
Verification
------------
pytest tests/test_plan_metrics*.py 18/18 pass
pytest tests/test_plan_contemplation*.py 17/17 pass (Phase 3)
pytest tests/test_discourse_planner_*.py 99/99 pass
pytest tests/test_articulation_demo.py all claims supported
pytest tests/test_narrative_example_intents.py pass
pytest tests/test_runtime_config.py pass
cognition eval OFF vs ON 45/45 surface byte-equal
45/45 trace_hash byte-equal
4/4 aggregate metrics
identical
core test --suite smoke 67/67 pass
core test --suite runtime 19/19 pass
Phase 5 (logged, not built)
---------------------------
Offline contemplation miner that consumes ``last_plan_findings``
+ ``last_plan_metrics`` streams across many turns and emits
reviewable pack-mutation candidates. Still SPECULATIVE;
review-gated; never auto-promoted to memory. Now unblocked by
the structured metric surface Phase 4 lands.
Wires deterministic, read-only contemplation OVER a completed
``DiscoursePlan`` BEFORE the renderer fires. This is the
"reasoning at meaningful checkpoints" capability — the system
now inspects the global shape of its own articulation plan and
emits SPECULATIVE findings about quality issues the move-by-move
planner couldn't see locally.
Doctrine alignment (ADR-0080)
-----------------------------
* **Read-only** — never mutates the plan, packs, vault, teaching
corpus, or runtime state. Returns findings as a tuple; the
runtime stores them on a read-only property.
* **SPECULATIVE-only** — every finding is stamped
``EpistemicStatus.SPECULATIVE`` by the schema's ``__post_init__``;
the doctrine pin ``test_findings_always_speculative`` keeps that
invariant visible.
* **Deterministic replay** — same plan → byte-identical findings
(same ``substrate_hash``, same ``finding_id``).
* **No parallel learning path** — findings flow to a read-only
observation surface (``runtime.last_plan_findings``). Promotion
to memory still goes through the existing proposal → review →
ratify chain. The offline contemplation miner (Phase 5 target)
is what eventually consumes the findings and emits reviewable
pack-mutation candidates.
v1 rules (``core/contemplation/plan_preflight.py``)
----------------------------------------------------
* ``PLANNER_GAP`` — non-BRIEF mode produced anchor-only depth.
Signals the teaching/cross-pack substrate for that lemma is too
thin for the planner to expand.
* ``WEAK_SURFACE`` — three or more moves share a predicate.
Signals the rendered surface will read mechanical (e.g. three
``belongs_to`` clauses in a row). Fires on today's compound
prompt ``"What is truth, and why does it matter?"`` — the
6-sentence plan uses ``belongs_to`` 3 times.
* ``COVERAGE_GAP`` — every move in a multi-move plan draws from
a single ``FactSource``. Signals one-sided substrate (e.g.
pack-only with no teaching enrichment).
Runtime wiring
--------------
* New ``RuntimeConfig.discourse_contemplation: bool = False`` —
opt-in for now. Default off keeps the cognition eval byte-
identical to Phase 2 (verified 45/45 surface + 45/45 trace_hash).
* New ``ChatRuntime.last_plan_findings`` property — read-only tuple
of ``ContemplationFinding`` records from the most recent turn.
Reset to ``()`` at the start of every plan-engagement call so
findings never leak across turns.
* Contemplation runs AFTER the planner produces a multi-move plan
and BEFORE the renderer fires; the plan itself is not modified.
Demo (config: discourse_contemplation=True)
-------------------------------------------
"What is knowledge?" → planner fast-path; no findings
"Tell me about memory." → 3 moves, distinct predicates;
no findings (good!)
"What is truth, and why does
it matter?" → 6 moves, ``belongs_to`` x 3:
[WEAK_SURFACE] subject='truth'
predicate='predicate_repeats_in_plan'
object='belongs_to'
proposed action: diversify the
relation inventory for 'truth'
(grounds / requires / reveals /
contrasts) so the planner has
more variety to draw from.
"Explain truth." → 3 moves, distinct predicates;
no findings
Tests
-----
* ``tests/test_plan_contemplation.py`` — 11 unit tests pinning
each rule, empty/trivial plans, determinism, and the
SPECULATIVE-only doctrine.
* ``tests/test_plan_contemplation_runtime.py`` — 6 end-to-end
tests proving the runtime wiring: disabled by default,
populated when enabled, reset across turns, deterministic
across runs, all findings SPECULATIVE.
Verification
------------
pytest tests/test_plan_contemplation*.py 17/17 pass
pytest tests/test_discourse_planner_*.py 99/99 pass
pytest tests/test_articulation_demo.py all claims supported
pytest tests/test_narrative_example_intents.py pass
pytest tests/test_runtime_config.py pass
cognition eval OFF vs ON 45/45 surface byte-equal
45/45 trace_hash byte-equal
4/4 aggregate metrics
identical
core test --suite smoke 67/67 pass
core test --suite runtime 19/19 pass
Phases roadmap (logged in commit, not built today)
--------------------------------------------------
* Phase 4 — articulation telemetry enrichment. Emit per-turn
metrics (grounding_ratio, anaphora_engagement, plan_completeness,
novelty, focus_consistency) to the existing telemetry sink so
the offline miner has structured signal.
* Phase 5 — offline contemplation miner. Extend
``core/contemplation`` with a miner that consumes
``last_plan_findings`` streams and emits reviewable
pack-mutation / teaching-corpus expansion proposals. Still
SPECULATIVE; review-gated.
* chore(evals, cli): contract standardization + bench --json stdout cleanliness
End-of-session shippability pass. Three concrete fixes:
1. core/cli.py — bench --json no longer pollutes stdout
Several bench paths call scripts.run_pulse.run_pulse which prints
verbose [pulse] traces unconditionally to stdout, breaking jq /
programmatic consumers of --json output.
New _bench_stdout_guard() redirects stdout → stderr for the
duration of the bench run when --json is set. Operator still sees
the pulse trace (on stderr), but --json consumers get a clean JSON
document on stdout. Applied to all four bench paths: cost,
articulation, default suite, and --suite all.
Verified: core bench --suite determinism --json now produces
parseable JSON; human path still shows 1140 [pulse] lines.
2. evals/{frontier_compare,realizer_guard}/contract.md (new)
core/contemplation/contract.md (new)
Each new contract follows the established pattern (37 contracts
already exist under evals/<lane>/contract.md):
- What it measures
- Why it matters (structural win)
- How to run
- How to read the output
- Pass criteria table
- When it has failed and why
- Runner / module layout
Coverage:
- frontier_compare: both Lane A (CORE-only suites) and Lane B
(cross-provider prompt_battery) with explicit guardrails
against mixing — operator asks for the wrong lane combination,
runner exits 2 with helpful error.
- realizer_guard: C1/C2 articulation safety boundary — synthetic
illegal candidates rejected directly by check_surface AND
former-bug runtime prompts now produce legal articulations.
- contemplation (ADR-0080): not under evals/ since it's runtime
infrastructure that consumes eval reports — contract lives at
core/contemplation/contract.md. Documents the read-only +
SPECULATIVE-only + deterministic-replay invariants and the
shared DiscoveryCandidateSink plumbing convergence (ADR-0080).
3. evals/CLAIMS.md — Tier 2 rows added
- frontier_compare Lane A: determinism.primary_score, max_versor_condition
- frontier_compare Lane B: prompt_battery.primary_score (CORE adapter),
cross-provider artifact persistence
- realizer_guard: all_claims_supported
- contemplation: SPECULATIVE-only invariant, deterministic replay,
additive sink path, no pack mutation (all CI-pinned by tests)
Verification
------------
$ core test --suite smoke -q
67 passed in 27.22s (no regression)
$ uv run pytest -q tests/test_contemplation_loop.py \
tests/test_contemplation_pipeline_convergence.py \
tests/test_frontier_compare_cross_provider.py
27 passed in 4.87s
$ core bench --suite determinism --json 2>/dev/null | jq .results[0].passed
true (was: JSONDecodeError on prior [pulse] pollution)
* feat(evals/ui): report viewer renders Lane B cross-provider + pass-rate chart
Stop-hook caught that #62 only covered contracts — the 929-line
report_viewer.html was never audited against the new cross-provider
report shape from #61. Two real gaps:
1. Lane-aware observation drawer
The drawer hardcoded Lane A (CORE-native) fields: surface,
grounding_source, anchor_lens_mode_label, versor_condition.
Lane B (cross-provider) observations carry different fields:
provider, model, elapsed_ms, error_type, error_message.
Loading a cross-provider report rendered only the surface row
with empty `grounding` — the provider + model + timing data
was unreachable without expanding "Show raw JSON".
Fix: detect Lane B (presence of `obs.provider`) and render the
appropriate field set. Lane A still renders identically (now
also surfaces trace_hash + register_id when present, which were
silently buried in the raw JSON before).
2. Pass-rate chart per suite
The summary strip showed one aggregate Primary % across all
suites, with no way to see WHICH suite is dragging the score.
Multi-suite runs (e.g. --suite all) had to expand each panel
individually to find the failing one.
Fix: new .passrate-chart element below the summary strip,
one horizontal bar per suite showing passed/total. All-pass =
solid green, all-fail = solid red, partial = green/red split
at the pass fraction. CSS only — no new dependencies.
3. SUITE_PREAMBLES gains the prompt_battery entry so the sidebar
shows the "side-by-side surface evidence across providers"
description when loading a Lane B report.
Verified
--------
- Brace/paren/div balance unchanged (308/308 / 380/380 / 54/54)
- One <script> tag pair preserved
- Generated a real Lane B report via
`python -m evals.frontier_compare --provider core --suite prompt_battery`
for visual confirmation
Out of scope (noted for future PR)
----------------------------------
Sampled 3 `core demo` targets:
- register-tour: clean schema (all_claims_supported, claims, grid)
- audit-tour: both scene_1_* keys AND an empty scenes:[] array — inconsistent
- anti-regression: no all_claims_supported key, uses all_gates_held instead
Demo schema standardization deserves its own PR — operator tooling
would benefit from a uniform top-level success field across demos.
* docs(evals) + chore(demos): systematic audit + uniform success field
Stop-hook caught two real gaps after the contract+UI PR:
- demos had divergent success-field names (all_gates_held vs
learning_loop_closed vs claim_supported vs nested claims_supported)
- no systematic look at the 48 eval directories had been done
Both addressed concretely; remaining work captured in audit doc
rather than vaguely deferred.
1. Demo schema standardization — uniform all_claims_supported field
----------------------------------------------------------------------
All 9 ``core demo`` targets now emit a top-level
``all_claims_supported: bool`` field. Existing per-demo fields
(``all_gates_held``, ``learning_loop_closed``, ``claim_supported``,
nested ``claims_supported``) are preserved for backwards compat —
the new field is an alias derived from the demo's existing success
signal, not a replacement.
Operator tooling and the CI gate can now target
``all_claims_supported`` without knowing each demo's idiomatic
field name.
Files touched:
- evals/anti_regression/run_demo.py — adds AND of all_gates_held +
active_corpus_byte_identical
- evals/learning_loop/run_demo.py — adds AND of learning_loop_closed +
active_corpus_byte_identical
- scripts/publish_pack_measurements.py — adds AND of the three
entries in the nested claims_supported dict
- evals/long_context_cost/comparison_runner.py — adds alias for
claim_supported (singular)
The 5 demos already using ``all_claims_supported`` (audit-tour,
register-tour, anchor-lens-tour, orthogonality-tour, articulation)
are unchanged.
Verified across all 9 demos:
audit-tour : True
register-tour : True
anchor-lens-tour : True
orthogonality-tour : True
pack-measurements : True ← new alias
anti-regression : True ← new alias
learning-loop : True ← new alias
articulation : True
long-context-comparison : True ← new alias
2. docs/EVAL_AUDIT_2026-05-20.md — systematic 48-lane audit
------------------------------------------------------------
Replaces the "future PR" deferral with a concrete document.
Contains:
- Method (what was inspected for each lane).
- Summary (40/48 have contract.md; 18/48 have saved results;
empty results/ ≠ broken — most lanes regenerate on demand).
- Cross-provider relevance triage:
* 9 lanes are cross-provider-relevant and could benefit
from the prompt_battery-style adapter pattern (cognition,
english_fluency_ood, hebrew_fluency, koine_greek_fluency,
grammatical_coverage, inference_closure, multi_step_reasoning,
discourse_paragraph, foundational_*_ood, etc.).
* 29 lanes are CORE-only by design (versor closure, anchor
lens, identity divergence, provenance, etc.) — wiring
providers would be category-erroneous.
- Demo schema standardization status (this PR closes that).
- UI/UX coverage matrix.
- 5 concrete follow-up items, each focused enough for a single
PR, none requiring architectural change.
Regenerated reports
-------------------
evals/long_context_cost/results/comparison_v1.json and
evals/results/phase2_pack_measurements.json now contain the new
all_claims_supported field (auto-regenerated when validating the
schema change).
evals/frontier_compare/results/sample_core_promptbattery.json
added as a reference Lane B report so the new viewer always has
something to load on first open.
Connects ADR-0080's read-only contemplation loop to the existing
teaching-pipeline plumbing without forcing a type collapse. The
SPECULATIVE-only invariant from #55 is preserved verbatim; what
changes is *where the findings flow*.
What was wrong with the prior shape
-----------------------------------
PR #55 shipped a parallel core/contemplation/ package whose findings
were written as one JSON blob per CLI invocation, with no consumer.
The SPECULATIVE-only invariant protected a write path that didn't
exist. My closed PR #56 (second miner) would have entrenched the
duplication.
What this PR changes
--------------------
1. Schema (core/contemplation/schema.py)
- Adds a BOUNDARY note documenting why EvidencePointer (teaching)
and ContemplationEvidenceRef (core) intentionally stay separate:
EvidencePointer.source is constrained to {corpus, pack,
vault_coherent} — pointers into reviewed in-process memory the
runtime trusts. ContemplationEvidenceRef points to external
report files that have NOT been reviewed. Converging them would
either widen the runtime-grounding enum (losing the "reviewed
memory only" guarantee) or force benchmark reports to masquerade
as vault_coherent. Both are worse than keeping them separate.
- Adds format_contemplation_finding_jsonl(finding) — the canonical
JSONL formatter mirroring teaching.discovery.format_candidate_jsonl.
2. Runner (core/contemplation/runner.py)
- Both runners gain an optional sink: DiscoveryCandidateSink | None
parameter. When supplied, each finding is emitted as one
canonical JSONL line via the SHARED protocol — same protocol
that backs DiscoveryBufferSink and DiscoveryMonthlyFileSink.
- Sink path is additive: the ContemplationRun blob is byte-identical
whether or not a sink is supplied (pinned by test).
- No sink supplied → existing in-memory behavior preserved exactly.
3. CLI (core/contemplation/__main__.py)
- Adds --lane {frontier_compare, contradiction_detection} flag.
Default unchanged.
- Adds --sink-root <path> flag. When set, instantiates a
DiscoveryMonthlyFileSink and findings land at
<root>/<YYYY>/<YYYY-MM>.jsonl — the SAME layout discovery
candidates use, so operators can grep one stream.
4. Miner (core/contemplation/miners/contradiction_detection.py)
- Restored from closed PR #56 under the unified pipeline.
- Failure-mode split preserved (missed_contradiction /
false_contradiction_flag) with asymmetric repair actions.
What this PR does NOT do
------------------------
- Does NOT unify ContemplationFinding with DiscoveryCandidate.
DiscoveryCandidate.trigger is Literal[would_have_grounded,
successful_comparison, hedge_acknowledged, oov_resolved_via_decomp]
— all turn-loop flavored. None describe "I parsed a benchmark
report." Forcing a 5th trigger that no turn-loop extractor
produces would pollute the turn-loop type for the schema's sake.
- Does NOT extend teaching/gaps.py. Gap aggregates DiscoveryCandidate
cells by (subject, intent) — domain nouns. ContemplationFinding
subjects are namespaced ("contradiction_detection/CON-PUB-002").
Different operator views. A sibling aggregator can come later
when an operator actually asks for it.
Why this is the right unification point
---------------------------------------
The honest convergence is at the *sink* (so all SPECULATIVE evidence
lives in one rooted append-only stream), not the *aggregator* (which
appropriately produces typed views per evidence family). The boundary
doctrine from #55 is preserved; it now connects to existing plumbing
instead of writing JSON to disk with no consumer.
Tests (tests/test_contemplation_pipeline_convergence.py, 10 cases)
------------------------------------------------------------------
- DiscoveryBufferSink satisfies DiscoveryCandidateSink (shared protocol)
- frontier runner emits findings to shared sink
- contradiction runner emits findings to shared sink
- sink is optional — no-op when absent
- emission is canonical JSONL (sorted keys, no newline, deterministic)
- DiscoveryMonthlyFileSink persists findings at <root>/<YYYY>/<YYYY-MM>.jsonl
- sink emission does not alter the ContemplationRun blob (additive)
- contradiction miner predicate split + repair-action asymmetry
- config_hash differs between lanes (replay can distinguish)
- BOUNDARY doc is present in schema.py (regression guard)
- ContemplationEvidenceRef field invariants
- format_contemplation_finding_jsonl is deterministic + canonical
All 18 tests pass (5 original ADR-0080 + 13 new convergence).
Live evidence
-------------
$ uv run python -m core.contemplation \
evals/contradiction_detection/results/v1_public_*.json \
--lane contradiction_detection \
--sink-root /tmp/sink_demo
/tmp/sink_demo/2026/2026-05.jsonl ← same layout as discovery candidates
predicate=missed_contradiction subject=contradiction_detection/CON-PUB-002
predicate=missed_contradiction subject=contradiction_detection/CON-PUB-004
predicate=false_contradiction_flag subject=contradiction_detection/CON-PUB-005
predicate=false_contradiction_flag subject=contradiction_detection/CON-PUB-006