Commit graph

2191 commits

Author SHA1 Message Date
f94dbd4045 Merge pull request 'feat(logos): bulk live morph authority on turn + teaching seams' (#97) from feat/logos-bulk-live-authority into main
Reviewed-on: #97
---
Delivered: bulk Logos live authority

Branch: feat/logos-bulk-live-authority @ 91bb0a1b
Claim: bulk impact on live seams + sealed proof — not full Logos complete.

What improved (provable)

┌───────────────────────────┬────────────────────────────────────────────────────────────────────────────────┐
│ Arm / path                │ Behavior                                                                       │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ canonical                 │ PASS (no morph authority)                                                      │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ metadata                  │ Bit-identical digest to canonical                                              │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ executable                │ ABSTAIN on observed HE plural + singular-exclusivity claim                     │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ adversarial               │ FAIL_CLOSED on OOV / missing surface                                           │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ TeachingStore.add         │ Same decision fn → proposal CONTESTED                                          │
├───────────────────────────┼────────────────────────────────────────────────────────────────────────────────┤
│ CognitiveTurnPipeline.run │ Same decision fn → abstention surface, authority_source=logos_morph_constraint │
└───────────────────────────┴────────────────────────────────────────────────────────────────────────────────┘

Typed IR: CanonicalConstraint → LogosConstraint (rule_id, morphology_id, source_pack_id, source_span) — no free meaning dict on this path.

Verification (plan gates)

• Four-arm: wrong_count=0, provenance complete, dual-run digests identical
• Unit: 40 passed (morph + governance + surface)
• Live entry: turn abstains with he_morph_v0.plural_abstain
• Smoke: 180 passed (~133s)
• Residual note: docs/analysis/logos-bulk-live-authority-residual-2026-07-20.md

Still residual (honest)

• linguistic_pipeline cue tables remain parallel (not pack morph authority)
• One rule type; micro HE pack only
• Legacy math IR not fully unified
• No holonomy-crown / full lexicon claims

---

Live Logos morph authority is wired into teaching + CognitiveTurn: executable HE morph can abstain/refuse; metadata matches canonical; fail-closed on ambiguity; sealed four-arm proof + typed constraint IR.

Prove / re-run:
• uv run pytest tests/test_observed_he_morph_constraint_v0.py -q
• uv run python -c "from generate.observed_he_morph_v0 import run_four_arm_ablation; print(run_four_arm_ablation().as_dict())"
• uv run core test --suite smoke -q
2026-07-20 23:42:48 +00:00
Shay
91bb0a1b8c feat(logos): bulk live morph authority on turn + teaching seams
Wire observed-HE morph constraint into CognitiveTurn answer path and
unify teaching store with the same pure decision function used by the
four-arm ablation. Typed LogosConstraint IR carries rule/morph/pack/span
provenance without free meaning dicts. Executable mode abstains on
plural-vs-singular exclusivity; metadata remains bit-identical to
canonical; adversarial OOV/missing fails closed. Residual dual-system
debt documented honestly (cue-table linguistic_pipeline still parallel).
2026-07-20 16:35:04 -07:00
25504c9d0b Merge pull request 'feat(cognition): fail-closed linguistic governance + trilingual constraint pipeline' (#96) from feat/linguistic-governance-fail-closed-ir into main
Reviewed-on: #96
2026-07-20 23:23:52 +00:00
Shay
e0d1b4754a feat(cognition): fail-closed linguistic governance + trilingual constraint pipeline
Close residual answer-authority debt: typed CoherenceRefusal/ContractViolation/
FieldFailure, structured ProofTrace (atoms→operators→closure), and Shadow Gate
paths that refuse certified answers when contract_assessment is None or geometry
is open. Add shared semantic primitives and Layers A/B/C (English/Hebrew/Koine)
as constraint producers only, with Cl(4,1) structure-sensitive embedding (no
unitize in generate/), three field outcomes, and an articulation firewall that
blocks payload-value citation bypass and uncertified content.
2026-07-20 16:16:56 -07:00
18c578d960 Merge pull request 'feat: Master Convergence Stages 1–4 (land closed stack + skeptic fixes)' (#95) from feat/observed-he-morph-constraint-v0 into main
Reviewed-on: #95
---
Cl(4,1) geometric sovereignty convergence (Master Blueprint Stages Pre→1–4) is on feat/observed-he-morph-constraint-v0 (tip 8c4221d4), with Forgejo PRs #90–#94.

• Validate: uv run core test --suite smoke -q (or full: uv run core test --suite full -q)
• Stage pins: uv run pytest tests/test_geometric_convergence_checklist.py tests/test_stage3_epistemic_inductive.py -q
• HE four-arm ablation: PYTHONPATH=. python3 -c "from generate.observed_he_morph_v0.ablation import run_four_arm_ablation; print(run_four_arm_ablation())"
2026-07-20 22:40:57 +00:00
Shay
8c4221d496 fix(stage3): gate inductive expansion on geometric admissibility
Non-admissible candidates no longer enter work/derived. Only admissible
base and derived edges seed fixed-point steps, so multi-step paths close
only under geometric conditions (Stage 3 exit). Updates tests that
previously locked in admissible=False-but-still-promoted behavior.

[Verification]: tests/test_stage3_epistemic_inductive.py 10 passed
2026-07-20 15:07:12 -07:00
Shay
9f85832baa fix: full-suite gates — suite SoT + multi-root depth pin
- Re-export core.cli._TEST_SUITES from cli_test.TEST_SUITES so argparse
  and tests cannot drift from the runtime packs/smoke/algebra pins
  (Stage dual-pack draft boundary was only on cli_test).
- Align relational ablation multi-root metadata test with Stage 3
  fail-closed policy: ≥2 unique roots do not emit [root:] notes.

[Verification]: Smoke suite passed locally (~133s, 180 passed);
postfix targeted 5/5; full pre-fix was 4 failed / 13281 passed
(2 env dirt/flake classified separately).
2026-07-20 14:58:21 -07:00
Shay
c71c00cca6 fix: skeptic remediations for Stage 3–4 exit gates
- Metadata arm returns baseline decision payload (bit-identical digests).
- TeachingStore auto-applies HE morph rule from compiled pack (live path).
- Inductive derived edges stamp geometric admissibility via pipeline versor
  grounding; refuse ungrounded promotions.
- VaultPromotionPolicy default residual_threshold=1e-6 for COHERENT.

[Verification]: skeptic_fixes 26 passed; stage4 ablation digests equal
2026-07-20 14:08:37 -07:00
3de431dcdc Merge branch 'feat/stage3-he-ambiguity-epistemic-closure' into feat/observed-he-morph-constraint-v0 2026-07-20 21:05:42 +00:00
7755fd4708 Merge branch 'test/stage2-physics-parity-hardening' into feat/stage3-he-ambiguity-epistemic-closure 2026-07-20 21:05:30 +00:00
2a32887809 Merge branch 'docs/stage1-governance-boundary-freeze' into test/stage2-physics-parity-hardening 2026-07-20 21:05:18 +00:00
68cf8258fa Merge branch 'main' into docs/stage1-governance-boundary-freeze 2026-07-20 21:05:03 +00:00
1f1a94a39b Merge pull request 'feat(physics,cognition): Cl(4,1) geometric sovereignty convergence' (#90) from feat/cl41-geometric-convergence-sovereignty into main
Reviewed-on: #90
2026-07-20 21:04:49 +00:00
Shay
121e9ebb3c feat(he): Stage 4 observed-HE morph constraint v0 + four-arm ablation
Vertical slice feat/observed-he-morph-constraint-v0:

- Load observed HE morphology from compiled packs/data/he_logos_micro_v1
  with exact source_span provenance.
- Authored PluralAbstainRuleV0 → language-independent CanonicalConstraint.
- TeachingStore.add consumer seam: ABSTAIN/FAIL_CLOSED → CONTESTED.
- Sealed four-arm ablation: canonical, metadata (inert), executable
  (decision change), adversarial OOV/missing (fail closed).
- No English-to-Hebrew pseudo-morphology.

[Verification]: observed-HE 5 passed; teaching regression 15 passed; smoke 180 passed
2026-07-20 14:00:01 -07:00
Shay
f9f94d8df8 feat(cognition): Stage 3 geometric coherence, inductive closure, OOV pins
Complete Stage 3 residual gates on live pipeline authority:

- GeometricCoherenceVerdict distinguishes field-closed vs unverified
  without inventing EpistemicState.COHERENT (dual taxonomy ownership doc).
- Bounded expand_relation_closure over teaching-store triples with budget,
  cycle safety, base multi-tail contradictions, replayable provenance in
  operator_invocation / trace_hash.
- OOV/egress authority tests: conformal neighbors context, cga_inner nearest,
  vault_hits not surface gate.

[Verification]: Stage 3 unit 18 passed; pipeline regression 27 passed; smoke 180 passed
2026-07-20 13:54:51 -07:00
Shay
aaa8503f0b feat(recognition): Stage 3 fail-closed multi-root depth ambiguity
Stop silent roots[0] commitment when multiple HE/GRC roots are observed.
Return typed RootSenseAmbiguity, mark runnable assessments non-runnable
with AMBIGUOUS_ROOTS provenance, and refuse multi-root node canonicalization
without authored single-candidate resolution.

[Verification]: stage3 root ambiguity + oov pipeline tests passed (39)
2026-07-20 13:47:02 -07:00
Shay
b1dc57ff22 test(physics): Stage 2 hardening pins for scale domain, L2 excision, parity
Executable Stage 2 exit evidence: shared fraction-decrease scale domain,
identity L2 helper absence, GoldTether geometric residual path, RATIFIED|
DEMOTED-only ratifier, active Cl(4,1) product/conformal/Gram/GT/sandwich
smoke, dual-backend vs Python-authoritative parity boundary (no false
rotor_power Rust claims), and LE f64 SHA-256 digests.

[Verification]: algebra suite green (includes Stage 2 pins); stage2 9 passed 1 skipped
2026-07-20 13:45:05 -07:00
Shay
8494a239e4 docs(governance): Stage 1 ADR collision freeze + dual-pack serve boundary
Reconcile Master Blueprint ADR-0240–0253 titles with the live Accepted
registry without renumbering history. Claim ADR-0253 for the freeze policy
itself; reserve 0254–0261 for Blueprint gaps. Document packs/data vs
packs/he|grc dual boundary and pin serve isolation in smoke/packs suites.

[Verification]: Smoke 180 passed (includes 4 dual-pack boundary pins)
2026-07-20 13:43:25 -07:00
Shay
6ffea249ec fix(cognition): break ContractAssessment import cycle in pipeline
Lazy-import ContractAssessment inside geometry contract assessment so
problem_frame_contracts → chat → cognition → problem_frame_contracts
cannot form a partial-init cycle. Unblocks GSM8K rat1/wave_a runners.

[Verification]: rat1+wave_a+pipeline+surface 38 passed; import path ok
2026-07-20 13:36:53 -07:00
Shay
120e9554ab fix(cognition): pre-stage land repairs after L2/PASSTHROUGH excision
Close residual failures from geometric sovereignty hardening without
restoring scalar-L2 or PASSTHROUGH authority:

- Ratify CORRECTION/COMPARISON via vocab-grounded tag and multi-token
  subject anchors so teaching capture and cognition intent accuracy hold.
- Keep observational wave leakage from vetoing teaching while
  identity_wave_gate is off; syntactic override still rejects.
- Surface GoldTetherViolationError as fail-closed tether residual rather
  than aborting lifecycle observation.
- Replace excised legacy identity-eval path with a geometry-blind baseline.
- Align sensorium corridor and lift instrument assertions with sandwich
  unitary close and honest PARITY when baseline already solves.

[Verification]: smoke 176 passed; cognition 122 passed 1 skipped;
teaching 109 passed; claim batch 34 passed
2026-07-20 13:34:59 -07:00
Shay
e7d116c924 feat(physics,cognition): Cl(4,1) geometric sovereignty convergence
Excise identity L2 dual-mode and pipeline PASSTHROUGH cold-starts; enforce
wave-only IdentityCheck with MissingWaveStateError; hard-fail GoldTether
transitions via GoldTetherViolationError; dual-competing Shadow Coherence
Gate with populated contract_assessment; auto-compile field packets before
intent ratify; conformal argmax ratification and content-addressed proof
atoms; sandwich multi-modality ingress with GoldTether digests.

Amend ADR-0243/0244 (docs/adr + research) for wave-only / SUPERSEDED sketches.
Pin binary gates in tests/test_geometric_convergence_checklist.py.

[Verification]: Smoke suite passed locally (~136–142s, 176 passed);
cognition suite passed (122 passed, 1 skipped); lifecycle suite 51 passed.
2026-07-20 12:20:05 -07:00
4b9dbe7236 Merge pull request 'fix(a2k): reject out-of-range fraction-decrease scales' (#89) from fix/a2k-fraction-decrease-scale-range into main
Reviewed-on: #89
2026-07-20 04:42:05 +00:00
Shay
ac4c3e04ad test(a2k): prove geometric path refuses out-of-range decrease scales
Regression coverage for the audit repro (5/4), zero/one/gt-one geometric
no-bypass, injected Fraction domain mutations, promotable resolver
refusal, and flip the PR #87 preexisting gap pin to require operator
refuse alongside baseline.
2026-07-20 04:40:35 +00:00
Shay
1a6efe0d13 fix(a2k): share fraction-decrease scale domain on geometric admission
Obligation assessment required 0 < scale < 1 (scale_out_of_range) while
_versor_binding_from_scale_value admitted any positive finite scale, so
resolve_promotable_fraction_decrease could still bind and return a
negative decrease for inputs like 5/4.

Add _is_valid_fraction_decrease_scale as the single semantic-domain
predicate and enforce it in both assess_fraction_decrease and
_fraction_decrease_scale_binding so geometric VersorBinding cannot
grant permission the obligation layer denies.
2026-07-20 04:40:35 +00:00
6b8a013280 Merge pull request 'docs(analysis): PR #87 post-merge integrity audit' (#88) from chore/postmerge-pr87-integrity-audit into main
Reviewed-on: #88
2026-07-20 04:37:04 +00:00
Shay
375da93e61 docs(analysis): PR #87 post-merge integrity audit
Evidence-led verification of merged relational-operator ablation:
canonical semantics, fixture independence, condition labels, scale>1
isolation, production/gate isolation, HE/GRC roadmap authority, and
reproducibility. Verdict: VERIFIED, no corrective PR needed.

[Verification]: Ablation 20 passed; runner digest match; organ pre-existing
37 passed; smoke 176 passed locally (~133s).
2026-07-19 21:24:34 -07:00
5fae5a67f2 Merge pull request 'feat(eval): deterministic relational operator ablation + HE/GRC Logos roadmap' (#87) from feat/deterministic-relational-operator-ablation into main
Reviewed-on: #87
---
Delivered

┌─────────────────────────────┬──────────────────────────────────────────────────────────────────────────────────┐
│ Item                        │ Location                                                                         │
├─────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────┤
│ Cartography ledger          │ docs/analysis/relational-operator-ablation-cartography-2026-07-19.md             │
├─────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────┤
│ Ablation dossier            │ docs/analysis/relational-operator-ablation-dossier-2026-07-19.md                 │
├─────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────┤
│ HE/GRC Logos roadmap        │ docs/analysis/hebrew-koine-greek-logos-pack-capability-roadmap-2026-07-19.md     │
├─────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────┤
│ Ablation code + sealed eval │ generate/relational_operator_ablation.py, evals/relational_operator_ablation/v1/ │
├─────────────────────────────┼──────────────────────────────────────────────────────────────────────────────────┤
│ Tests                       │ tests/test_relational_operator_ablation.py                                       │
└─────────────────────────────┴──────────────────────────────────────────────────────────────────────────────────┘

Git

Branch: feat/deterministic-relational-operator-ablation
Remote: Forgejo core-labs/core (5 commits)

1. df029a84 — cartography ledger
2. 1d4d51bc — ablation runners
3. 87b01868 — sealed eval lane v1
4. 84f7a6f7 — tests + ablation dossier
5. 08881ab6 — HE/GRC Logos roadmap + cross-links

Pull request

#87

• Title: feat(eval): deterministic relational operator ablation + HE/GRC Logos roadmap
• Mergeable: yes
• [Verification]: smoke 176 + ablation 20 + runner n=8 wrong=0 documented on the PR

Fresh start for next PR

Roadmap recommends:

┌──────────┬──────────────────────────────────────────────────────────────────────────────┐
│ Field    │ Value                                                                        │
├──────────┼──────────────────────────────────────────────────────────────────────────────┤
│ Branch   │ feat/observed-he-morph-constraint-v0                                         │
├──────────┼──────────────────────────────────────────────────────────────────────────────┤
│ Worktree │ ../core-observed-he-morph-constraint                                         │
├──────────┼──────────────────────────────────────────────────────────────────────────────┤
│ Scope    │ Observed-Hebrew morph constraint only (not pack bulk-load, not English math) │
└──────────┴──────────────────────────────────────────────────────────────────────────────┘
2026-07-20 04:16:08 +00:00
Shay
08881ab665 docs(analysis): land HE/GRC Logos pack capability roadmap
Audit-only dossier: current-state matrices, dual pack systems, holonomy
limits, no-go list, and recommended next branch
feat/observed-he-morph-constraint-v0. Cross-link cartography and ablation
dossiers; no pack expansion or ablation fixture changes.
2026-07-19 21:14:21 -07:00
Shay
84f7a6f7cc test(eval): prove relational operator ablation invariants + dossier
Unit, integration, and adversarial tests; isolate pre-existing A2k scale>1
gap; final engineering dossier with measurement table.
2026-07-19 21:00:15 -07:00
Shay
87b01868fa feat(eval): sealed relational operator ablation lane v1
Eight-case fixture, measurement runner, and committed report for baseline
/ operator / depth / metadata_only / invalid conditions (wrong=0).
2026-07-19 21:00:15 -07:00
Shay
1d4d51bc9b feat(generate): deterministic relational operator ablation runners
Baseline scalar frame path, live geometric operator path, inactive depth
contribution provenance, metadata-only control, and fail-closed invalid
condition for fraction_decrease — without inventing EN→HE/GRC role maps.
2026-07-19 21:00:15 -07:00
Shay
df029a8485 docs(analysis): cartography ledger for relational operator ablation
Evidence matrix, live-path maps, claim corrections, and go/no-go for a
narrow fraction_decrease ablation (not ancient-language GSM8K magic).
2026-07-19 21:00:15 -07:00
d8d62b8ea3 Merge pull request 'feat(trackb): S1–S4 symbolic SME, selector, pure-S1 coverage gain (Inc2)' (#86) from feat/trackb-symbolic-sme-s2s4 into main
Reviewed-on: #86
2026-07-20 03:12:47 +00:00
Shay
2544fda6f6 fix(trackb): pure-S1 extract refuse multi-clause overmatch (0361 wrong=0)
The inverted-seed P2 regex treated optional "times" as optional, so
additive "8 more" on multi-entity zoo case 0361 extracted as pure S1 and
the three-gate corridor emitted 1.125 vs gold 114.0 (wrong≠0).

- Require explicit multiplicative surface (times as many / twice the …)
- Fail-closed purity gate: multi-clause markers, multi-numeric, multi-mult
- Regression tests for 0361 and bare "N more"
- measure --mode full-holdout-wrong0 over all 500 holdout cases

[Verification]: full-holdout emit_ok=9 wrong=0 refused=491; unit 37 passed
2026-07-19 20:04:27 -07:00
Shay
e5643454d9 feat(trackb): S1–S4 symbolic SME, selector, pure-S1 coverage gain
Increment 2 for ADR-0252 Track B. Extends Increment 1's structure-mapping
slice with pure-family canonicals S2–S4, overlapping-waves selector, and a
structure-mapping-owned pure-S1 text extract that recovers four real
holdout cases the serving reader misses (0148, 0228, 0234, 0441).

Bar results (command-backed via scripts/measure_trackb_inc2.py):
- generalization ratio S1 = 9.0 (9 holdout cases / 1 template)
- coverage gain organ 5 → trackb 9, wrong=0
- selector routes S1/S2/S3; refuses empty; surface≠structure

Off-serving: no organ retirement, serving reader untouched. S2–S4 holdout
ratios are 0 (parser frontier). S3/S4 emit refuses at multi-register scope
without weakening the three-gate wrong=0 path.

[Verification]: Smoke suite passed locally (~132s, 176 passed);
trackb unit tests 35 passed; measure_trackb_inc2 all modes green.
2026-07-19 19:49:07 -07:00
f6fd5030bb Merge pull request 'feat(trackb): S1 symbolic structure-mapping vertical slice (off-serving)' (#85) from feat/trackb-symbolic-sme-s1 into main
Reviewed-on: #85
---
This is the verification I promised — and Grok's report is a different animal from the last three: it self-flagged the vacuous-FP trap, found and fixed a real wrong-answer bug, and stopped at the gate. That's what honesty looks like. But the whole point is that I check the code, not the report. Pulling the integrity-critical files.Verified against the code — and this one **passes.** First PR in the entire arc where the code matches the report. Credit to Grok: it did the real thing.

What I actually confirmed in the source, not the summary:

**Blindness holds.** `map_to_s1` takes only a `RoleGraph`; it references `S1_CANONICAL.structure_id` (allowed prior knowledge) but never a gold label. In the measure script, `score_label` is called strictly *after* `_map_decision` returns. The mapper cannot see the answer. ✓

**`wrong=0` is protected by genuine defense-in-depth — three independent fail-closed gates**, and I traced each:
1. *Map-time:* the total's part-set must **equal** `{a,b}` exactly (not superset), and any `contain` on `b` or a third entity refuses. The "true 115 → emit 15" bug class can't map.
2. *Corridor agreement:* classical `solve`+`verify` AND multi-register `certified` AND `|classical − mr| ≤ 1e-6`, else refuse.
3. *Original-graph backstop:* after emit, it re-solves the **original** graph and refuses on `original_graph_disagreement`. So even a mapper bug that drops an entity gets caught here.

Every path is conjunctive and refuses by default — I found no fail-open, and the `except Exception` explicitly refuses rather than counting the crash as signal. The measure script even returns a nonzero exit code if `wrong != 0`, so it can't print success while failing. That's the discipline we asked for, actually implemented.

**Separability is honestly bounded.** Grok flagged the exact vacuous-FP trap I warned about — only 5 holdout cases parse, so it prints the vacuity note and measures real separability on synthetic negatives + **135 real non-compare graphs from the other corpora (0 false positives)** instead of hiding behind a vacuous holdout number.

Now the honest calibration, so we don't over-celebrate — this is Increment 1, and it proves the *architecture*, not yet the *value*:

- **No coverage gain yet.** It ties the S1 organ on the organ's own 5 cases (5/5 match). It doesn't yet solve anything the organ couldn't.
- **0148 still doesn't parse** — the motivating case. That confirms the old diagnosis: the bottleneck is the *reader/parser*, and structure-mapping downstream can't rescue a parse that never produces a graph.
- **The §4 bar — generalization ratio > 1 — is untested.** What we have is a precise, surface-invariant *template match* for one graph shape, not yet Gentner-style discovery of correspondence across dissimilar problems, and no families have collapsed or organs retired.

So my ruling: **this is a real GO for Increment 1 — merge it.** It's off-serving, additive, blind, fail-closed, and verified. One 1-command confirmation before you merge, so my off-serving sign-off is evidence-based rather than assumed:

```
git diff --stat main..feat/trackb-symbolic-sme-s1
```

Every path should be under `generate/structure_mapping/`, `evals/structure_mapping/`, `scripts/`, `tests/` — nothing in `math_candidate_graph.py` or the organ dispatch. If so, merge #85.

The next increment is where the paradigm has to *earn its keep*: extend to S2–S4 **and measure whether the mapper generalizes past the organs** — coverage the organs don't have, families collapsing to shared structure, at least one organ retired under the 3-part proof. If it only ever ties the organ on the organ's turf, the paradigm isn't paying yet. Want me to draft the Increment-2 handoff to that bar?
2026-07-20 02:33:41 +00:00
Shay
85781f0f4c fix(trackb): pure-S1 gate — refuse total supersets and extra contains
Mapper required only that total include {a,b}; rebuild-from-binding then
dropped third entities and could emit a certified wrong answer (e.g. 15
instead of 115). Require exact total part-set {a,b}, refuse any contain
outside reference a, and agree with original-graph classical solve before
emit. Adds multi-entity refuse unit tests.
2026-07-19 19:10:05 -07:00
Shay
943614c522 feat(trackb): S1 symbolic structure-mapping vertical slice (off-serving)
ADR-0252 §5 geometric SME is NO-GO; this is Track B Increment 1.
Adds role-predicate conversion from MathProblemGraph, S1 canonical skeleton,
blind symbolic mapper (match/refuse + binding), and solve via classical
verify plus multi-register certificate. Research report and holdout measures
included. Serving reader unchanged; no S2–S4 generalization.
2026-07-19 19:01:23 -07:00
Shay
1ccef49130 docs(adr): land ratified ADR-0252 governing problem-solving paradigm 2026-07-19 16:09:52 -07:00
Shay
8b995ce0fb docs(paradigm): dead-code audit — zero removals, all candidates traced and verified
Trace every module/class the six retired paradigm documents introduced or
proposed, plus a general zero-importer sweep of generate/, core/, evals/.
Every candidate is one of: never built, already removed in a prior pass, or
live (imported by the serving reader / a protected dependency chain, and/or
test-covered). No code removed; no behavior change.

Baseline == final (no code touched):
  smoke: 176 passed
  full:  13163 passed, 111 skipped, 1 xfailed, 0 failed
  holdout_dev v1: correct=5 wrong=0 refused=495 (n=500)
2026-07-19 15:15:52 -07:00
Shay
7a6556e852 docs(paradigm): retire six competing unratified problem-solving paradigms under ADR-0252
Archive the six competing, unratified problem-solving-paradigm formulations
(semantic-substrate master plan, state-transition blueprint, binding-graph
proposal, capability roadmap v2, eight sprint lookbacks, and ADR-0174's
deprecated dispatch prescription) to docs/paradigm-archive/, each prefixed
with a SUPERSEDED header. ADR-0164 and ADR-0174 stay in place (Accepted,
permanent record) with a Superseded-by: ADR-0252 line added. ADR-0252 itself
is a decision record only — it records the retirement and reserves the
governing paradigm content for a separate, human-led step.

No code changes; no behavior changes.
2026-07-19 14:46:47 -07:00
1de6dbc945 Merge pull request 'docs(reader-arc): recalibration ADR-0251 + preserved knowledge + overfit inventory' (#81) from docs/reader-arc-recalibration into main
Reviewed-on: #81
2026-07-19 20:46:39 +00:00
Shay
f8bf960d8d docs(reader-arc): recalibration ADR-0251 + preserved knowledge + overfit inventory
Records the reader-arc recalibration: halt bespoke-per-case regex work,
verified clean-base reset to main@854ce023, and eradication of
feat/reader-inc2-caseband's (PR #80) debt (loosened _token_in, empty
proof file, bundled provenance) without merging that branch.

Preserves the genuinely-general knowledge before the branch is retired:
the has_numeric_token multiplier-blindness bug (root cause + minimal fix
idea, deliberately not landed) and case 0148's correct-for-the-right-reason
decomposition, both independently re-verified against source rather than
copied from the branch's own claims (which included a stale narrative
about a graph-dump proof file that was actually committed empty).

Adds a file:line-cited overfit inventory (31 bespoke/single-case surface
patterns) as inventory only -- nothing removed, the existing regex reader
keeps serving unmodified.

ADR-0251 proposes (not builds) a geometric surface->canonical
normalization spike using existing off-serving primitives
(conformal_procrustes, VocabManifold, the cognition pipeline's Phase C
anti-unification telemetry), wrong=0-gated against holdout_dev/v1, with
Smith-chart algebra and any Fibonacci "dimensional cascade" explicitly
disclaimed. Awaiting ruling -- no implementation started.
2026-07-19 13:42:47 -07:00
854ce023d9 Merge pull request 'docs(reader-arc): increment-1 BAND plan — {seed + comparison(#78) + question-arithmetic} (for ruling)' (#79) from feat/reader-band-increment-1 into main 2026-07-19 04:58:00 +00:00
Shay
81688465a5 test(reader-arc): Q3 hazard-close — lock q:difference fail-closed (no guessed direction)
Discharges Josh's Q3 ruling: defer the q:difference capability, close the
hazard now. _pattern_b_comparative_candidates is inert by construction
(returns []), so 'how many more/fewer … than …' questions refuse end-to-end
even when every statement injects cleanly. This suite LOCKS that so a future
D.5 wiring cannot reopen the guessed-direction hole without positively
determining the more/fewer direction. Verified inert-safe by 5 direct tests.
2026-07-18 21:57:36 -07:00
Shay
2da1a33839 fix(reader-arc): compound extractor refuses no-digit compare clauses (wrong=0 hazard)
The compound discrete-count extractor's all-or-nothing tail guard is digit-
based, so a no-digit compare clause in a conjunctive list ('buys 4 lbs beans,
6 lbs milk, and twice the amount of carrots as beans') was silently DROPPED —
injecting 4+6 and losing the carrots (a wrong=0 landmine, masked only by the
case refusing elsewhere). The #78 substrate widens the compound extractor's
verb reach, so harden it: refuse the whole compound on any '<factor> the
amount/number of' / 'twice/thrice/times/half the' / 'double/triple' surface.
Verified: 0082 compound now refuses fail-closed; tune+measure wrong=0; PARSED
unchanged (3/2); smoke 176 green.
2026-07-18 21:44:34 -07:00
Shay
c217af31ee docs(reader-arc): band-solve=0 outcome + minimal-convertible-band measurement
Records the increment-1 outcome straight (band-solve=0, ~30-band bet
falsified; foundations = banked byproduct not success) + the minimal-band
measurement: no small band converts; taxonomy not at bedrock (GSM8K statements
are per-sentence multi-capability); multi-compound highest-leverage (70/106
tractable needed-sets) but insufficient alone; q:complex biggest intractable
wall (~35 cases). Increment-2 recommendation = CASE-FIRST (target specific
closest cases, build their exact end-to-end needs) not capability-first, since
the taxonomy fragments. Foundations verified full-500 wrong=0, smoke 176.
2026-07-18 21:37:16 -07:00
Shay
387f065649 feat(reader-arc): #78 positive-polarity allowlist substrate (seed + comparison) — foundations, band-solve=0
The two reusable foundations of increment 1, both verified wrong=0:
- NEUTRAL_COUNT_VERBS (math_roundtrip): the shared positive-polarity ALLOWLIST
  (production/acquisition/possession-count verbs), fail-closed. Extends ADD_VERBS
  with curated neutral production verbs (score/write/teach/plant/harvest/produce/
  draw/sing). Depletion/transfer (SUBTRACT/TRANSFER_VERBS) excluded by positive
  determination — 'Alice lost twice as many' refuses, not via a blocklist.
- Seed reader: discrete_count matcher acquisition path aligned to the allowlist
  (made/baked/grew/scored/wrote/taught now inject; depletion refuses).
- Comparison reader: _comparison_anchor_verb() allowlist extended to the same
  NEUTRAL_COUNT_VERBS (keeps fail-closed; captures scored/caught/wrote frames).

Verified: production injects, depletion refuses; 71 pinned comparative tests
pass (polarity guards now via allowlist); tune wrong=0; smoke 176 green.

FINDING (measured 4 ways: yield harness, band map, band oracle, capability
histogram): band-1 as scoped {seed + forward-comparison + q:simple/summation}
converts 0 tune cases. Its otherwise-in-band cases STRAND on multi-compound (25)
and compare-additive (17) — capabilities outside the ~30 band. The conjunction
recurs at sub-capability granularity; the band-map taxonomy was too coarse.
Foundations validated + reuse-ready (done-when part 2); band-solve>0 (part 1)
needs the band to grow to +compare-additive(+multi-compound) — a scope/done-when
ruling for Josh. No merge until ruled.
2026-07-18 21:18:02 -07:00
Shay
12098cb3d3 docs(reader-arc): increment-1 BAND plan — {seed + comparison(#78) + question-arithmetic} for ruling
Design-first plan, ruling before any build (same gate as #76). Band =
foundational spine core: seed-simple + comparison + question-arithmetic
(simple/summation/difference), on the shared #78 polarity substrate.

Bakes in Josh's four refinements: (1) deliberate ~30 smaller-band rationale
(first end-to-end reader conversion ever + two reusable foundations, zero new
hard capabilities beyond #78); (2) done-when reset to band-solve>0 wrong=0 +
foundations validated, ~40 bar cumulative across increments 1-3; (3) haircut
factor (real/upper-bound) as a named output; (4) TWO wrong=0 direction drivers
— #78 statement polarity AND q:difference question direction, both positively
determined + fail-closed. Grounds design on what EXISTS (question layer partly
built: summation works, q:difference partial). Schedules q:complex decomposition
next (may be corridor-tractable, could move the ~48% ceiling).
2026-07-18 20:54:28 -07:00
e1eb2a5ca1 Merge pull request 'feat(reader-arc): compare_multiplicative increment 1 — compiler tier + frame-anchored reader (measured delta=0, reroutes to shared-layer)' (#77) from feat/compare-mult-increment into main 2026-07-19 03:07:06 +00:00