core/docs/sessions/2026-05-23-pedagogy-research-and-teaching-loop-pivot.md
Shay 310aed9ff0
chore: Refactor CLI and Governance Anchors (#926)
* docs: consolidate governance anchors and clean up test registries

* refactor(cli): decompose cli into dedicated modules

* test: fix broken test baselines and formatting

* docs: add domain boundary READMEs for governance anchors

* test: update baseline for determination lane

* test: fix capability_pass expectation

* test: fix CORE_SHOWCASE_SKIP_BUDGET enforcement

* chore: cleanup CLI extraction and unreachable code
2026-07-03 12:34:56 -07:00

336 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# SESSION 2026-05-23 — Pedagogy Research & Teaching-Loop Potential Pivot
**Date:** 2026-05-23
**Status:** Research note; load-bearing for ADR-0129 + ADR-0130
**Trigger:** Operator-supplied review of *Beyond Traditional Pedagogy:
Research-Based and Emergent Techniques for Deep, Durable Learning*
(`/Users/kaizenpro/Downloads/Beyond Traditional Pedagogy ...md`,
2026-05-23)
**Branch:** `docs/pedagogy-review-and-teaching-backlog`
---
## Why this session exists
CORE's mid-2026 work has concentrated on the GSM8K-math substrate arc
(ADRs 0114a → 0119 → 0120 → 0121 → 0122 → 0123 / 0123a / 0123b → 0126
candidate-graph topology → 0127 units pack → 0128 numerics pack). The
last three substrate ADRs each produced **zero sealed-holdout lift**
despite being correct work, leading to an architectural pivot (ADR-0126)
and a substrate-substrate (ADR-0127 / 0128) reframing.
That sequence has been **all about the truth-articulation path**
parse → graph → solve → verify → realize. The orthogonal axis — how
CORE *learns* from reviewed corrections — has not received the same
load-bearing attention since the ADR-0040-series teaching-substrate
work. The operator surfaced a pedagogy literature review as a sanity
check on whether the teaching loop, considered on its own merits,
has structural gaps that the GSM8K-math focus has been deferring.
This session is the result of that check: the literature review of
the supplied document, follow-up confirmation research on contested
claims, and the resulting two backlog ADRs (0129 and 0130).
---
## The reviewed document
**Title:** *Beyond Traditional Pedagogy: Research-Based and Emergent
Techniques for Deep, Durable Learning*
**Structure:** Executive summary + ~10 themed sections + a
synthesis table + 22 reference URLs. ~300 lines, well-cited within
the established cognitive-psychology / learning-science canon
(Bjork, Roediger & Karpicke, Kapur, Mayer, Collins / Brown / Newman,
Freeman et al., etc.).
**Headline claims:**
1. Active learning > passive lecture (Freeman et al. 2014 PNAS
meta-analysis as exemplar).
2. Retrieval practice (effortful recall) drives durable learning;
spacing + interleaving amplify.
3. Productive failure (Kapur) produces larger conceptual gains than
instruction-first ("3x" rhetoric in some references).
4. Embodied cognition: gesture, manipulation, handwriting matter for
acquisition.
5. Multimedia learning (Mayer): coordinated verbal + visual channels
subject to cognitive-load management.
6. Cognitive apprenticeship (Collins / Brown / Newman): modeling,
coaching, scaffolding, articulation, reflection, exploration.
**Treatment quality:** sound at the survey level; weak on
calibration of contested findings.
---
## Literature confirmation pass
To avoid uncritical adoption, three areas with known replication
or boundary concerns were searched against 20242025 literature:
### 1. Productive failure — calibration of the "3x" rhetoric
**Anchor:** Sinha & Kapur 2021 meta-analysis (166 experimental
comparisons, ~12,000 participants), [SAGE](https://journals.sagepub.com/doi/full/10.3102/00346543211019105).
| Claim | Reality |
|-------|---------|
| "3x conventional gains" | Headline from high-fidelity PF studies; meta-analysis average is **d = 0.36**, rising to **d = 0.58** at high design fidelity. Real but more modest. |
| "Broadly applicable" | **Largely a STEM finding.** Non-STEM evidence scarce; domain-general skill transfer not supported. |
| "Works for all learners" | Better effects for **older students** (secondary onwards); prior knowledge is a strong moderator (PMC 2023 study on prior math achievement). |
**Verdict for CORE:** PF is the doc's most-overstated technique.
The structural analog inside CORE (let-attempt-then-review)
already exists in adversarial generation (ADR-0119.5), but with
a different mechanism — adversarial generation is a wrong-answer
*rejection* tool, not a learning-from-attempt tool. Adopting PF
shape inside CORE would mean intentionally allowing the engine
to attempt with knowingly-insufficient grounding and learning
from the gap. **This is the deliberate inverse of CORE's
`wrong==0` doctrine** and would require structural justification
beyond "the literature supports it."
### 2. Retrieval practice — transfer limits
**Anchor:** Pan & Rickard 2018 transfer meta-analysis;
Cognitive Research 2024 follow-up on far-transfer mechanisms,
[Cognitive Research](https://cognitiveresearchjournal.springeropen.com/articles/10.1186/s41235-024-00598-y).
| Claim | Reality |
|-------|---------|
| "Retrieval drives transfer" | **Near transfer: yes (d = 0.4). Far transfer: weak/null (Pan & Rickard d = 0.16, n.s.).** |
| "Works for complex material" | Strongest for simple materials learned by rote; complex / educationally relevant materials show smaller, more contingent effects. |
| "Universal mechanism" | Recent work (Cognitive Research 2024): far-transfer benefits appear specifically when **rule-based learning** is the underlying mechanism + after delay. |
| "Lecture-hall ecological validity" | Glaser & Richter 2025 ([Teaching of Psychology](https://journals.sagepub.com/doi/10.1177/00986283231218943)): testing effect transfers poorly to studied-but-not-practiced content. |
**Verdict for CORE:** Retrieval practice IS the most robust
finding *for retention of practiced material*. CORE's vault recall
already encodes the exact-recall ceiling of this technique. The
*spaced-retrieval* extension (spacing across time) is the part
not currently modeled in CORE's teaching loop — see ADR-0129.
### 3. Embodied cognition — replication crisis
**Anchor:** Machery 2024 chapter on the embodied-cognition
replication crisis,
[Routledge Handbook of Replication](https://www.taylorfrancis.com/chapters/edit/10.4324/9781003322511-50/replication-crisis-embodied-cognition-research-edouard-machery);
Frontiers in Education 2026 STEM-learning integrative review,
[Frontiers](https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2026.1811569/full).
| Claim | Reality |
|-------|---------|
| Embodied learning effects | **Known replication crisis.** Foundational findings have failed independent replication. |
| Handwriting > typing | Strongest for very early literacy acquisition; broader generalizations are contested. |
| Universal benefit | "Embodiment sometimes facilitates learning and sometimes does not" — boundary conditions matter (Frontiers 2026). |
**Verdict for CORE:** Not applicable directly (no body, no
sensorimotor system). Structural analogs (e.g., the
algebra/field/vault substrate as "grounding in a non-symbolic
representation") exist but the analogy is too weak to load-bear
design decisions.
---
## What the doc missed (frameworks worth knowing)
These should be on the radar even though they weren't in the
reviewed document:
| Framework | Why it matters |
|-----------|----------------|
| **Worked-example effect** (Sweller, Paas, van Merriënboer) | Strong evidence for novice instruction; counter-evidence for experts (see expertise-reversal) |
| **Expertise-reversal effect** | Techniques that help novices actively hurt experts and vice versa. Directly relevant to CORE's `apprentice → audit-passed → expert` promotion contract (ADR-0120) |
| **Cognitive load theory** (Sweller) | Distinct intrinsic / extraneous / germane load distinction. Operationally useful for designing teaching corpora |
| **Deliberate practice** (Ericsson) | Specific goals + immediate feedback + repetition at the edge of capability. Better lens than "active learning" for skill domains |
| **Self-explanation effect** (Chi) | Narrow but strong evidence, particularly for science learning from worked examples |
| **Bloom's 2-sigma problem** (1984) | Unsolved benchmark: 1:1 tutoring delivers ~2 SD gains over conventional instruction. Most "evidence-based" techniques are attempts to approach this asymptote without the staffing cost |
| **Feedback science** (Hattie & Timperley 2007; Wisniewski et al. 2020) | Type / timing / specificity of feedback dominate effect sizes |
| **Pre-testing effect** (Carpenter, Richland) | Testing *before* studying primes attention. Distinct from retrieval practice |
---
## Cross-walk to CORE architecture
This is the load-bearing section: not "what does the literature
say" but "what does the literature say that maps onto a structural
move CORE could make."
| Pedagogy concept | CORE analog | Status |
|------------------|-------------|--------|
| Retrieval practice | `teaching/correction.py` + vault recall | **Structurally aligned.** Every reviewed correction IS a retrieval+strengthen event. Exact-recall ceiling already met. |
| Spaced retrieval | (none) | **Genuine gap.** No deterministic spaced re-verification of past corrections. → ADR-0129 |
| Interleaving | Cross-pack chains (ADR-0064 / 0067) | **Aligned.** Cross-pack chains force discrimination across domains. |
| Metacognition / calibration (prediction vs outcome) | (none at teaching layer; partial at runtime via ADR-0035) | **Genuine gap.** No prediction-vs-outcome capture in teaching loop. → ADR-0130 |
| Cognitive apprenticeship | Ratified packs as articulated expert ontology | **Strong analog.** Packs ARE the encoded expert representation; ratification IS the "fade scaffolding" step. |
| Worked examples → fading | Teaching corpora → unsupervised generation | **Partial.** Corpora encode correct answers; less so the reasoning chain that produced them. Could be more first-class. |
| Productive failure | Adversarial generation (ADR-0119.5) | **Different mechanism.** Adversarial generation is rejection; PF would mean attempt-before-grounding. Inverse of `wrong==0`. Not recommended for direct port. |
| Pre-testing | (none) | Genuine gap. CORE always grounds before articulating; never the reverse. Adopting would conflict with `wrong==0`; not recommended. |
| Self-explanation | `SolutionTrace` provenance chain | **Structurally present.** Every answer has its derivation. Could be more first-class in teaching-store records. |
| Cognitive load theory | Substrate hierarchy: algebra → field → vault → realizer | **Implicit alignment.** CORE's layering matches CLT separation of intrinsic structure from extraneous load. |
| Expertise reversal | Pack-tier promotion (ADR-0120) | **Already encoded.** The `apprentice / audit-passed / expert` contract already knows that what helps an apprentice can ossify an expert. |
| Desirable difficulties | `wrong == 0` discipline | **Inverse mapping.** CORE refuses *undesirable* difficulty (confabulation under uncertainty). A teaching-side concept of *desirable* difficulty (challenging-but-not-impossible curriculum sequencing) is not yet first-class. |
| Feedback science | `teaching/review.py` | **Partially aligned.** Reviewed corrections ARE structured feedback. Timing / specificity dimensions could be more first-class. |
---
## The two structural gaps worth addressing
Distilled from the cross-walk, two design moves are both
*pedagogically supported by robust literature* AND *consistent
with CORE's existing determinism + provenance discipline*:
### Gap 1 — Spaced reviewed-correction replay
**Mapped to:** retrieval-with-spacing literature (most robust
finding).
**ADR:** [ADR-0129](../adr/ADR-0129-spaced-correction-replay-deferred.md)
**Status:** Deferred.
**Summary:** Periodic deterministic re-run of past reviewed
corrections to verify they still produce intended outcomes
under current state. Defense against silent regression as the
correction store and pack set evolves.
### Gap 2 — Pre-articulation calibration logging
**Mapped to:** metacognitive calibration / prediction-outcome
comparison literature.
**ADR:** [ADR-0130](../adr/ADR-0130-pre-articulation-calibration-deferred.md)
**Status:** Deferred.
**Summary:** When a correction is proposed, log CORE's
pre-correction prediction; on acceptance, emit the gap.
Provides empirical answer to "is CORE actually getting better"
across pack-version cohorts; supports operator triage.
---
## What is NOT proposed (and why)
| Considered | Rejected because |
|------------|------------------|
| Adopt productive-failure mechanism inside CORE | Inverse of `wrong==0`; would require structural justification beyond pedagogy literature. Adversarial generation (ADR-0119.5) covers the related "wrong-answer rejection" use case without the conceptual conflict. |
| Adopt pre-testing in articulation | Same conflict with `wrong==0`. CORE grounds before articulating by design. |
| Add embodied / sensorimotor layer | No body. The structural analogy (substrate as "grounding") is too weak to load-bear. |
| Add peer-learning multi-agent loop | Out of scope. Multi-agent coordination is a separate architectural question; not driven by this pedagogy review. |
| Adopt cognitive-load-theory load-balancing in realizer | Already implicit in the substrate hierarchy. Making it more explicit risks decoration without integration. |
---
## Why both ADRs are deferred, not accepted
Both ADR-0129 and ADR-0130 are **proposed but deferred**, following
the established ADR-0121 / ADR-0122-deferred pattern. The deferral
reasons compose:
1. **Path-B uncertainty.** The active GSM8K-math arc
(ADR-0126 / 0127 / 0128) may resolve to a benchmark
re-targeting. If so, the correction-store population
characteristics change, and the right cadence (ADR-0129) /
cohort structure (ADR-0130) may differ.
2. **No observed incident.** Neither ADR has a triggering
incident. They're defensive infrastructure — useful if a
regression occurs (0129) or calibration drift develops (0130),
but speculative without that evidence.
3. **Cost/benefit unmeasured.** Both add telemetry volume and
operator review surface. Worth it only if the signal proves
load-bearing.
4. **Composition argument.** If either is un-deferred, the other
should be re-evaluated jointly — spaced-replay events
naturally yield calibration evidence; the two share
infrastructure. Deferring both together preserves that
composition.
The exit criteria for un-deferral are documented in each ADR's
"Exit criteria for un-deferral" section.
---
## Sequencing recommendation
1. Land ADR-0126 (PR #161) — architecture.
2. Land ADR-0127 (Gemini in flight) — units pack.
3. Land ADR-0128 (Opus #2 in flight) — numerics pack.
4. Re-run train sample with both packs mounted → real Path-A vs
Path-B verdict.
5. If Path A: continue along the math expert promotion path.
ADR-0129 / 0130 remain deferred until an incident or
bandwidth pressure surfaces them.
6. If Path B: benchmark re-targeting becomes the work; ADR-0129 /
0130 may become more relevant if the new benchmark's
correction-store characteristics are different enough to
warrant proactive verification.
---
## Reference list (additional to the original document)
- Sinha, T. & Kapur, M. (2021). When Problem Solving Followed by
Instruction Works: Evidence for Productive Failure.
[SAGE](https://journals.sagepub.com/doi/full/10.3102/00346543211019105)
- Pan, S. C. & Rickard, T. C. (2018). Transfer of test-enhanced
learning: meta-analytic review and synthesis. *Psychological Bulletin*.
- Glaser, J. & Richter, T. (2025). The Testing Effect in the
Lecture Hall: Does it Transfer to Content Studied but Not
Practiced? [Teaching of Psychology](https://journals.sagepub.com/doi/10.1177/00986283231218943)
- Cognitive Research: Principles and Implications (2024). Far
transfer of retrieval-practice benefits: rule-based learning
as the underlying mechanism.
[Springer](https://cognitiveresearchjournal.springeropen.com/articles/10.1186/s41235-024-00598-y)
- Machery, E. (2024). The Replication Crisis in Embodied Cognition
Research. *Routledge Handbook of Replication*.
[Taylor & Francis](https://www.taylorfrancis.com/chapters/edit/10.4324/9781003322511-50/replication-crisis-embodied-cognition-research-edouard-machery)
- Frontiers in Education (2026). Embodied cognition in STEM
learning: an integrative review.
[Frontiers](https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2026.1811569/full)
- Sinha & Kapur (2023). Prior math achievement and inventive
production predict learning from productive failure.
[PMC](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10185511/)
- Bloom, B. S. (1984). The 2 Sigma Problem.
*Educational Researcher 13(6)*.
- Hattie, J. & Timperley, H. (2007). The Power of Feedback.
*Review of Educational Research 77(1)*.
- Wisniewski, B., Zierer, K., Hattie, J. (2020). The Power of
Feedback Revisited: A Meta-Analysis.
- Ericsson, K. A., et al. (1993). The Role of Deliberate Practice
in the Acquisition of Expert Performance.
*Psychological Review 100(3)*.
- Sweller, J., van Merriënboer, J. J. G., Paas, F. G. W. C. (1998).
Cognitive Architecture and Instructional Design.
---
## Open questions surfaced (not resolved this session)
These are noted for future sessions; not items I'm advocating
for action:
1. **Should teaching-corpus records carry "why" structure, not
just "what"?** Self-explanation literature suggests reasoning
chains in corpora may be more useful than answers alone.
`SolutionTrace` already exposes provenance; pushing this into
teaching corpora is a separate question.
2. **Is there a deliberate-practice analog at the curriculum
level?** ADR-0120's promotion contract already encodes
"stretch-but-pass" structure (correct_rate ≥ 0.60 floor).
Whether sub-curricula should also encode this is open.
3. **Could the pack-mutation-proposal pathway adopt a worked-
example pattern?** When a pack mutation is proposed, today
the operator sees the diff; could they also see a small
worked example showing the behavioral implication?
Speculative.
4. **Is Bloom's 2-sigma a meaningful target for CORE?** A
deterministic engine with exact recall has structural
properties that may exceed 1:1 tutoring on some axes
(consistency, replay) while underperforming on others
(adaptation, social affordances). Whether to claim this
target is an architectural framing question, not a
technical one.
---
## End-of-session state
- **ADRs added:** 0129 (deferred), 0130 (deferred).
- **Session note:** this file.
- **Branch:** `docs/pedagogy-review-and-teaching-backlog`.
- **PR plan:** single docs-only PR for the three files; lands
independently of the in-flight ADR-0126 / 0127 / 0128 chain.
- **No code changes.** No regression risk.