Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Learning-Science Grounding & Joint Study/HPC Design

draft first full version, 2026-07-18. Tracked as sail-wm67.

The bandit-arms-gat chapter covers what sail-judge actually does and how its reward signal can fail. This chapter is about the other half: where the arm constructs themselves came from, why they’re the right ones (or aren’t yet), and the actual, live ask that follows from taking that seriously — a joint pre-registered study and a shared compute application, not a hypothetical.

Why these constructs, specifically

The v1 arms (curious/concrete/reflective; synthesize/contrast/extend) were rhetorical postures — plausible-sounding, ungrounded in anything about how people actually learn. The current arms replace each one with a named construct from the learning-science literature, chosen deliberately, not just relabeled:

  • Generation effect (Slamecka & Graf) — people retain material better when they generate an answer themselves rather than being shown it. Grounds the asker’s generation-effect arm directly.
  • Retrieval-practice / testing effect (Roediger & Karpicke) — the act of recalling something, not just re-exposure to it, is what strengthens memory. Grounds retrieval-practice.
  • Desirable difficulty (Bjork) — conditions that make learning feel harder in the moment can produce better long-term retention than smooth, easy practice. Grounds desirable-difficulty — and, as the bandit-arms-gat chapter’s Goodhart table already shows, is exactly the construct that breaks any live reward proxy, by definition: a proxy that rewards smooth immediate success is rewarding the opposite of what this construct is.
  • Metacognitive calibration — asking for a confidence estimate before checking an answer, tied directly to the mini-study’s own KPI (below).
  • Wood, Bruner & Ross (1976)’s scaffolding functions — marking critical features, reduction in degrees of freedom, direction maintenance — ground the discourse-driver’s three arms. This is a facilitation paper, not an instruction paper, which is part of why the discourse-driver role reads as the “mentor” half of the instructor/mentor split named in the previous chapter.

None of this is a settled taxonomy — it’s a first, citable pass, offered specifically so it can be argued with by someone whose actual field this is.

Where the current design is weakest

Two separate weaknesses, not one:

  • The trigger modelwhen either role fires — is still just a raw ‘?’-count and a fixed turn interval. It never got the same research-grounding treatment the arm-selection side did. Nobody has proposed what a learning-science-grounded triggering signal would even look like yet. The Question Mark Problem (Appendix: Artifact Series) walks through where that raw count came from and why nobody has revisited it — this weakness and that history are the same open question, not two separate ones.
  • The reward system — covered in full in the bandit-arms-gat chapter’s Goodhart table, but worth restating the punchline from a learning-science angle specifically: five of the seven arms are regressional or causal risks (a cheap proxy can drift from the real construct without anyone noticing), two are extremal risks (the proxy holds only until optimization pushes hard against it), and one — desirable-difficulty — has no honest live proxy by construction, not as an engineering gap to close. That last one can only ever be validated against a delayed, real outcome measure. Which is exactly what the mini-study below exists to provide.

Grounding further: Hummel (2025)

Sandra Hummel’s own completed study — Higher Education Under Generative AI: Biographical Orientations of Democratic Learning and Teaching, Education Sciences 15(12):1572 (n=151: 122 students, 29 lecturers, grounded-theory analysis of written articulations) — is real data to build on, not just a citation for KPI-realism. Her five reconstructed orientations: pragmatic (coping with workload), adaptive (learning under opacity), relational (authority and resonance), ambiguous (improvisation and fragility), recognition (voice and visibility) — synthesized into three axes: temporal sovereignty, epistemic opacity/accountability, recognition ecologies.

Honest status: this grounds the next arm-taxonomy pass, not the current one. Nobody has yet checked whether these five orientations map onto the asker/discourse-driver arms more directly than the generation-effect/ testing-effect/Wood-Bruner-Ross taxonomy already does, or whether they operate at a different level entirely — learner disposition rather than agent action — and should inform something else (the reward signal, or a future persona model) instead of the arms themselves. Real, unresolved, not glossed over.

The actual ask: a joint study and a compute application

Two concrete, live pieces, not a hypothetical collaboration:

  • A pre-registered scaffold-vs-answer mini-study, to Hummel’s own KPI-realism standard: d ≥ 0.4, a 2-week retention follow-up, a Zerbe-style item set, n in the dozens of vocational trainers. This is the study that would actually settle whether desirable-difficulty (and the rest of the arm taxonomy) tracks real learning gain, not just a proxy for it.
  • A joint HPC compute application, naming Hummel as PI on her ScaDS.AI Young Investigator standing (which resolves an eligibility gap a solo application would have hit), scoped to a concrete, modest ask: a short Alpha Centauri (A100) allocation for a GPT-2-scale SAITO ablation — training actual model weights via DPO to compare a downstream-competency reward against an immediate-satisfaction reward, reusing sail-judge’s own prompt templates and Goodhart-variant table as the interpretive template for whatever divergence shows up. Not the €3M version of this work — a small, real pilot sized to what a bootstrap can actually run.

This is a real, sent ask, not a pitch — the compendium link went out alongside it, this chapter is what that link points to.

Open questions

  • Do Hummel’s five orientations map onto the existing asker/discourse-driver arms, or do they belong somewhere else in the design (the reward signal, a future learner model)?
  • What would a learning-science-grounded triggering signal (replacing the raw ‘?’-count) actually look like — is there a construct as citable as the ones already grounding arm selection?
  • Is the pre-registered mini-study’s power (n in the dozens) actually sufficient to detect d ≥ 0.4 given real-world attrition over a 2-week follow-up — worth a real power calculation before this goes further, not an assumption.

Sources

Slamecka & Graf (generation effect); Roediger & Karpicke (testing effect); Bjork (desirable difficulties); Wood, Bruner & Ross (1976, scaffolding functions); Hummel, S. (2025), Education Sciences 15(12):1572 — references/hummel-2025-biographical-orientations.pdf.