Writing
14 min readResearchLearningEducationEdTech

How people actually learn: the evidence, the effect sizes, and what to build

A working reference on retrieval practice, spacing, interleaving, cognitive load, feedback, and assessment — with sources next to the claims and the unverified figures flagged.

ShareShare
0 views

This is the reference I compiled before writing an SAT course, and it's published in the same state I use it in: claims with sources next to them, confidence tags, and an explicit list at the end of every figure I could not verify well enough to rely on.

How to read the tags. Robust — multiple meta-analyses, replicates in real classrooms, survives its critics. Solid — meta-analysed, with known moderators. Contested — real in the lab, disputed in magnitude or generality. Myth — popular, and unsupported as usually stated.

Figures were checked against abstracts, meta-analysis records and primary PDFs where those were reachable. Where a number rests on a single secondary description or on a vendor's own marketing, it says so.

1. Retrieval practice

Retrieving something from memory beats being exposed to it again.

FindingValueSource
Lab: testing vs restudyg = 0.50, 159 studies, 81% favour testingRowland 2014
Transfer to new formatsd = 0.40; d = 0.58 across formatsPan & Rickard 2018
Classroomg ≈ 0.3–0.5Yang et al. 2021

Confidence: Robust. Dunlosky et al. (2013) rated practice testing "high utility" — one of only two techniques out of ten to earn that.

The crossover is the operationally important part. Roediger & Karpicke (2006): at five minutes, restudy wins. At one week, the tested group recalls roughly 1.5× as much. Any experiment that measures learning at the end of a session will conclude retrieval practice doesn't work.

Note also that transfer (d = 0.40) is smaller than verbatim retention (d = 0.50). Retrieval practice on SAT vocabulary transfers well. Retrieval practice on "the five levers of a DTC P&L" transfers less well to actually running one. Don't promise past what that supports.

Where it fails

Practice success below ~50% with no feedback kills the effect. Effects strengthen as success rises, particularly above 75%.

Complex, highly interconnected material may attenuate it. Van Gog and Sweller argue the effect shrinks as element interactivity rises; Karpicke and Aue rebut. Contested. Expect it to work best on facts, formulas and short procedures — which is most of SAT prep — and less reliably on integrated strategic judgement.

2. Spacing

Distributed practice was the other high-utility technique in Dunlosky. Robust, hundreds of studies across a century.

The useful subtlety, from Cepeda et al. (2006, 2008): the optimal gap is not a fixed fraction of the retention interval. It shrinks proportionally as the retention interval grows. A gap that's right for an exam next week is wrong for one in three months, and not by a constant.

This is the argument for a scheduler over a rule of thumb. "Review after 1, 3, and 7 days" is a guess that happens to be right for one horizon.

3. Interleaving

Mixing problem types rather than blocking them. Rohrer et al. ran a randomised classroom trial in real maths lessons: d = 0.83 on a delayed test.

Solid, with two hard moderators:

The material has to be confusable. The mechanism is discrimination — learning which problem is which. Interleaving unrelated topics buys nothing.

It has to be gated behind basic competence. Interleaving before the individual procedures are learned is just noise.

And learners dislike it. Both spacing and interleaving feel worse and test better — which is why satisfaction ratings cannot be allowed to drive scheduling decisions.

4. Spaced repetition scheduling

The modern answer is FSRS, which replaced SM-2's fixed ease multipliers with a three-variable model per item per user:

Scheduling inverts the forgetting curve for a chosen target retention, rather than multiplying the previous interval by a constant. Difficulty updates on each review with mean reversion toward the item's initial difficulty, so one bad day doesn't permanently mark an item as hard. Stability gains are largest when retrievability is low at review time — the model formalises the desirable-difficulty finding rather than restating it.

Anki introduced FSRS in 23.10. Parameter counts have grown across versions (19 in v5, 21 in v6). The widely repeated claim that it produces 20–30% fewer reviews than SM-2 for equal retention is a community figure — what the published benchmark measures is prediction accuracy, which is not the same claim.

The engineering consequence dominates everything else here: you must store the complete review log from the first user. Item, timestamp, grade, elapsed interval, scheduler state at the time. Per-user parameter fitting cannot be done retroactively. It is the one genuinely irreversible decision on this list.

5. Cognitive load and multimedia

Working memory is the bottleneck. Sweller's distinction: intrinsic load is the material's inherent difficulty, extraneous load is imposed by how it's presented, germane processing is the useful work.

The single most actionable result is the coherence effect: adding interesting-but-irrelevant material reduces learning, at around d = 0.70. Not "adds nothing" — subtracts. This includes the good anecdote that doesn't serve the point.

Also well supported: spatial contiguity (put the label on the diagram, not in a legend), segmenting (learner-paced chunks), and the worked-example effect — novices learn more from studying worked solutions than from solving problems, and that reverses as expertise rises. Which is why examples should fade: fully worked, then missing the final step, then the final two, then a problem.

Caution on Mayer's effect sizes generally. Many derive from single-lab, short-lesson, immediate-test studies. Directionally sound; treat the magnitudes as indicative.

Dual coding deserves a specific warning. It's well supported as cognitive psychology and weak as the instructional prescription usually drawn from it. Dunlosky rated imagery for text learning low utility. "Add a picture to the slide" is not what the theory says.

6. Feedback

The finding that should be on the wall: Kluger & DeNisi (1996) meta-analysed 607 effect sizes and found that over a third of feedback interventions decreased performance.

The mechanism is where attention lands. Feedback directed at the task helps. Feedback directed at the self — praise, ability attribution, comparison to others — hurts, because it moves attention off the work.

So: "Correct." "Not quite — here's why." Never "you're a natural at this" and never "you're struggling with algebra".

Hypercorrection is the most useful positive result: errors made with high confidence are corrected more reliably than errors made with low confidence. Capture confidence alongside each answer — one tap, three levels — and confidently-wrong items become the highest-value thing to route into review.

7. Metacognition

Learners are poor judges of their own learning, and the direction of the error is consistent: fluency is mistaken for knowledge. Material that reads easily feels learned. Judgements of learning made immediately after study are badly calibrated; made after a delay they improve considerably.

Practical consequence: a confidence rating collected at the moment of study is nearly worthless. Collected at review, it's diagnostic.

8. Motivation and completion

The blunt baseline: across large-scale MOOC data, around half of enrolments never open a single lesson, and completion showed no improvement over six years of the format maturing.

What has the best evidence-to-effort ratio anywhere in this document: implementation intentions. A single prompt asking the learner to specify when and where they'll study — "I'll study at 7pm at the kitchen table" — shows around d = 0.65 on follow-through. It is an afternoon of engineering.

Deadlines and checkpoints work through the same mechanism. Self-paced with no dates is where completion goes to die.

Streaks and reminders are useful as commitment devices. The engagement multiples quoted by the apps that use them are first-party and not causally interpretable — a person who maintains a streak was already going to show up.

9. Assessment

For anything that predicts a score, the maths matters.

Item response theory models the probability of a correct answer as a function of ability. The 2PL adds a discrimination parameter to difficulty; the 3PL adds guessing (about 0.2 for a four-option question, not 0.25 — weak examinees are drawn to good distractors).

Item information is what adaptive testing is built on: I(θ) = a² · P(θ) · (1 − P(θ)). It peaks at an item's own difficulty and scales with the square of discrimination. A highly discriminating item measures precisely, but only near its own difficulty. Summing across items gives the test information function, and the standard error is one over its square root — which tells you exactly where your item bank is thin.

Start classical, though. Before you have volume, two statistics do most of the work: the proportion correct as a difficulty proxy, and the point-biserial correlation as a discrimination proxy. Flag anything below 0.2 and delete anything negative. A negative point-biserial means strong students are getting it wrong, which nearly always means the item is mis-keyed or ambiguous. This one check finds more broken questions than any review process.

How the digital SAT adapts

Multistage adaptive testing, not item-level adaptivity (College Board). Two separately-timed modules per section. Module 1 is identical for everyone; performance routes you to an easier or harder Module 2, independently per section; scoring accounts for which path you took.

MST rather than full CAT because it's robust to one fluky early answer, permits review within a module, and needs a smaller item bank.

For an SAT prep product this is a requirement rather than a design choice. Linear practice tests produce score estimates that don't correspond to the real instrument, and score prediction is the entire value proposition.

Reliability and validity

Report standard error in score units. "1280 ± 40" is actionable; "α = 0.91" is not — and Cronbach's alpha equals reliability only under assumptions real tests violate, in which case it underestimates. There is a suspicious pile-up of published alphas at exactly .70.

The highest-value measurement available to a prep company: collect students' actual reported scores after the real exam and regress them on final practice score. That produces an honest marketing claim, a calibration curve, and a moat, because almost nobody does it.

10. Advanced techniques worth building

Pretesting. Questions asked before teaching improve learning of that content even when nearly every answer is wrong — provided the answer is then studied. A 2025 multilevel meta-analysis found g = 0.66 on prequestioned information and g = 0.01 on non-prequestioned information from the same lesson. Sharply targeted, with no spillover, so prequestion exactly what you most want retained. Solid.

Generation effect. Producing beats reading: d = 0.40 across 445 effect sizes. Robust. Counterintuitively, heavily constrained generation outperforms open-ended — which favours cloze deletions over free response.

Self-explanation. Bisra et al. (2018), 69 effect sizes, g = 0.55. Prompting learners to explain beat providing explanations. Periodic prompts beat one at the end. Badly undervalued, and free: it doesn't need grading.

Elaborative interrogation. d ≈ 0.42–0.56 — but the moderator is prior knowledge and it can flip the sign. Learners without it generate wrong explanations and consolidate them. Asking "why?" of a total novice is a misconception factory.

Transfer-appropriate processing. Morris, Bransford & Franks (1977): semantic encoding beat rhyme encoding on a standard test, and rhyme encoding beat semantic on a rhyming test. Memory depends on the match between encoding and retrieval, not on absolute depth. Robust, and the most under-applied principle here.

Deliberate practice. Macnamara et al. (2014) found it explained 26% of performance variance in games, 21% in music, 18% in sports, 4% in education and under 1% in professions. Ericsson rebutted that they had broadened the construct well past his definition. Contested. What survives is the components — work at the edge of ability, immediate informative feedback, decomposition, repetition with correction. What doesn't survive is any claim that N hours produces expertise.

11. What does not work

Learning styles. Pashler et al. (2008) tested the meshing hypothesis — that instruction works best when matched to a stated style preference. The evidence bar is a crossover interaction, and essentially no methodologically adequate study cleared it. Myth. A learning-styles quiz will convert well and make the product worse.

Re-reading, highlighting, summarisation, keyword mnemonics, imagery for text. All rated low utility by Dunlosky, and all among the most-used student strategies. The mechanism of failure is fluency: they raise ease, ease raises confidence, confidence ends study. Fine as comfort features; never as study methods, and never counted toward progress.

The learning pyramid / Dale's Cone. Dale's 1946 cone contained no percentages and made no retention claims. The percentage version is attributed to a study that has never been produced and is described as lost. The round numbers are the tell. Myth. Never cite it.

70-20-10. A 1980s exercise asking around 200 executives to recall where they had learned things. Retrospective self-report, non-representative sample, not a measurement of learning. Myth as a quantitative model; the qualitative version needs no numbers.

The Ebbinghaus curve as usually drawn. The curve is real and has been replicated — with an interesting upward jump around 24 hours, consistent with sleep consolidation. What isn't supported is applying its specific percentages to arbitrary content. Ebbinghaus used nonsense syllables and one subject. The shape generalises; the numbers don't.

12. What to build, ranked

Ordered by evidence strength divided by build cost.

Tier 1 — before anything else

  1. Mandatory correct-answer feedback on every question. Trivial.
  2. Store the complete review log from day one. Low cost, irreversible if skipped.
  3. Retrieval-first lessons — quiz, teach, quiz.
  4. Implementation-intention onboarding. Best evidence-to-effort ratio in this document.
  5. Coherence pass — delete decorative media and tangents.
  6. Segmented lessons and a daily volume cap.
  7. Item screening — flag point-biserial below 0.2, delete anything negative.
  8. Self-explanation prompts after worked examples.
  9. Ban self-level feedback in all copy.

Tier 2 — real engineering

FSRS scheduling with per-user fitting · confidence capture feeding hypercorrection routing and a calibration chart · interleaving within confusable families, gated by mastery · faded worked examples driven by mastery · MST practice tests mirroring the digital SAT's routing · delayed retention as the primary product metric · deadlines and checkpoints.

Tier 3

Collecting real SAT scores and regressing them on practice scores — the highest-value business item here · misconception-tagged distractors · IRT calibration once volume allows · test-out for advanced learners · progress bars over mastery rather than content consumed · live cohorts.

Tier 4 — do not build

Learning-styles assessment · leaderboards · participation badges · any marketing citing the learning pyramid or 70-20-10 · watch time or lessons completed as a headline outcome metric · multiple choice without feedback, anywhere, ever · a raw retention-target slider exposed to learners.

The one-paragraph version

Build a platform where every piece of content is followed by retrieval practice with mandatory feedback; where a recall-probability model decides what to show and when, rather than a fixed multiplier; where the practice format matches the criterion test exactly; where worked examples fade to problems as mastery rises; where confidence is captured and confidently-wrong answers get the strongest correction; and where the metric being optimised is delayed retention rather than practice accuracy or completion percentage. Then add an implementation-intention prompt at onboarding, delete every decorative element, and put a deadline on everything.

Most of Tier 1 is about a week of work.

Flagged as unverified

These are the figures I could not confirm against a primary source. They are leads, not facts, and nothing published should depend on them.

  1. Roediger & Karpicke's exact one-week percentages (the commonly quoted ~61% vs ~40%). The 1.5× ratio is well attested.
  2. Yang et al. 2021's classroom effect size — sources give both 0.50 and 0.33.
  3. Cepeda et al.'s headline effect size. Study counts verified; magnitude not.
  4. The exact MOOC certification rate (~3%). "Half never start" and "no improvement over six years" are attested.
  5. Cohort versus self-paced completion rates — vendor first-party data.
  6. Streak-engagement figures — first-party, and not causally interpretable.
  7. "FSRS produces 20–30% fewer reviews than SM-2" — community claim; the benchmark measures prediction accuracy.
  8. FSRS parameter count for v4.5. The v5 and v6 counts are attested.
  9. FSRS stability-increase and post-lapse formulas — the difficulty-update and same-day formulas were corroborated verbatim; verify the rest against the reference implementation.
  10. IRT calibration sample sizes (roughly 200, 500, and 1,000 responses per item for 1PL, 2PL, 3PL).
  11. Mayer's effect sizes generally — largely single-lab, short-lesson, immediate-test.
  12. The specific study reported as showing elaborative interrogation worsening learning of statistical probability.

The essay version of this, with what it meant for my own studying, is here.

More writing