Everything that feels like studying isn't
I got an A* using techniques the research rates as low utility. Then I read the research, and had to rebuild the course I was writing.
I am building an SAT course. Before writing the lessons I did what I should have done years earlier and went through the actual literature on how people learn — meta-analyses, effect sizes, the arguments where the researchers disagree with each other.
It went badly for me personally. Most of what I did to get an A* in Economics is rated low utility by the people who study this for a living.
The thing I got wrong for two years
I re-read. I highlighted. I made summary notes and then made better summary notes.
All three are in the bottom tier of Dunlosky et al.'s review — the one that graded ten common study techniques on the strength of the evidence behind them. Re-reading, highlighting, and summarisation all came out low utility. They are also, by a distance, the three most popular things students do.
The failure isn't that they do nothing. It's the mechanism by which they do nothing, which is worth understanding because it generalises well beyond studying.
Re-reading raises fluency — how easily the material goes down. Fluency feels like understanding. It is a completely different thing, and the brain does not distinguish them. So you read the chapter a third time, it flows, you feel you know it, and you stop.
The technique's failure mode is that it terminates study early while making you confident. It doesn't just waste the hour. It costs you the hours afterwards.
I wrote in another piece that I spent three months getting better at writing essays I thought were excellent and had no evidence about. I framed it then as a feedback problem. It was also this problem. I had no external check, and my internal one was reading fluency, which is calibrated to say yes.
The finding that makes self-teaching hard
Here is the result that reorganised how I think about all of it.
Testing yourself on material beats re-reading it — around g = 0.50 across 159 studies, with 81% of them favouring testing. That much is well known.
The part that isn't well known is the timing. Roediger and Karpicke found that at five minutes, re-reading wins. At one week, the tested group recalls roughly one and a half times as much.
There's a crossover. The method that works looks worse right up until it doesn't.
The same shape shows up everywhere in the good techniques. Spacing your practice feels less productive than a block of it. Mixing problem types up feels like chaos next to doing twenty of the same kind — and beats it by d = 0.83 in a classroom trial. Researchers call these desirable difficulties, which is a polite name for "the effective option is the one that feels bad".
Now put that next to a student sitting alone deciding what to do tomorrow, using how today felt as the input.
The signal is inverted. Not noisy — inverted. Sessions that felt productive were often the wasted ones, and the sessions where I kept blanking on things I'd read three times were the ones that worked. Left alone with your own judgement, you will reliably select against the techniques that work.
That is the actual case for structure, and it is not the one usually made. It isn't that people are lazy. It's that the instrument they're using to steer reads backwards.
The one thing I did right, by accident
Every fortnight I sat a real past paper under timed conditions, whether or not I felt ready.
I did that because I needed a deadline I couldn't negotiate with. It turns out to be three separate evidence-backed things at once, and I got all three without knowing any of them.
It was retrieval rather than review. It was spaced, by two weeks. And it had format fidelity — same paper, same clock, same conditions as the thing I was actually going to be assessed on.
That last one is the most under-applied idea in the whole field. Morris, Bransford and Franks showed that memory depends on the match between how you encoded something and how you're asked to retrieve it, not on how deeply you processed it. Semantic encoding beat rhyme encoding on a normal test — and rhyme encoding beat semantic encoding when the test was about rhymes.
So: the SAT is four-option multiple choice, under time pressure, on a screen, in Bluebook. If your practice is flashcards building free recall of vocabulary, you are training a retrieval process the exam never asks for. It will feel like progress. Some of it even is. It is not the one being marked.
A lot of test prep fails exactly here, and it's invisible, because "we cover all the content" is true at the same time.
Three things I've now removed from the plan
A learning-styles quiz. I had one sketched. It's a great onboarding step — everyone believes in visual and auditory learners, so everyone completes it and feels seen. Pashler et al. went looking for studies meeting the minimum bar to test it and found essentially none supporting it. Worse than useless: it justifies giving the "visual learner" less text, when what should determine the format is the content and the exam.
Anything citing the learning pyramid. The one with 10% of what we read, 90% of what we teach. Dale's original 1946 cone contained no percentages and made no retention claims at all. The numbers are attributed to a study that has never been produced and is described as lost. The suspiciously round figures are the tell — real data does not come out in tens.
70-20-10. Traced back, it's a 1980s exercise asking around 200 executives to recall where they felt they'd learned. Retrospective self-report from a non-representative sample. Even the ATD has published a piece asking where the evidence is. The qualitative claim underneath — you learn a lot on the job, training alone isn't enough — is obviously true and doesn't need invented numbers propping it up.
I'd have shipped at least two of these. They're in every competitor's marketing, which is precisely why they were on my list.
What the course does instead
Concretely, and none of it is expensive:
Every lesson opens with two or three questions on material that hasn't been taught yet. Getting them wrong is the mechanism, not a failure — prequestioned material comes out at g = 0.66, and the rest of the same lesson at g = 0.01. No spillover at all, which means you have to prequestion precisely what you most want remembered.
Every question shows the correct answer, always, whether right or wrong. Below about 50% success with no feedback, the testing effect stops working entirely.
Worked examples fade. Fully worked, then missing the last step, then the last two, then it's a problem.
After each one: "in one sentence, why does step three work?" Nobody grades it. Prompting the explanation beat providing one at g = 0.55, and the writing is where the benefit lives.
Scheduling is a model of your recall probability per item, not a fixed multiplier on the last interval — and the review log gets stored from day one, because you cannot fit that model to data you didn't keep.
And the metric is delayed retention. Not lessons completed, not time on site, not practice accuracy. Given the crossover, any measurement taken at the end of a session will tell you the correct features are broken.
The part I keep coming back to
The uncomfortable version of all this is that the confidence you feel after a good study session is generated by the same process that ends it.
You cannot fix that by trying harder to be honest with yourself, because it isn't a character problem. You fix it by putting something outside your own head — a real paper, a timer, a question you have to answer before you're allowed to look — and then believing that instead.
Which is what school was doing, badly, on a fixed calendar, for thirty people at once. I removed it, then spent two years rebuilding it by hand.
The full research is up as a separate piece — effect sizes, sources, the arguments researchers are still having, and the thirteen figures I could not verify well enough to quote.
More writing
One Kinase, Two Answers
Endurance training and hypertrophy don't merely compete for time. They send opposing instructions to the same signalling node, and the conflict has a name: AMPK. Here's what it actually does, how big the effect is, and how to schedule around it.
Measurement Without a Decision Is a Hobby
More data makes the signal-to-noise problem worse, not better, unless the analysis accounts for it. A counterweight to my own tracking article, and the rule that decides whether a metric earns its place.