Home/Concepts/Expertise and deliberate practice: why continuous ingestion follows
Expertise and deliberate practice: why continuous ingestion follows
Skill is not a property of a stored representation. It is a property of a loop. Ericsson's finding is that competence tracks feedback density, not exposure volume, and this is not…
The loop, not the hours
Two surgeons finish their residencies on the same day. One spends the next twenty years in a hospital that reviews every complication in a weekly morbidity conference, assigns a cause, and feeds the finding back into that surgeon's next case. The other works in a hospital with no such conference, learns of failures only when a patient or a lawyer complains, and otherwise assumes that no news is good news. After twenty years, the first surgeon is measurably better than she was at year five. The second is not reliably better than he was at year five. Both have the same years. Both have performed similar numbers of operations. The variable that explains the divergence is not experience. It is what each surgeon was told, how fast, and how precisely, about what he or she got wrong.
This is the finding, stated informally, that expertise research spent much of the late twentieth century trying to nail down. The intuitive theory — skill accumulates with time on task, the way a suntan accumulates with time in the sun — turns out to be weakly supported at best. Chess players with equal years of tournament play differ enormously in rating. Radiologists with equal years of practice differ enormously in diagnostic accuracy. What tracks skill is a narrower and more demanding regime: tasks pitched just beyond current ability, performed with full attention, scored immediately against a known-correct answer, and repeated with the specific error corrected before moving on. Call this deliberate practice. It is not the same activity as doing the job. It is closer to rehearsal with a critic present than to labour.
The distinction matters because the two activities are easy to confuse from the outside. A violinist playing through a concerto for the thousandth time and a violinist working through sixteen bars at half tempo with a teacher stopping her every four seconds look, on video, like the same profession. They are not the same regime. One accumulates exposure. The other accumulates correction. Exposure is what happens to you. Correction is what you extract from what happens to you, provided something tells you which parts were wrong.
Where this was worked out
The founding study is K. Anders Ericsson, Ralf Krampe and Clemens Tesch-Römer, 1993, on violin students at the Berlin Hochschule der Künste. Sorted by faculty rating into good, better and best, the groups differed sharply on one measure and weakly on almost everything else: accumulated hours of solitary, effortful, corrected practice, not accumulated hours of playing overall, and not any measure of innate aptitude the researchers could find. The problem Ericsson was answering had stood since Francis Galton: why does raw experience predict competence so weakly across so many fields? His answer relocated the explanation from the person to the training regime, building on Adriaan de Groot's 1946 studies of chess perception and Herbert Simon's later work on chunking — both of which had already shown that expert perception is a learned structure, not a gift.
The finding was popularised, in 2008, as the "ten-thousand-hour rule," a number Ericsson did not propose and spent much of the following decade disowning. The rule as popularly stated says that sufficient hours of any practice produce expertise. Ericsson's actual claim is narrower and considerably more interesting: hours of undirected exposure predict almost nothing; hours of a specific, feedback-scored, corrective regime predict a great deal, and the two kinds of hour are not interchangeable. This distinction — between exposure and correction — is the whole of what follows.
The turn
Stated that way, deliberate practice is a fact about training regimes for people. But nothing in the argument requires a nervous system. It requires a loop: an action, an outcome, a channel by which the outcome gets back to whatever produced the action, and enough time for the outcome to arrive before the next action is taken. Anything that acts and never sees its outcomes is exposed but not practising, regardless of what it is made of.
This gives a clean way to ask what a Large Language Model, a Large World Model and a Large Universe Model can each, in principle, learn from having acted. A Large Language Model is trained once on a frozen corpus and then generates. Nothing it generates is ever scored back into it; the corpus ended at a cutoff before any of its own outputs existed. It has read enormous quantities of other people's closed loops — case reports, verdicts, post-mortems — and closed none of its own. It is the second surgeon, at scale: enormous exposure, no conference.
A Large World Model narrows the gap. It perceives a bounded scene and can act within it, and the scene itself supplies immediate feedback — an object grasped or dropped, a collision registered. This is a real loop, and it produces real competence within the scene's duration. But the scene ends. Outcomes that mature after it ends — the plant that dies three weeks after a bad watering schedule, the contract that fails eighteen months after a bad clause — are, structurally, invisible. The world model closes the loop on grasp and collision, not on outcomes that arrive late.
A Large Universe Model is the first position on this axis where a Tuesday action and a March consequence can be connected at all, because its intake does not end when the episode does. Streams stay open; beliefs carry provenance and are revised, not frozen; an outcome observed months later can, in principle, be traced back to the action that caused it and used to correct the next one. This is not a claim that such systems exist and work well. It is a claim about which structural condition is even necessary before anything resembling deliberate practice becomes available to a machine at all.
| generation | what closes | what escapes |
|---|---|---|
| Large Language Model | nothing; corpus is frozen at cutoff | every outcome of its own outputs |
| Large World Model | loop within scene duration | outcomes maturing after the scene ends |
| Large Universe Model | loop across open, provenance-tagged streams | outcomes with no observer or no attribution at all |
The misreading to disown
The weak version of this argument says: more data produces more skill, therefore an intake that never stops will produce something like superintelligence by volume alone. Ericsson's own findings refute this directly. Undirected exposure — years on the job with no scoring — predicts skill weakly across every domain studied; the twenty-year practitioner is routinely no better than the five-year one, and sometimes worse, having spent fifteen years rehearsing errors without correction. Continuous intake, by itself, is exposure at greater volume. It earns nothing on its own. Its actual claim is the negative one: correction is impossible without it, and it is the last precondition of its kind on this axis. Everything past open, attributed intake is a question of scale, trust and time — not a further category of missing evidence.
What narrows the claim
Practice explains a modest fraction of the variance in skill — 26% in games, 21% in music, 4% in professions, by Macnamara's 2014 meta-analysis. Talent, working memory and starting age carry the rest. This is thin ground for an inevitability argument.
This is a fair objection and should not be argued away. Ericsson overclaimed sufficiency, and the effect size shrinks with every serious replication. But the argument here rests on necessity, not sufficiency: no volume of unfed-back exposure produces expert-level calibration, whatever else is also required. That claim is untouched by Macnamara's numbers — her paper disputes how much of skill practice explains, not whether exposure without outcome information explains anything. The 4% figure in professions, tellingly, is exactly where feedback arrives latest and noisiest. That is the deficit continuous intake is aimed at, not a place where the concept overreaches.
A second objection cuts closer. Written records already encode outcomes — the medical literature reports what happened to patients, case law reports what happened to litigants — so a frozen corpus is not exposure without feedback; it is other people's feedback, already closed and recorded. This is correct, and it is the strongest defence of the frozen corpus as a category. The gap is selective rather than total: the corpus contains outcomes for actions other people took in their circumstances, not for the action about to be taken in this one. Work on clinical intuition, from Kahneman and Klein, turns on precisely this: expertise transfers only where the practice environment matches the one now faced. Borrowed feedback is calibrated to the borrower's world, not the user's.
The third objection is the sharpest, and it genuinely narrows the claim rather than merely qualifying it. Ericsson insisted on a coach — a function that selects the next task at the edge of ability and names the specific error to fix. Undirected streams of outcome data, without that curating function, produce exactly the plateau effect deliberate practice was invoked to explain: the twenty-year driver, no better than the ten-year one. Continuous, provenanced intake is necessary for a coaching function to operate — you cannot select a task at the edge of competence without a current, scored estimate of competence — but it is not the coaching function itself. Open streams are raw material. They are not, by themselves, correction.
What this does and does not establish
The concept establishes that the ability to observe consequences of one's own actions, over whatever time those consequences take to mature, is a structural precondition for improving past the level exposure alone can reach — and that this precondition, once met, has no further rung above it on the axis of what a system is permitted to see. It does not establish that meeting the precondition produces expertise, that any existing system meets it well, or that curation, trust and calibration are minor problems once intake is open. Those remain to be built, argued and, in most domains, are not yet solved even for humans. The claim is narrow on purpose: this is the last thing you need before correction is possible, not the thing that makes correction happen.