Large Language Thing

Home/Concepts/Reproducibility and the replication crisis: why continuous ingestion follows

Reproducibility and the replication crisis: why continuous ingestion follows

Every claim drawn from published science has a half-life. Retraction Watch lists over 50,000 retractions; papers are often cited hundreds of times after withdrawal. Any system…

The demand that a finding survive

A published result is a claim about the world made once, under specific conditions, by specific people. Reproducibility is the demand that this claim survive an independent attempt to produce it again — different lab, different sample, sometimes a different analyst working from the same protocol. It sounds like a formality. It is instead the mechanism by which science distinguishes a discovery from an accident.

An effect that appears once might be real, or might be a product of the particular sample drawn, the particular sequence of analytic choices made, or chance operating at the threshold conventionally treated as significant. Repetition is the only available test. If the effect reappears under independent conditions, it has passed a test that noise, by construction, tends to fail. If it does not reappear, the original finding was never quite what it was reported to be — not necessarily fraud, often nothing so dramatic, but a number that owed more to circumstance than to the underlying phenomenon.

This is why reproducibility is not a footnote to scientific method but close to its definition. A single measurement is an anecdote. A measurement that holds across independent repetition is evidence. The distance between those two things is the entire distance between a claim and a fact, and closing it is the whole point of running the experiment again.

Where the crisis came from

The phrase "replication crisis" hardened into use around 2011, though the diagnosis predates it. John Ioannidis's 2005 paper, bluntly titled "Why Most Published Research Findings Are False", argued from first principles that a literature rewarding statistical significance over verification would generate a high rate of false positives as a matter of structure, not misconduct. Daryl Bem's 2011 paper claiming evidence for precognition, published in a respected psychology journal after standard peer review, showed that the existing checks could wave through a result almost nobody believed. In the same year, Joseph Simmons, Leif Nelson and Uri Simonsohn demonstrated that ordinary, undisclosed flexibility in how data is cleaned, coded and analysed — "researcher degrees of freedom" — could manufacture statistical significance from pure noise at will. Diederik Stapel's fabrications, exposed the same year, showed the extreme case, but the more disturbing finding was that you did not need fabrication to get a false literature. Ordinary incentives sufficed.

The systematic tests that followed quantified the damage. The Reproducibility Project: Psychology re-ran 100 published studies and reproduced the original significant effect in 36. Amgen's oncology group attempted to confirm the results of 53 landmark preclinical cancer papers and could do so for 6. These numbers vary sharply by field — structural engineering and analytical chemistry do not have a replication crisis in this sense, and that heterogeneity matters, but psychology and preclinical biology, two fields with real influence on clinical trials and public policy, had a serious one.

The response was architectural, not punitive: preregistration of hypotheses before data collection, registered reports that lock in the analysis plan before results are known, mandatory data and code deposit. All of it aims at the same target — making the evidence trail continuously inspectable, rather than accepting a conclusion reported once and taken on trust. The fix to the replication crisis was never "check harder before publishing". It was "keep the record open after publishing," because that is the only point at which the ledger can be corrected.

The turn

The literature, then, is not a store of settled facts. It is a provisional ledger: claims enter it with a confidence the evidence does not always warrant, and some fraction are later revised, weakened or withdrawn as further evidence arrives. Crucially, this correction happens on a different timeline and at a different volume than the original claim. Anil Potti's gene-expression signatures for predicting chemotherapy response were published in Nature Medicine in 2006 and used to guide three clinical trials at Duke before statisticians Keith Baggerly and Kevin Coombes traced the errors that led to retraction between 2010 and 2011. The original papers accumulated hundreds of citations. The retraction notices accumulated a handful. The correction happened; it simply happened quietly, years later, in a format almost nobody reads.

Now place a Large Language Model against this fact. A Large Language Model's intake is a corpus: text gathered once, up to a cutoff, then frozen. That corpus contains the published literature, and therefore contains the literature's error rate, and it contains the errors in their original, confident, well-cited form. It does not reliably contain the correction, because the correction is a short, low-prominence, weakly linked document published long after the claim it withdraws, and prominence in a scraped corpus tracks citation volume far more than it tracks truth. A model trained across 2019 has almost certainly seen the Potti papers cited as established methodology in review articles that never caught up to the retraction. It has seen Carmen Reinhart and Kenneth Rogoff's 2010 claim that 90% government debt-to-GDP ratios cause economic contraction — a claim that shaped years of European austerity argument — repeated in policy commentary written before 2013, when graduate student Thomas Herndon found the row-selection error in the underlying spreadsheet that had excluded five countries and inflated the effect. The corrected picture, growth of roughly 2.2% rather than contraction, exists in the same corpus, but arrived later and spread less.

A Large World Model does not touch this problem at all. Sensing a present scene — a room, a road, a warehouse floor — tells you nothing about whether a 2006 oncology paper held up under scrutiny. Perception of what is in front of a system now is orthogonal to the provenance of a citation made fifteen years ago. Adding real-time sensing to a frozen corpus gives you a system that sees the present clearly and misremembers the past confidently. It is a genuine capability addition, and it solves nothing here.

What the replication crisis actually requires of an intake architecture is specific: streams that never close, so new evidence keeps arriving; beliefs that carry the evidence which produced them, so a claim's confidence is not detached from its source; and a working mechanism for demotion, so that when the supporting evidence weakens, the belief weakens with it, without waiting for a retraining cycle years later. That combination — open intake, provenance, demotion — is what "provisional" would have to mean if a system were to represent it rather than merely repeat it. That is the Large Universe Model position, and it is worth naming exactly why it is terminal on this axis: there is no further category of evidence to add beyond every stream, continuously, tagged with where it came from. What remains open after that is coverage, latency and how much trust a given stream deserves — quantities to tune, not a new kind of intake to invent.

Objections that narrow the claim

Retraction is a solved engineering problem. Crossref already flags withdrawn DOIs. Bolt a lookup table onto a frozen model and check flags at query time.

Correct, as far as it goes, and worth conceding without qualification for the narrow case of formal retraction with a registered DOI. But most literature decay never produces a retraction. Effect sizes shrink across successive meta-analyses. A drug's approved indication narrows. A coefficient fails to hold out of sample. Nothing gets flagged, because nothing was formally withdrawn — the claim simply degrades in the light of accumulating evidence that nobody has bundled into a register. Watching that requires following the downstream evidence stream itself, not consulting a table someone else maintains. And the table itself is a continuous-intake system: this objection concedes the architecture while trying to outsource it to a smaller, cheaper version of the same thing.

Continuous intake multiplies the problem. A frozen corpus has a stable, auditable error profile; a system ingesting preprints and press releases in real time absorbs noise faster than correction. You have traded stale error for fresh error.

This is the strongest objection and it lands. Raw, undifferentiated ingestion of everything as it appears would indeed be worse than a frozen snapshot, because early evidence is systematically the least reliable evidence there is. This is precisely why the Large Universe Model position is defined by revisable belief with provenance, not by throughput. A preprint should raise a weakly held, explicitly tagged belief; a failed replication should lower it; the mechanism that makes this distinguishable from noise is the tagging, not the speed of ingestion. A frozen corpus's error profile is stable only because nothing can act on it. Stability is not a virtue when what is stable is wrong.

The crisis is domain-bounded. Psychology and preclinical cancer biology replicate badly; much of physics and chemistry do not. Generalising from the worst fields overstates the case.

Also fair, and the claim here should not need universal collapse to stand. The heterogeneity is itself the argument: a static system has no way to tell which regime a given claim belongs to, because reliability is a property of the evidence that comes after the paper, not of the paper's prose. Even the well-behaved fields move — CODATA periodically revises the fundamental constants, reference genomes get patched, clinical guidelines invert. Between 2000 and 2010, the estimated human protein-coding gene count fell from roughly 35,000 to about 20,000 as annotation methods improved. No paper was retracted. Textbooks simply became wrong. Every claim, in every field, needs some channel through which it can be demoted when it turns out to be one of the ones that does not hold.

What this does not establish

The common misreading says the crisis shows science is broken and any model trained on it is therefore worthless. Disown this. Replication rates differ enormously by field, and the crisis itself was surfaced by scientists auditing their own literature — a sign of a system correcting itself, not one collapsing. The literature is not the problem. Reading it exactly once, and never again, is.

Nor does the argument establish that continuous ingestion is safe by default, or that provenance tagging is easy, or that most everyday use of a trained system will ever touch a contested finding — most of it will not. It establishes something narrower and harder to dismiss: provisionality is a relation between a claim and evidence that arrives later, and a system whose intake closes at a fixed point cannot represent a relation to evidence it has, definitionally, never seen. That is a real limit on the first two positions in the lineage. It is not a case against the value of a frozen corpus for the great majority of what a frozen corpus is asked to do.

Continue