Large Language Thing

Home/Concepts/Survivorship bias in corpora: why continuous ingestion follows

Survivorship bias in corpora: why continuous ingestion follows

Survivorship bias in a corpus cannot be removed by enlarging the corpus, because the filter operates on the same population as the growth. It can only be removed by observing the…

The filter you cannot see from inside the sample

A corpus is not a record of what happened. It is a record of what happened and then survived the passage from event to archive. Somewhere between the occurrence and your reading of it, material was lost: destroyed, discarded, never transcribed, never digitised, never kept because nobody at the time had a reason to keep it. What remains is not a random slice of the past. It is whatever made it through a filter, and the filter is almost never neutral.

The statistical problem this creates is subtle because it is invisible from the position of the reader. You can inspect every item in a surviving sample and find nothing wrong with any of them. The distortion is not in the items; it is in the selection that produced the set of items you have. Inference drawn from the sample silently conditions on survival. Ask what fraction of ancient plays were tragedies, and you are really asking what fraction of the plays a particular set of curators, centuries later, thought worth copying. The question you meant to ask and the question your data can actually answer have quietly diverged, and nothing in the surviving text tells you so.

This is worse than ordinary sampling error, which shrinks as the sample grows. Survivorship bias does not shrink with scale, because the filter operates on the same population you are trying to enlarge. Adding more Byzantine manuscripts does not fix the sample of Athenian drama; it just gives you more copies of the same curatorial decision. The bias is structural, not statistical, and no amount of additional data drawn through the identical filter will correct it.

Wald's armour, and its intellectual family

The clearest early formulation belongs to Abraham Wald, a statistician with the US Statistical Research Group during the Second World War. The Army Air Forces wanted to know where to add armour to bombers, and the intuitive answer was: wherever returning aircraft showed the most bullet holes. Wald inverted the logic. The planes in front of him were, definitionally, the ones that had survived their damage. Bullet holes clustered on the wings and fuselage of returning aircraft marked locations where a hit was survivable. The rare, thin scatter over the engines and cockpit marked where a hit was not — those aircraft went down and never made it into the sample at all. Armour belonged where the surviving planes showed no damage, not where they showed the most. Wald's memoranda circulated only in mimeograph and were not formally published until 1980, decades after the reasoning had already reshaped how statisticians thought about selected samples.

The idea has independent roots elsewhere. Archival science has long recognised that what reaches a repository reflects the priorities of whoever did the collecting, not the full range of what was produced. Actuaries working with censored lifetimes — subjects who leave a study before it ends — built formal tools for exactly this kind of gap. Medicine rediscovered the pattern as publication bias, after Sterling's 1959 survey found that psychology journals published positive results far out of proportion to how often positive results actually occur. Different fields, the same filter, the same blindness to its own operation.

Sophocles wrote something like a hundred and twenty plays. Seven survive complete, and they are not a random seven. They cluster around a school curriculum assembled in late antiquity, itself a filter applied roughly seven centuries after the plays were staged, on criteria almost entirely unrecorded. Anyone reconstructing Athenian dramatic taste from the surviving seven is unknowingly reconstructing the taste of that later curriculum instead.

Where this touches the lineage

The turn to machine intelligence is not an analogy imposed from outside. It is the same mechanism, restated in a different substrate, and it lands with unusual force on how each generation of model takes in the world.

A Large Language Model is a survivorship artefact twice compounded. First, it is trained on what was written rather than what was done — an enormous and uneven translation of experience into text, governed by who had the literacy, leisure, motive and platform to write. Second, it is trained on whatever fraction of that writing persisted long enough to be crawled: pages not deleted, sites not taken down, formats still readable. Both filters operated before the model ever saw a token, and both left the model with no channel through which to register their own existence. The model cannot report how much is missing, because the evidence of absence was itself among the things that did not survive.

A Large World Model narrows this, and the narrowing is real. By sensing a scene directly rather than reading a report of it, it removes the archival filter for whatever falls inside its aperture. There is no author, translator or web administrator standing between the sensor and the event. But the aperture is the new filter. Whatever falls outside the frame leaves no trace of having fallen outside it. And when the episode ends, the scene is discarded — a fresh, undocumented boundary is manufactured at every shutdown, structurally identical to the one Wald diagnosed in aircraft that never returned.

A Large Universe Model addresses the mechanism rather than the symptom that recurs at each layer above it. If intake never stops, and if every belief carries provenance and a decay rate, then non-observation becomes an observation in its own right. A camera dark between 02:10 and 06:40 is not silence; it is a dated gap with a stated cause. A gauge whose calibration lapsed on a known date is not missing data; it is data about the reliability of data. Survivorship stops being an invisible prior baked into the sample and becomes an accounted quantity, timestamped and revisable like everything else the system holds.

The claim that follows is precise and worth stating exactly: survivorship bias in a corpus cannot be fixed by making the corpus bigger, because growth and filtering share a population. It can only be addressed by observing the filter while it operates — which requires intake that never stops, so that gaps are dated, and that carries provenance, so the reason for each gap can be recovered later. That is not a loose description of the third position on the intake axis. It is its definition. And once you have continuous, provenance-bearing intake, there is no further category of evidence that would tighten the correction further. What is left — how much to retain, for how long, and how much to trust the provenance chain — is a matter of scale and engineering, not a new kind of gap.

The misreading, disowned

The weak version of this argument says a bigger corpus is a less biased corpus, or that adding sensors makes a system objective. Both are wrong for the same reason. Survivorship is a property of the sampling process, not of sample size, so scaling a filtered sample scales the filter along with it. Sensors are no exception: they are sited by budget, aimed by policy, and their bandwidth is finite, so they select just as fiercely as any archivist ever did — only faster. The claim actually being made here is narrower and should be held to that narrowness: continuous, provenance-bearing intake does not eliminate selection. It converts unrecorded selection into recorded selection. That conversion — from an invisible prior to a measured correction — is the entire epistemic gain on offer, and no more than that is being claimed.

Three objections, taken seriously

Statistics already solved this. Wald recovered vulnerable areas from returning planes alone; Heckman selection models and censored-data likelihoods correct for missingness without observing what went missing.

True as far as it goes, and it goes quite far — but Wald's correction worked because the filter was known and physically continuous: planes were lost by being hit, and hit locations vary smoothly. Every such correction needs an identifying assumption about the selection mechanism, and that assumption cannot be tested from the surviving sample itself. Heckman-style estimates are famously fragile to their exclusion restrictions. Where the filter is heterogeneous and historically contingent — an archive, a crawl, a curriculum — no such assumption is credible. Observing the filter directly does not replace statistical correction; it supplies the one input the correction cannot infer on its own.

Continuous intake does not escape selection, it relocates it. Sensors are sited by budget and politics, bandwidth forces compression, and retention policy deletes on schedule.

This is correct, and it narrows the claim rather than defeating it. Continuous intake does not yield an unbiased sample of the world; siting, compression and retention are all filters too. But they are filters that occur inside the observed period, which means the siting decision, the compression ratio and the erasure request are themselves events that can be logged, dated and later interrogated. Archival survivorship is unmeasurable because the deletion left nothing behind. Contemporaneous survivorship is measurable because the deleting act is itself on the record. Bias with a known magnitude is not the same epistemic object as bias with none.

Much of what fails to survive is redundant, and aggressive preservation costs more than the error it prevents.

Often true, and worth conceding without qualification for the bulk of ordinary cases: the seventh copy of a shipping manifest adds nothing. But redundancy is not evenly spread, and its scarcity concentrates exactly where the stakes are highest — rare events, marginal populations, systems already failing, the CLIWOC logbooks whose captains steered around the storms that later climatologists most want to model. Turner's 2008 audit of antidepressant trials found 31 per cent of FDA-registered studies never published, and the published subset ran 94 per cent positive against 51 per cent in the full registry — a gap invisible without an independent register to compare against. The missing tail dominates inference even when it is a negligible share of total volume, which is precisely why an outage log that coincides with the event under study is not housekeeping. Correlated missingness is the dangerous kind, and only a record of the outage exposes it.

What this does and does not establish

The concept establishes that any corpus, however large, encodes an unmeasured selection process, and that enlarging the corpus does not touch the process. It establishes that removing the bias requires watching the filter operate in real time, with reasons attached to every gap — which is a structural argument for continuous, provenance-bearing intake, not an incidental feature of it.

It does not establish that such intake yields a neutral or complete picture of anything. Siting, bandwidth and retention remain filters, and they will keep shaping what gets kept. It does not establish that the correction is free: dating gaps and preserving provenance chains costs storage, governance and trust, and that cost is real. And it does not establish that intelligence built this way is thereby wiser, safer or more capable in any general sense. It establishes something narrower and more useful: that this particular ladder, the one measured by how a system takes in the world, has a rung above which the missing categories run out, because there is no observation that lies outside continuous, dated, provenance-bearing intake. That is a claim about a filter. It is not a claim about a mind.

Continue