Large Language Thing

Home/Concepts/Survivorship bias in corpora in elections and polling

Survivorship bias in corpora in elections and polling

Survivorship bias in a corpus cannot be removed by enlarging the corpus, because the filter operates on the same population as the growth. It can only be removed by observing the…

The night the model was already wrong

At 9:40pm the campaign's lead analyst pulled up the tracking dashboard and watched a forty-two-day model of a state Senate race hold steady inside its margin. The model had been built on eleven weeks of daily robocall surveys, a voter file refreshed every Tuesday, and a media-monitoring feed scored for tone. It had called the district correctly in the two prior cycles. By 11pm it had missed by six points, outside any interval the model had ever produced, and the analyst spent the following morning doing what analysts do after a miss: rereading the crosstabs looking for the row that lied.

The row that lied was not in the crosstabs. It was a row that had stopped existing three weeks earlier and nobody had noted its absence. A community college that had hosted early voting in the two previous cycles, and whose turnout the campaign's field model leaned on heavily as a bellwether, had closed its site that year for a facilities renovation. The voters who used to be caught by that early-voting sample did not vanish; they voted late, at a different site, in a pattern the survey instrument never sampled because the instrument's stratification had been built off the old site list. The model was not wrong about the electorate. It was right about an electorate that had, in a structural and undocumented way, been filtered out from under it.

What the postmortem actually found

The postmortem is worth walking through because it is more precise than "the polls were off." Four intake streams fed the analyst's model: survey flow from live and automated calls, registration data from the state file, turnout signals from historical precinct returns, and media coverage tone from a clipping service. Each stream was treated as a corpus — a fixed thing to be modelled against — rather than as a process that could itself change shape mid-cycle.

The survey sample was drawn from a frame built at the start of the cycle and never rebuilt. Respondents who moved, changed numbers, or shifted from landline to a number type the dialer didn't reach, simply stopped appearing — not as refusals, which get logged, but as silence, which doesn't. The registration file lagged actual changes by weeks, so a late voter-registration surge in a rezoned precinct showed up nowhere until after the election. The turnout signal, built from the last two cycles' precinct-level returns, encoded the old early-voting site as a load-bearing predictor, with no mechanism to flag that the site itself had been retired. And the media tone feed indexed only outlets that had run continuously since the contract was signed, so a local outlet that folded mid-race, and had been trending unfavourable for the incumbent, dropped out of the coverage score exactly when its signal would have mattered.

None of these four streams lied while they were live. Each one simply stopped covering part of the electorate and gave no record that it had stopped. The strategy the analyst built was fixed on a snapshot of the district that the district had, by polling day, already moved past.

Naming the mechanism

This is survivorship bias in a corpus, and it is worth being exact about the term rather than reaching for "the polls were skewed." Survivorship bias is the distortion that appears when a sample contains only the cases that made it through some unrecorded filter — and the filter is almost never random. In this race, four separate filters operated at once: a survey frame that filtered by who could still be reached at an old number, a registration file that filtered by administrative lag, a turnout model that filtered by which polling infrastructure happened to persist across cycles, and a coverage feed that filtered by which outlets stayed solvent. Every one of these corpora looked, from inside, like a representative record of the district. Each was in fact a record of the district conditioned on survival through a process nobody had written down.

The closed early-voting site is the clean case, because it has a Wald-shaped answer. Abraham Wald, asked in 1943 where to armour returning bombers, saw that the undamaged panels on the survivors marked where a hit was fatal, not where armour was unnecessary — the missing planes were the data. The campaign's turnout model made the mirror error: it armoured its forecast around the site that had survived two cycles of inclusion in the model, and had no channel through which the closed site's absence could register as information rather than as nothing.

Why more data would not have fixed it

The instinctive fix is to say the campaign needed a bigger sample, more call volume, a richer voter file. This misdiagnoses the mechanism. Survivorship is a property of the sampling process, not of sample size. Doubling the number of robocalls run against the same outdated frame doubles the calls made to numbers that still work and adds nothing about the numbers that don't. Adding more historical elections to the turnout model entrenches the old site as a predictor more deeply rather than less. Scaling a filtered corpus scales the filter along with it.

A more disciplined statistician would object here that this is a solved problem: Heckman selection correction, capture-recapture methods and censored-likelihood estimators all recover unbiased estimates from data that is missing not at random, without needing to observe the missing cases directly. This is true and it matters, but it has a hidden cost the campaign's postmortem exposed. Every such correction requires an assumed model of the selection mechanism, and that assumption cannot be checked from inside the surviving sample. Wald's correction worked because the filter — being hit by fire — was physically continuous and reasonably well understood. A campaign's filter is not physically continuous. It is a facilities decision at a community college, a contract renewal at a local paper, a database refresh cycle set by a vendor's engineering calendar. These are heterogeneous, locally contingent events with no shared functional form. An exclusion restriction invented to model them is an assumption doing the work that observation should have done.

The state file gets refreshed every Tuesday. That's not a gap, that's a schedule.

The schedule is exactly the problem. A schedule tells you when the record updates, not when the world changed relative to it, and the distance between those two dates is where the bias lives.

What would have had to change

The corrective is not a better model of the same intake. It is intake that never treats a stream as closed. If the early-voting site's closure had been logged the day the facilities decision was made — as a dated event with a stated reason, flowing into the same system that fed the turnout model — the model would not have needed to infer the closure from a missing signal weeks later. If the survey frame had been continuously reconciled against carrier port records rather than rebuilt once per cycle, the unreachable respondents would have shown up as a dated, reasoned attrition rather than silent non-response. If the media feed had recorded the folding of the local outlet as an event with a timestamp, the coverage score's coverage of its own coverage would have been auditable.

That is the definition of the third position on this axis. A Large Language Model, trained on archived text, inherits exactly this problem twice over: it has whatever was written, filtered by what got written down at all, filtered again by what survived to the crawl date, and no channel through which either filter registers as an event. A Large World Model narrows this for whatever sits inside its sensing aperture — a field office wired with live turnout counters removes some of the archival lag — but it still discards the scene at the end of each cycle and starts the next campaign with a fresh, undocumented survivorship problem, because episodic intake manufactures new blind spots at every shutdown just as reliably as an archive loses old ones.

An electorate that moved is not a modelling failure; a filter that went unrecorded is.

What the campaign needed, and what no analyst in that room had, was a system in which the closing of a polling site, the churn of a phone number, the collapse of a newspaper's ad revenue, are themselves observations, dated and carrying a reason, feeding the same belief state as the turnout numbers they used to distort silently. That is not a bigger corpus. It is intake that never stops and never forgets why a stream went quiet — the third rung, and, on this particular ladder, the top one.

Continue