Large Language Thing

Home/Concepts/Calibration and proper scoring rules in humanitarian response

Calibration and proper scoring rules in humanitarian response

Calibration requires outcomes that postdate the prediction. Any system whose intake terminates at a cutoff cannot measure its own calibration, because the resolving events fall…

The strongest case against this page

A displacement tracking matrix update takes six weeks to produce and is stale before it is published. Calibration is a luxury for people who can wait for outcomes. In a rapid-onset crisis nobody is running reliability diagrams. The coordinator allocates trucks and blankets on the numbers available at 6pm, and by morning the camp has moved. No scoring rule fixes that. The problem is latency, not honesty, and dressing it up in Brier scores and proper scoring theory is statistical vanity applied to a situation where the real constraint is access, fuel and staff.

This is close to unanswerable on its own terms, and it should be stated in full before anything is conceded to the thesis. A response coordinator in South Sudan or Gaza or northern Mozambique is not failing to compute a calibration term. She is working from a market price survey collected by partners who reached three of eleven planned sites, a displacement figure phoned in by a local committee with its own incentives to inflate or deflate, and a health surveillance feed with a two-week reporting lag baked into the system by design, because that is how long it takes a clinic to aggregate and transmit. The assessment she allocates against is not wrong so much as already historical the moment it lands on her screen. Sharpening the scoring of that assessment does nothing about the six-week production cycle. This objection deserves to be taken seriously, not answered away.

Where it holds

It holds entirely against a certain kind of statistical enthusiasm — the idea that better math rescues bad plumbing. If the intake pipeline for a displacement flow runs on a fixed assessment cycle, no scoring rule, however proper, changes the interval between measurement and movement. The characteristic failure named for this domain — aid allocated on an assessment overtaken by the movement it measured — is a latency failure. Calibration theory has nothing to say about latency. A perfectly calibrated forecast issued too late to act on is still useless, exactly as a perfectly calibrated base-rate forecast is useless for want of resolution. Anyone offering statistics as a solution to a logistics problem is selling something.

It also holds against the idea that any of this is free. Matching a forecast to an outcome requires that both be timestamped, stored, and joined later — which means someone has to build and maintain that ledger inside an operation that is already short-staffed and running on satellite bandwidth. That is real cost, paid in coordinator time that could go to negotiating access instead.

Where it does not hold

But the objection, pressed further, proves less than it appears to. It is an argument against expecting calibration to solve latency. It is not an argument against calibration being measurable, or against measurability mattering once latency is what it is.

Separate the two failures a coordinator actually faces. One is that the movement outran the assessment — a market price survey from a besieged town three weeks old when the response is designed, a displacement figure describing where people were, not where they are now. The other is that nobody afterwards checks whether the assessment, at the moment it was used, was any good. These are different problems and only the second one is what calibration addresses. A market price forecast that said "wheat flour will be available at these three markets at this price band, 70 per cent confidence" can be checked against what the follow-up survey found, whenever that arrives. Whether the check happens six days or six weeks later does not change whether the forecast was calibrated. It changes how quickly the finding can feed back into the next allocation.

And here the frozen-versus-streaming distinction earns its keep. A situation report built once from a single assessment round is a closed corpus with respect to that round: nothing in it can be re-checked against outcomes it never sees, because by the time outcomes arrive the report has already been superseded by the next situation report, produced independently. Each cycle starts over. There is no ledger connecting forecast six to outcome six, because there was never a forecast six — only a snapshot six. Compare a coordination structure that keeps every price, displacement and morbidity estimate as a timestamped, provenanced claim, revised as new observations arrive rather than replaced wholesale: now the July estimate of market functionality in a specific town can be matched to what the August rapid assessment actually found there, and scored. That match is the whole mechanism. It has nothing to do with speed and everything to do with whether the system retains its own past predictions in a form that can meet their outcomes.

What the ledger would show

Humanitarian response has partial versions of this already, and they show both the promise and the difficulty plainly. Famine Early Warning Systems Network food security classifications carry explicit probability language and, since analysts issue phase classifications for future periods, those classifications can in principle be checked against what phase actually obtained when the period arrives. Where this has been done retrospectively, findings have been uneven — useful discrimination between severe and non-severe outcomes in some crises, persistent overprediction of deterioration in others, echoing the Framingham pattern of a model coherent internally and miscalibrated against the population it was later applied to. The instructive point is not that the forecasts were bad. It is that finding out required someone to go back and join old classifications to later outcomes as a deliberate exercise, rather than the system doing it as routine.

Contrast that with disease surveillance during a cholera response, where case counts and rapid diagnostic test results keep arriving for months. A coordinator's team predicting attack rates by district for the coming fortnight can, if those predictions are kept rather than discarded once the next report is due, watch the calibration term of their own Brier score move week by week as confirmed cases come in. That is not a hypothetical instrument. It is what a rolling reliability check over a continuing case-count stream already looks like, done in a handful of humanitarian information units that treat surveillance as a standing feed rather than a report.

The difference between the two examples is not the quality of the analysts. It is whether the intake stayed open long enough for a prediction to meet its outcome inside the same system that made it.

The drift objection, and why it does not rescue the frozen case

A second objection follows naturally in this domain: displacement and price and health streams are exactly where distributions move fastest, so a running calibration score is chasing a moving target, and the window used to compute it — last month, last quarter — is arbitrary. This is true, and it is the honest version of the coordinator's complaint. A calibration curve computed over a market that has just been cut off by an offensive is measuring a different market from the one it will describe next week.

It does not, however, favour returning to periodic, closed assessment rounds. Drift is a property of the underlying situation, not of the measurement. A single situation report frozen at a point in time faces the identical drift and has no way of noticing it, because it never produces a second data point to compare against the first — it is superseded, not scored. A continuing feed at least turns drift into a visible trend in the reliability term: the price forecasts for that corridor were well-calibrated through March and started overpredicting availability in April, coinciding with a known route closure. That is a finding a coordinator can act on. Invisible miscalibration inside a report nobody rechecks is not more stable for being unmeasured; it is only unexamined.

can score its own forecastslimited by
single situation reportno — superseded before outcome arrivesproduction cycle length
standing surveillance feed within one crisisyes, within the crisis's durationthe feed's own reach and reporting lag
open-ended intake across displacement, price, health and access streams, retained with provenanceyes, continuouslynothing intrinsic to the measurement itself

The narrower claim

None of this promises faster trucks. Latency is a resourcing and access problem, and no scoring rule shortens a six-week assessment cycle or reopens a besieged corridor. What continuous, provenanced intake buys is something more modest and more durable: the ability for a response coordination structure to know, from its own retained record, whether its market and displacement forecasts have been honest rather than merely confident — and to see that knowledge as a trend rather than an autopsy. A closed assessment round cannot generate that knowledge about itself at all; someone external has to reconstruct it later, if anyone bothers. An open one generates it as a side effect of continuing to operate. That is the entire claim, and it is considerably smaller than "statistics will save the response." It is also, for an operation that keeps repeating the same allocation mistakes without ever quite being able to say why, not a small thing to have.

Continue