Large Language Thing

Home/Concepts/Signal detection theory in humanitarian response

Signal detection theory in humanitarian response

The optimal cut point is a function of prior odds. Prior odds move. Therefore any system whose intake stopped cannot hold an optimal cut point except by luck, and the luck decays…

What arrives

A response coordinator running an operation in, say, a flood-affected region of the Sahel does not receive one stream. She receives four, on different clocks. Displacement tracking arrives from field monitors and satellite imagery, updated every few days, with a lag between a movement happening and someone logging it. Market price data comes weekly from vendor surveys in functioning markets, and stops arriving altogether from markets that close. Health surveillance arrives through facility reporting, which is only as good as the facilities still standing and staffed. Access constraints — which roads are mined, which checkpoints are extorting, which areas are contested this week — arrive as unstructured field reports, sometimes hours old, sometimes weeks stale, with no consistent format.

None of these streams is a signal by itself. Each is evidence bearing on a yes/no judgement the coordinator must keep making: is the caseload in this locality rising past the threshold that justifies diverting a convoy here rather than there. That judgement is a detection problem in the exact technical sense. There is a distribution of readings you'd see if nothing unusual were happening, and a distribution you'd see if a genuine surge were underway, and they overlap. The coordinator's job is to say yes or no under that overlap, repeatedly, at a pace the crisis sets.

What is held

The naive version of the job treats each new report as a fresh judgement. The working version holds a belief: a running estimate of how many people are actually moving, how tight food access actually is, how strained the health system actually is, updated as reports arrive and decayed as they age. This is the part easy to skip under pressure and expensive to skip in practice. A report from a market survey three weeks ago is not false, but it is not current either, and treating it as current is exactly the failure mode that recurs across humanitarian operations: aid gets allocated on an assessment that the population itself has already outrun.

Holding the belief properly means attaching provenance to every figure that feeds it — which source, what method, what date, what confidence — because the figures disagree with each other constantly, and resolving the disagreement requires knowing which report to trust more, not just averaging them. A displacement estimate built from satellite imagery and one built from local informant counts will diverge by a factor of two in active phases of a crisis. Discarding either loses information; blending them without provenance loses the ability to tell later which one was right, which is the only way the blending rule improves.

What triggers revision

Revision is triggered by movement in the underlying rate, not by the arrival of a single alarming report. A single report of a spike in acute malnutrition presentations at one clinic might be noise — a stockout elsewhere pushing caseload sideways, a new clinic opening and capturing cases that were always there. What should move the belief is convergence: the same direction showing up across displacement figures, market prices and health data inside a compressed window. That convergence is the closest operational proxy for a shift in the true prior odds of a surge, and prior odds are the quantity the whole judgement pivots on.

This is where the criterion — the threshold at which "elevated risk" becomes "divert the convoy" — has to move too, and move for a defensible reason. If the base rate of need in a locality has genuinely doubled, holding the old threshold means missing cases that a lower bar would have caught; the old threshold was calibrated to a population that has since walked away. If the coordinator drops the threshold reflexively every time one noisy report comes in, she starts allocating scarce transport capacity to noise, which starves the localities where the surge is real. Getting this right requires exactly what the streams are meant to supply: a continuously updated read on where the true rate actually sits, distinct from where any one report claims it sits.

What the operator sees

In an operations room this shows up as a dashboard that is deliberately unglamorous: a map with confidence bands rather than point estimates, arrows for trend rather than snapshots, and a provenance tag on every number that lets the coordinator ask "how old is this and from whom" before acting on it. What she is watching for is divergence between the assessment and the movement — the gap between the caseload the last full survey recorded and the caseload the trend lines imply exists now. That gap is the operational meaning of criterion drift. It is measured in days, and in active displacement crises it accumulates fast: a locality assessed as stable eight days ago can be hosting three times its recorded population by the time a convoy scheduled against the old figure arrives.

The characteristic failure of the sector is precisely this: aid allocated on an assessment overtaken by the movement it measured. It is not usually a data quality failure. The individual survey was probably accurate the day it was taken. The failure is temporal — treating a snapshot as a standing estimate, in a domain where the underlying population does not hold still for the survey to remain true.

The report is not wrong; it is simply no longer describing the place it describes.

What it costs

Running the loop this way is expensive in a specific sense worth being honest about. It requires standing capacity to keep every stream open simultaneously, staff whose job is reconciliation rather than fresh collection, and a tolerance for dashboards that show uncertainty rather than clean numbers, which is a harder sell to donors than a single confident figure. Coordinators are frequently pressured to report point estimates because point estimates are fundable and confidence bands are not. That pressure pushes toward exactly the failure mode the loop is built to prevent.

Two objections deserve a straight answer here, because both are live in this domain and neither is a strawman.

The first: distribution shift in a crisis is not always a matter of the base rate moving while the detection machinery stays sound. Sometimes the machinery itself degrades — informant networks collapse when an area becomes too dangerous to access, and the reports that do arrive are no longer representative of the population, they are representative of whoever could still reach a phone. No amount of re-estimating prior odds fixes a survey instrument that has stopped sampling the right people. This is correct, and it is a real limit. But the two failures are distinguishable by timescale and by remedy. A base-rate shift is repaired by updating a single estimate once convergent evidence arrives, often within a reporting cycle. A collapsed sampling frame requires rebuilding data collection — new informant networks, new access arrangements — which takes months. Continuous intake does not fix the second problem, but it is what reveals that the second problem exists at all; a coordinator working from a frozen baseline assessment has no way to notice that her sampling frame has quietly stopped representing the population it once did.

The second: keeping every stream open, continuously, for an entire operation is a disproportionate answer to what is, mathematically, a one-number problem — the current prevalence of need. A leaner alternative exists in principle: periodic, well-designed rapid assessments, drawn at intervals, feeding a threshold that gets recalculated on a fixed schedule rather than watched continuously.

Why maintain four live streams when a monthly rapid assessment gives you the same number for a fraction of the cost?

The answer turns on who gets assessed. In access-constrained crises, the localities that get rapid assessments are disproportionately the ones still reachable — which correlates strongly with the ones not experiencing the worst access constraints. A monthly sample built this way is not a random draw from the affected population; it is a draw from the subset the current access situation permits, and that subset is exactly what shifts fastest during a surge. Breaking that bias requires knowing, with provenance, which localities were assessed, which were skipped and why, and when access conditions changed — which is a record kept over a continuing stream, not a periodic snapshot. The rapid assessment is cheap. Knowing whether it can be trusted this month is not, and that knowledge only exists if something was watching in between.

Continue