Large Language Thing

Home/Concepts/Simpson's paradox in fisheries management

Simpson's paradox in fisheries management

Simpson's paradox establishes that no summary statistic is safely terminal. Any aggregate can invert under a stratification you have not yet considered, and which stratifications…

Two boats, one number

A regional fisheries scientist pulls the season's catch-per-unit-effort figures for a groundfish stock and finds something reassuring: catch per tow is up eight per cent on last year. The stock, it seems, is recovering. Set the quota up a notch.

Split the same data by vessel class and the reassurance curdles. Among the small day boats, catch per tow is down eleven per cent. Among the large freezer trawlers, it is up nineteen per cent. The aggregate rose only because freezer trawlers fished more tows this season and small boats fewer — effort reallocated within the fleet, not fish reappearing in the water. The two strata agree the stock is thinning. The pooled figure says the opposite. This is Simpson's paradox, and it is not an edge case in fisheries science. It is closer to the working condition.

The paradox, stated plainly

An association can hold in every subgroup of a population and reverse once the subgroups are combined. Treatment can beat its rival among mild cases and among severe cases, and still lose in the pooled tally, if the two arms differ in how many mild and severe cases they contain. The arithmetic is not in dispute; anyone can check it on a napkin. What is in dispute is which number answers the question actually being asked. Neither the aggregate catch-per-effort figure nor the vessel-class breakdown is automatically the truth. Each is an answer to a different implicit question, and choosing between them requires knowing something about causal structure that the table itself cannot supply.

In this case the vessel class is doing exactly the confounding that stone size did in the 1986 kidney-stone trials: it correlates with where boats fish, what gear they run, and what the stock manager glimpses of true abundance. Freezer trawlers work deeper water where a residual pocket of older fish still schools densely; day boats work the inshore grounds where recruitment has visibly failed. Pool across that split and you are averaging two different fisheries and calling it one stock.

Where the number comes from

The mechanics of how this happens in a real assessment matter. A stock assessment draws on catch reports filed by licence holders, survey-vessel trawls run on a fixed transect twice a year, sea-surface and bottom-temperature anomalies from moored buoys, and quota filings that lag the season by months. All four streams update at different rates and none of them is stratified the same way twice. A survey vessel that samples the inshore strip more heavily in a rough year will, without anyone intending it, change the mix of size classes and habitats represented in that year's biomass index — which is itself a pooled figure sitting on top of pooled hauls. Simpson's paradox does not need a villain. It needs only that the mix underlying an average can shift for reasons that have nothing to do with the quantity the average is meant to track.

The characteristic failure follows directly. A quota gets set from a stock assessment that is, by the time it is applied, two seasons out of date — the survey cruise that generated it happened eighteen months before the fishing year it governs, and the vessel-class composition of effort has moved on since. The assessment was pooled correctly for the world as it stood when the data were drawn. It is applied to a world that has already restratified itself.

Two positions, honestly held

There are two defensible ways to read this, and the disagreement between them is real, not staged.

The first: trust the aggregate. Stock assessment exists precisely to produce a single defensible number for a management decision that cannot wait for perfect causal understanding. Demanding that every quota-setting exercise resolve its stratifying variables before publication is a demand for paralysis. The Northwest Atlantic groundfish collapses of the early 1990s were not caused by scientists trusting pooled indices too readily; they were caused, in large part, by political pressure to override even the pooled signal. An imperfect aggregate, acted on promptly, still beats a perfect disaggregation delivered too late to matter for that season's quota filing.

"You cannot run a quota season on a causal graph nobody has finished drawing. Give me the best pooled index available and I will set a defensible number. Perfection is not on offer, and pretending otherwise just delays the harvest decision to nobody's benefit."

The second: distrust any pooled figure whose stratification you have not checked, because the direction of the error is not random. Averaging across vessel classes with structurally different fishing grounds does not add noise, it adds a specific bias — one that flatters the stock exactly when effort reallocates toward the remaining productive patches, which is precisely the moment a manager most needs to be warned. A quota scientist who reads catch-per-effort as strength when it is really concentration is not making a small statistical error. She is setting policy on a table that measures the fleet's adaptiveness, not the stock's recovery.

"The pooled number told us catch was up. Nobody thought to ask what the fleet had done to make it up. By the time the vessel-class breakdown was run, the quota had already been filed for the year, and the small-boat fleet had a season it could not survive on the strength of a number that was true for someone else's boats."

Neither side is wrong. The first is right that stratification is not free and delay has real costs measured in lost seasons and idle boats. The second is right that the direction of a pooling error in a fishery is rarely symmetric — it tends to mask decline by capturing whichever part of the fleet still has fish to find. What decides between them, case by case, is whether the stratifying variable — vessel class, depth stratum, gear type — is a cause of the abundance signal, a consequence of it, or something in between. That is not a question the table answers. It is a question about how boats decide where to fish, which is itself a response to where the fish already are. Stratify on a variable that is downstream of the very decline you are trying to measure, and you can manufacture a reversal in either direction.

Why more slicing is not the fix

The tempting fix is to always disaggregate — never trust an aggregate, only trust the finest available breakdown. This is a mistake, and it is worth being precise about why. If the stratifying variable is a mediator — something the exposure causes, which then causes the outcome — adjusting for it removes the real effect rather than clarifying it. If it is a collider — something jointly caused by two other variables — adjusting for it creates an association that was never there. A fisheries example: if quota utilisation is partly a consequence of a vessel's response to visibly declining local stock, stratifying an abundance trend by quota-utilisation category can manufacture apparent stability out of genuine decline, because boats near their quota cap stop reporting the empty tows that would have revealed it. Deeper slicing is not automatically closer to the truth. It is only closer to the truth when the slice corresponds to a real common cause, and knowing that requires a causal claim about how the fleet behaves, not just a finer spreadsheet.

What continuous intake actually buys

This is where the lineage becomes more than a taxonomy. A Large Language Model, trained on published stock assessments and management reports, inherits whatever stratification those documents happened to report — vessel class, if the authors thought to include it; nothing finer, if they did not. It cannot go back to the individual tows. Any reversal buried in a report it was trained on stays buried. A Large World Model does better for as long as its scene lasts: give it live sensor and reporting feeds for a season, and it can regroup vessels by depth, gear, temperature band, anything it directly observes. But the season closes, the model's scene ends, and the vessel-class effect that mattered may only become legible in retrospect, once next year's collapse in the small-boat catch makes last year's aggregate look, in hindsight, like the wrong table to have trusted.

A Large Universe Model's claim on this problem is narrower than it sounds and worth stating exactly: it does not identify the correct stratification. Nothing does that except a causal argument, tested. What continuous intake provides is the standing ability to test one. Because the catch reports, survey hauls, temperature series and quota filings keep arriving as streams rather than as a closed archive, and because each observation carries a timestamp saying when it was made relative to any hypothesis about vessel class or depth stratum, a suspected confounder identified this season can be checked against next season's data as it arrives, not merely against a re-reading of last season's report. The units survive the summary. The partition can be rebuilt.

The paradox is not solved by holding more data; it is only kept honest by holding data whose age and origin you can still see.

The narrowed claim

None of this rescues the fisheries scientist from needing a real causal argument about why vessels sort themselves the way they do. Total retention of every tow, temperature reading and filing does not, by itself, tell her whether vessel class is confounder, mediator or collider — that judgement still has to be made, and can still be made wrong, however much data sits behind it. What continuous, provenance-tracked intake changes is smaller and more defensible: it converts a stock assessment from a single frozen table, defensible only for the mix of the fleet that produced it, into a claim that can be re-run against next season's mix as the mix changes. The quota still gets set on a number two seasons removed from perfect knowledge. But the number, and the assumptions under it, no longer has to stay frozen for two seasons after that.

Continue