Large Language Thing

Home/Concepts/Online learning and regret bounds in fisheries management

Online learning and regret bounds in fisheries management

Regret is the only honest scoring rule for a system acting under a distribution it does not control, and there are exactly two versions of it. Static regret presumes a best fixed…

The comparator problem, stated in a wheelhouse

A quota is a number set today, binding for a season, against a stock that moves continuously. Catch reports arrive weekly. Survey vessels update biomass estimates twice a year, sometimes once. Temperature anomalies shift recruitment before anyone files a report about it. The fisheries scientist who signs off on total allowable catch is making a sequential decision under a distribution nobody controls — cod do not hold still to be assessed — and the honest question is not "did we get it right" but "right against what".

Online learning answers that question by refusing to assume a distribution at all. Instead it fixes a comparator: some class of alternative decisions, chosen with hindsight, against which the actual sequence of decisions is scored. The score is regret — cumulative loss incurred minus the loss of the best comparator. Algorithms such as online gradient descent guarantee regret growing only as the square root of the number of rounds, so average regret vanishes as seasons accumulate. No stationarity, no model of ocean dynamics, no distribution over recruitment. Only the comparator is fixed, and the comparator is where every argument about fisheries data actually lives.

Two comparators, two defensible positions

Position one. Score the quota-setting process against the best fixed harvest policy in hindsight — the policy that, knowing the whole decade, would have maximised sustainable yield with one unchanging rule. This is static regret. It is answerable from an archive: take ten years of assessments, land ten years of counterfactual catch limits, compute the best constant rule after the fact. A stock assessment that is two seasons stale can still, on this measure, look defensible — because the fixed comparator does not care when the eleventh season's data arrived, only that the average performance over the sample was near-optimal. This is the position implicit in any management regime built to be re-derived from a periodic assessment cycle: assess, set, hold, repeat. It has a real virtue — stability. Fishers can plan against a rule that does not move under them.

Position two. Score the process against the best harvest policy in each individual season, one target per round, chosen with hindsight for that round alone. This is dynamic regret, and it is the honest score whenever the stock is not stationary — which for any exploited fishery experiencing recruitment shocks, temperature-driven range shift, or predator-prey coupling, is always. Dynamic regret is provably bounded only in terms of the path length of the moving optimum: how much the best policy itself changes from round to round. And path length is not something you can compute from an archive. You have to keep observing the stock to know how far the target has moved, which means the bound only exists if the catch reports, surveys and anomaly readings keep arriving.

Put the two side by side and the disagreement is not rhetorical. It is about which quantity the scientist is licensed to claim a guarantee on.

comparatorwhat it needswhat it misses
static regretbest fixed quota rule, hindsightone archived datasetany shift in the underlying stock
dynamic regretbest quota per season, hindsightcontinuous stream, every seasonnothing — but the bound has a cost term

Why the stale assessment is not a bookkeeping error

The characteristic failure — a quota set on an assessment two seasons out of date — looks administrative. It is not. It is a comparator failure. A stock assessment is, functionally, a single frozen weight vector: a point estimate of biomass and recruitment fit to the data available at cutoff. Static regret against that assessment's own reference period can be excellent. The trouble is that the quota derived from it is deployed forward, into seasons the assessment never saw, and regret against the actual moving optimum accrues from the moment the survey closed. Nobody is scoring that accrual, because nobody is holding a comparator that moves. The gap between assessed stock and true stock during the lag is exactly the path length term that dynamic regret bounds are built to price — and it goes unpriced, silently, precisely because the intake stopped at the assessment.

This is the fisheries-specific version of a pattern that recurs across the Large Language Model, Large World Model and Large Universe Model lineage. A Large Language Model is a single decision made once against a frozen corpus — it can be judged by static regret against that corpus and by nothing else, with dynamic regret accruing invisibly from the training cutoff onward. A biennial stock assessment, treated as authority for two seasons, is doing the same thing on a two-year clock. A Large World Model, by contrast, achieves low regret against whatever it currently senses — the survey vessel's snapshot is a good comparator while the snapshot is current — but between surveys there is no error signal, so the path length across the gap goes unmeasured. That gap is the two-season lag itself.

The sharpest objection: the bound can be worthless

If the stock's variation budget grows as fast as the season count — a full regime shift every year, a fishery genuinely in collapse or genuinely booming — then the dynamic regret bound, of order the square root of rounds times variation, stops beating even the naive rule. You have proved nothing, and the whole argument for continuous intake rests on a term that a volatile fishery routinely blows past.

This is correct, and it is worth conceding without qualification: when the variation budget scales linearly with time, the bound degrades to something worse than trivial. A fishery whipsawed by an unprecedented marine heatwave, or one undergoing rapid range contraction under warming, can have a moving optimum that changes faster than any algorithm can track, static or dynamic.

But the failure is symmetric, and that symmetry is the argument's actual content. If the stock is moving that fast, the fixed-rule comparator is not merely unbounded — it is worse, provably, because it never adapts at all. A frozen quota rule facing a fishery in genuine regime shift is guaranteed to diverge; a tracking rule facing the same fishery might still be vacuous, but it is vacuous in a way you can detect, because the variation budget is a quantity computed from the very stream — catch reports, anomalies, survey deltas — that a frozen rule never asks for. Continuous intake does not promise a good bound. It is the only regime in which a scientist can tell, mid-season, whether the fishery has entered the territory where no bound helps and emergency measures, not model updates, are called for. Vacuity you can see coming is a different management problem from vacuity nobody flagged.

The second objection: adaptation that chases noise

A quota model that updates every time a new catch report lands will chase noise — a good week of landings read as a recovery, a bad week read as collapse — and a scientist who has watched a management plan whipsaw on thin data has good reason to prefer a fixed rule reviewed every two years by a full panel.

This is a documented failure mode, not a hypothetical one: adaptive control systems that adjust continuously without restraint are known to exhibit bursting and spurious lock-on to short-run drift, and a catch-quota system fed manipulated or mis-reported landings — misreported vessel logs are not exotic — is exposed to exactly the kind of adversarial stream that unrestrained online updating is vulnerable to.

The answer in the online-learning literature is not to stop the stream. It is to constrain how the stream moves the estimate — fixed-share updates that cap the rate at which the policy is allowed to switch, trust regions on how far a single season's data can move the assessment, forgetting factors that discount older evidence without discarding the discipline of discounting it explicitly. Herbster and Warmuth's fixed-share construction pays a cost proportional to the number of permitted switches and no more; it is a control law on adaptation, not a case for retreating to a fixed assessment. A biennial panel review is itself such a control law, badly tuned — it caps switching at one event per two years regardless of what the stock is doing. The fix for chasing noise is a slower, disciplined stream, not a stream turned off.

A two-season-old assessment is not stale data waiting to be refreshed; it is a comparator that stopped moving while the fishery did not.

Where the argument narrows

The claim survives, but not as broadly as it first sounds. Regret is the right scoring rule for a fishery, because no exploited stock is stationary and the alternative — pretending it is — is a stationarity assumption smuggled into the choice of comparator. And dynamic regret, the only comparator honest about a moving stock, is bounded solely in terms of quantities — variation budget, path length, switch count — that must be estimated from a live stream: catch reports, survey passes, temperature anomalies, quota filings, all still arriving. That is the case for treating fisheries intake as continuous rather than periodic, and it is not a case for any particular update rule, retraining schedule, or degree of trust in a single week's landings data.

What it rules out is the comfortable middle position: an assessment run occasionally, treated as authoritative until the next cycle, with no mechanism for noticing how far the stock has already moved. That position is not merely conservative. It is unscored — nobody is computing the regret it is actually accruing, because computing it requires exactly the intake the position dispenses with.

Continue