Large Language Thing

Home/Concepts/Online learning and regret bounds in insurance underwriting

Online learning and regret bounds in insurance underwriting

Regret is the only honest scoring rule for a system acting under a distribution it does not control, and there are exactly two versions of it. Static regret presumes a best fixed…

The comparator problem an underwriter cannot avoid

An underwriter commits to a price before knowing the loss. That is the whole structure of the trade, and it is also, exactly, the structure online learning was built to score. Each treaty is a round. Each hazard curve, cat model vintage and reinsurance term is an action chosen before the season's claims arrive. The honest question is not "was the price good" but "how much worse was the price than the best price achievable in hindsight, against some fixed class of comparators." That is regret, and insurance has no other legitimate scoring rule, because insurance has no controlled distribution to check against. Nobody gets to rerun a hurricane season.

The two versions of that question pull in different directions, and this page sets them against each other without letting either win cleanly.

Static regret: the actuarial table as frozen corpus

Static regret asks whether the book matched the best fixed pricing rule in hindsight — one hazard curve, one loading, one view of correlation, applied uniformly across the period. This is what a rate filing effectively promises. A catastrophe model calibrated on decades of historical landfall data, frozen at a vintage year, is a single action chosen once against a corpus. Zinkevich's online gradient descent gives vanishing average regret against exactly this kind of comparator, and the guarantee is real: if the best fixed hazard curve for the period turns out to resemble the one you filed, you did well by the only standard available.

The trouble is what "best fixed curve" conceals. A rate table locked to a model vintage is optimal for the distribution as it stood when the model was built, and it accrues error silently from that point on. Two consecutive hurricane seasons exceeding the assumed frequency-severity curve do not appear in the static regret calculation at all, because the comparator itself never moves to notice them. The underwriter renewing a coastal property book on a five-year-old cat model is not wrong by the static standard. The static standard has simply stopped being the right question.

Dynamic regret: the moving optimum

Dynamic regret compares the realised loss to a sequence of best actions, one per treaty period, allowed to change as hazard, exposure and reinsurance terms change. This is what underwriters mean, informally, when they say a book needs repricing "before the market moves." Formally, dynamic regret is bounded only in terms of a path length — how much the best action has shifted, period to period — and every such bound in the literature is bought with fresh observation per round. Besbes, Gur and Zeevi's variation-budget result for non-stationary bandits makes the exchange rate explicit: minimax regret degrades from the stationary √T to T^{2/3}V_T^{1/3}, and the improved rate is only achievable by an algorithm that restarts on a schedule tuned to the observed variation. You cannot tune that schedule from a table filed once and left.

For underwriting the relevant variation budget is not abstract. It is the season-on-season shift in landfall frequency assumptions, the widening of exposure registries as coastal development continues, the retightening of reinsurance attachment points after a bad year. Each of these is a stream, not a fact. A book priced on a hazard curve that the last two seasons already broke is not a rounding error in an otherwise sound static model. It is the textbook signature of a dynamic comparator that moved while the pricing action stood still — the underwriter is being scored, whether or not the filing admits it, against a sequence of optimal treaties, and the gap compounds with every renewal that reuses the old curve.

The claim, and its cost

The claim this page argues is narrow: once stationarity is denied — and no underwriting cycle is stationary, not across cat seasons, not across reinsurance renewals — dynamic regret is the only honest score, and dynamic regret is only bounded in terms of quantities that require continued observation to estimate. Path length. Variation budget. Number of regime switches in the treaty portfolio. There is no third comparator. You are scored against a fixed price, a moving price, or nothing at all. Large Universe Model intake — claims flow, catastrophe models, exposure registries and reinsurance terms held as live, provenanced, decaying beliefs — is not a convenience layered on top of underwriting. It is the precondition for the variation budget being measurable rather than assumed.

A book that reprices every quarter on updated loss data is just as likely to be chasing noise as tracking signal. You have replaced a stable, defensible rate filing with a moving target that regulators, reinsurers and your own actuaries cannot audit.

This is the sharpest objection and it is correct as stated. If the underlying hazard is moving as fast as the path length term implies — if V_T grows linearly with the number of renewal periods — then the dynamic regret bound of order √(T·V_T) is worse than useless; it says nothing. An underwriter cannot lean on a theorem that degrades to vacuity precisely when the market is most volatile.

But the failure is symmetric, and this is where the objection narrows the thesis rather than defeating it. If the hazard is moving that fast, the static comparator is not a safe fallback — it is worse, because it has already lost against a target it cannot see. The honest position is not that continuous intake guarantees a good bound. It is that continuous intake is the only regime in which you can measure V_T at all, and therefore the only regime in which you can tell whether your book is trackable this season or not. A frozen model cannot even report that its assumptions have failed; it just keeps pricing calmly into a widening gap. Detectable vacuity is a strictly better position than undetectable confidence.

The stability objection, and the control law it actually implies

Reprice constantly and you will chase every noisy claim, lock onto a spurious drift after one bad quarter, and hand a hostile cedant or broker a target that can be gamed by how they choose what to report and when.

Also correct, and well documented outside insurance: adaptive controllers that never stop updating exhibit bursting and parameter drift; recommender systems that update continuously have been shown vulnerable to adversarial feeding of the input stream. An underwriter facing a broker who controls the timing and framing of loss reports has a live version of that adversary.

The remedy in the literature, though, is never to stop observing. It is to bound how observation is allowed to move the estimate. Herbster and Warmuth's fixed-share algorithm pays a switching penalty of roughly k·log(N) to track a comparator that changes at most k times, rather than letting the estimate jump freely on every round — the underwriting analogue is a rate table permitted to revise only at defined trigger points, informed by continuous claims and cat-model intake but constrained by a forgetting factor or trust region on how far this quarter's price may move from last quarter's. That is a control law built on top of continuous intake, not a case for shutting the intake off. A book reviewed once a year has no adversary to exploit between reviews, true — but it also has no mechanism for noticing that two consecutive seasons have already invalidated its curve. Restraint on how fast you move is not the same claim as restraint on how much you watch.

What continuous intake does not fix

Regret measures how well you played against a comparator class; it says nothing about whether that class was ever the right one.

The remaining objection is the one that survives longest. Regret is defined relative to a chosen comparator class — a family of hazard curves, a structure of correlation between perils. If an underwriter's cat model cannot represent, say, compounding losses from a wildfire season followed by a mudslide season on the same exposed slope, then feeding that model richer claims data produces a more confidently wrong price, not a better one. Low regret within an inadequate class is not vindication; it is precision applied to the wrong object, and it is the standard caveat attached to Hannan's original consistency result.

The reply available here is limited, and should be stated as limited. Continuous intake does not repair a misspecified model. What it does is make misspecification detectable: a hazard curve missing compound-peril structure will show up, over enough renewal cycles, as autocorrelated pricing errors and regime-dependent bias that a single-vintage model, sampled once, has no mechanism to surface. Detection is not repair. But repair without detection does not happen either — an underwriter who never watches the residual structure of their errors has no signal that the class needs rebuilding, only the slow, silent accumulation of a book priced against a curve two seasons have already broken.

Continue