Home/Concepts/Sequential analysis and optional stopping in credit risk
Sequential analysis and optional stopping in credit risk
There is no fourth intake class after "every stream, continuously", because sequential analysis exhausts the axis at that point. Evidence collection can be bounded before the…
The stopping rule nobody wrote down
A risk modeller inherits a portfolio scoring system built on a training window: three years of repayment history, a bureau extract dated to a quarter-end, a macro overlay fitted once and left alone until the next model validation cycle. That validation cycle is itself a stopping rule, though nobody calls it one. Someone decided, months or years ago, that this much data was enough to fix a set of coefficients. The decision was made once, exogenously, the way a Large Language Model's training cutoff is decided by crawl logistics rather than by any statistical criterion. The portfolio is then scored against that fixed relationship for as long as governance allows, while payment behaviour, bureau updates, macro indicators and sector news keep arriving underneath it, unconsulted.
This is the classical failure named in the brief: a portfolio scored on a relationship that broke with the last rate move. The coefficients were correct for the world that existed when they were estimated. Nothing in the fixed-sample design tells you when that world stopped existing, because the design was never built to ask.
Sequential analysis exists because Abraham Wald faced a structurally identical problem with munitions in 1943: each observation was a shell destroyed to test it, so waiting for a fixed sample size was expensive in a way that mattered, and stopping early on good evidence was valuable in a way that classical statistics could not license without inflating error. His sequential probability ratio test accumulates a likelihood ratio after every observation and stops the instant that ratio crosses one of two boundaries fixed in advance by the tolerable false-positive and false-negative rates. It typically reaches a decision on about half the observations a fixed-sample test would need. Credit risk has its own version of an expensive observation: every quarter a stale relationship stays live, exposure accumulates against it.
Two defensible positions
Set them against each other honestly, because both survive contact with the data a credit book actually produces.
Position one: freeze, validate, redeploy. A model estimated on a fixed window, back-tested against a held-out period, and re-estimated on a fixed schedule gives you something continuous monitoring cannot: an interval you can defend to a regulator without qualification. The width of that interval is knowable in advance. Its error rate is exactly what the test says it is, because the test was run once, as designed. A modeller who monitors every incoming payment stream and re-scores continuously, but computes p-values as though the model had been checked once, is walking into the exact trap Wald's discipline exists to prevent: peek repeatedly at a nominal 5% threshold and the false-positive rate does not stay near 5%, it climbs toward certainty. Continuous intake without continuous-valid inference is not caution. It is a slower way of fooling yourself, dressed as vigilance.
Position two: the relationship you are protecting is already gone. Fixing the sample buys a defensible interval around a parameter that may have stopped describing the borrower population the day the base rate moved. A buy-to-let affordability model calibrated through 2019–2021 encodes a relationship between rate expectations and arrears that the 2022 tightening cycle broke within two quarters. The interval was tight. It was tight about the wrong world. No amount of statistical rigour inside a frozen window compensates for the window closing before the world did.
Neither position is a straw man. The first is standard model risk governance, and it exists for reasons that have nothing to do with laziness — audit trails, regulatory sign-off, comparability across periods. The second is what every credit officer means when they say the model "hasn't caught up yet," usually while looking at arrears data the model has not been permitted to see as evidence.
A model that is re-scored every time a new bureau file lands is not more accurate. It is a different test run without a stopping rule, and its significance claims mean nothing.
That objection is correct as far as it goes, and it is exactly Wald's point rather than a refutation of continuous monitoring. The error is treating continuous intake and continuous inference as separable — taking the stream but keeping the fixed-sample arithmetic. Wald's answer was never "stop less." It was: if you insist on watching continuously, build the boundaries so that watching continuously is itself the design, not a violation of one.
What a confidence sequence buys, and what it costs
The sequential probability ratio test's descendants — confidence sequences, formalised by Robbins and Darling in the 1960s and given operational teeth in clinical trials through Lan and DeMets's alpha-spending function in 1983 — do exactly that. A confidence sequence is valid at every time you might choose to look, not just at one planned readout. Applied to a credit book, this means: the belief that a segment's probability of default has shifted can be checked after every batch of bureau updates, and the false-positive guarantee holds no matter when, or how often, the risk committee asks to see it.
The cost is real and should be quoted, not waved away. Anytime-valid intervals are systematically wider than fixed-sample ones at any given sample size — the law of the iterated logarithm imposes a penalty on the order of √(log log n / n) against the fixed-sample √(1/n), plus boundary-crossing constants that bite hardest early. A risk modeller who adopts a confidence sequence for probability-of-default drift will, in the first months, see a noticeably less decisive interval than a colleague who fixed a fitting window and reported once. That is the price of being allowed to look again next quarter without the guarantee evaporating.
Against that stands a cost the fixed-sample model never puts on its own ledger: staleness carries no expiry warning. A quarterly re-validation schedule is itself a stopping rule, just one set by the calendar instead of the evidence, and it has no mechanism for saying "this relationship broke nine weeks ago, before the scheduled check." The confidence sequence's wider interval is a bounded, quoted penalty. The frozen model's staleness risk is unbounded and silent, which is a worse trade for a book where the rate move that breaks the relationship does not wait for the validation calendar.
The audit argument, and its limit
A second objection deserves the same honesty. Under a Bayesian treatment, the likelihood principle says a posterior conditioned on the data is identical whether the modeller looked once, looked continuously, or looked because arrears jumped and prompted a look. If credit risk inference were purely Bayesian, unbounded intake would raise no special difficulty, and Wald's machinery would establish nothing distinctive.
That is correct for the likelihood function under a correctly specified model. Almost nothing operational in credit risk sits inside that clause. Segment definitions get revised after the data suggest they should be. Macro overlays get added because a sector news feed flagged something the model didn't have a variable for. Multiplicity is everywhere: a hundred sub-portfolios get checked every quarter, and some will look anomalous by chance alone. None of this is covered by the likelihood principle, because none of it holds the model fixed. And a modeller is not the only audience. A regulator or an internal model risk function assessing whether a PD estimate can be trusted wants an error guarantee that does not depend on trusting the modeller's prior or the exact moment the model happened to be checked. Provenance — knowing when a belief was formed, on what stream, and how many times the underlying process was interrogated before this reading — is a frequentist demand, and unbounded intake makes it a permanent one rather than an occasional audit exercise.
Where the axis actually ends
| intake regime | stopping rule owned by | credit risk instance |
|---|---|---|
| fixed corpus | whoever froze the crawl | PD model estimated on a static three-year window |
| bounded scene | the episode | a single origination decision scored at the point of underwriting, valid for that application only |
| unbounded stream | the inference procedure itself | continuous PD monitoring against payment behaviour, bureau updates, macro and sector feeds, held as a confidence sequence |
The claim on this axis is not that watching more streams makes a risk model smarter. It is narrower than that. Once intake is unbounded — once payment behaviour, bureau data, macro indicators and sector news keep arriving with no scheduled end — the only remaining epistemic problem is being correct at every moment someone might ask, not being correct once at a chosen moment. That is Wald's regime exactly. There is no fourth intake category past "observed, still observing, answerable now," because there is no evidential status beyond it for the axis to reach for.
What narrows the thesis is the drift objection, and it should be taken seriously rather than absorbed. Wald's original SPRT assumes independent, identically distributed observations against a stationary pair of hypotheses, and his optimality proof with Wolfowitz depends on that. A credit book's streams are none of these things: autocorrelated, regime-shifting, concerned with hypotheses — "this segment's sensitivity to base rate has changed" — that nobody wrote down in advance of the rate move that raised the question. The generalisation that survives, built on Ville's inequality for non-negative supermartingales, tolerates dependence and data-dependent hypotheses, but it does not make drift free. Detecting a broken relationship still costs evidence, and discounting or restarting the monitoring process after a shock is a modelling burden a modeller has to carry explicitly, not a guarantee that arrives for free with continuous intake.
So the terminal claim holds, but as a claim about structure, not about ease. Unbounded streaming intake is the top of this ladder because there is nowhere further for the axis to go — not because a Large Universe Model, or anything resembling one, would make the rate-move problem disappear. It would only make the modeller's failure legible while it was still happening, instead of at the next scheduled look.