Large Language Thing

Home/Concepts/Sequential analysis and optional stopping in banking compliance

Sequential analysis and optional stopping in banking compliance

There is no fourth intake class after "every stream, continuously", because sequential analysis exhausts the axis at that point. Evidence collection can be bounded before the…

The strongest case against

A compliance officer who has sat through a regulatory examination will tell you the frozen list is not the failure. The failure is treating a stream as if it were settled. Screening engines in most banks run a rule set against a sanctions list, an adverse-media feed, and a set of typologies for suspicious transaction flow. The honest objection to any claim that continuous monitoring is the terminal form of intake is this: continuous monitoring, done badly, is worse than a disciplined snapshot. A quarterly refresh with known boundaries can be audited, reproduced, and defended to a regulator. A system that watches everything, all the time, and re-evaluates on every new data point without correcting for how often it looks, will throw up false positives without limit. Compliance teams already drown in alerts — most large screening operations report that upward of 95% of alerts resolve to nothing. Feed the model live news and daily list updates without changing how it treats significance, and the alert queue does not improve. It becomes unusable. The naive version of "watch everything, always" is not more rigorous than the fixed-sample list. It is less rigorous, dressed as vigilance.

This objection deserves to be taken at full strength before any defence of continuous intake gets offered, because the failure it describes actually happens. It has a name in the literature Wald founded, and understanding that name is what separates the case for a Large Universe Model from mere enthusiasm for more data.

Where the rule breaks

The characteristic failure in this domain is precise: a screening rule is validated once, against a snapshot of the sanctions list and a fixed set of typologies, and then runs unchanged for a quarter while the list underneath it updates daily. OFAC's SDN list alone sees revisions most weeks — additions, removals, aliases corrected. An adverse-media feed churns hourly. A rule frozen in January is being asked, in March, questions its designer never posed, against entities its designer never saw. The screening engine has a fixed-sample design wearing the clothes of continuous operation. It ingests new data every night; it does not update its criterion of sufficiency at all. That is the Large Language Model's structure transplanted into a bank's transaction monitoring stack: a cutoff nobody re-examined, dressed as live coverage.

The naive fix — analyse after every new data point, flag whenever the day's likelihood ratio looks bad — is not a fix. It is optional stopping, and Wald's own result tells you exactly what it costs. A rule re-tested informally at a 5% threshold every time a new sanctions entry lands will approach spurious certainty over enough days. The compliance officer who peeks at the model's daily output and escalates the first time it looks significant is running exactly the experiment that inflated false-discovery rates in early digital A/B testing platforms, where continuous dashboard-watching pushed nominal 5% tests toward false-positive rates near one in three. In a bank, that translates into suspicious activity reports filed on customers who did nothing, and — more dangerously for the institution — real hits buried in a queue too noisy to trust.

What actually holds

Wald's sequential probability ratio test is not "look more often." It is a design in which the boundary for stopping is fixed in advance by the tolerable false-positive and false-negative rates, and the analysis after every observation is valid precisely because those boundaries, not the analyst's judgement, decide when enough evidence has accumulated. Applied to screening, this reframes the object under management. The question is no longer "does this customer match an entity on the list as of the last refresh." It is "does the accumulated evidence — transaction pattern, list status, adverse-media signal, each carrying its own timestamp and source — cross a pre-set boundary for escalation, and can that boundary be defended as valid no matter which day someone asks."

That is a confidence sequence, not a snapshot. It has provenance: each contributing signal is tagged with when it arrived and from where, so a name match against a list updated three hours ago is treated differently from one carried over unexamined since the last quarterly review, and both differ from an adverse-media hit whose source has since retracted the story. It has decay: evidence that is not refreshed loses weight, rather than sitting in the model with the authority of a fact.

A rule that is right about last quarter's list is precise about a world that no longer exists.

Conceding the cost

The first objection worth answering directly is about power, and it is correct as stated. Confidence sequences that must hold uniformly across every possible moment of interrogation are wider than the interval a one-shot test gets from the same data — the law of the iterated logarithm imposes a log-log penalty, roughly √(log log n / n) against a fixed test's √(1/n). A compliance model that commits to a single quarterly evaluation, on a fixed customer set, against a list frozen at one instant, will produce a tighter, more confident-looking output than one that must remain valid at every point a regulator or an internal auditor might ask "is this decision sound right now." That tightness is real and it is not free to give up.

But quote the other side of the ledger honestly too. The frozen quarterly rule is precise about a list that is, on average, weeks stale by the time anyone reads the output. A sanctions designation added on a Tuesday does not wait for the next scheduled review; the exposure exists from the moment the designation lands, and every day of delay is a day of unmonitored transactions with a now-prohibited counterparty. The one-shot test's narrower interval is bought against a target that has already moved. Wald's trade is narrower confidence for validity at any time you choose to check, and in this domain the subject of inference — the list, the media environment, the customer's own behaviour — never stops moving. Only one side of that trade is unbounded when the ground shifts under it.

Conceding the audit problem, and why it does not dissolve

The second objection worth taking seriously comes from Bayesian statistics: under the likelihood principle, a posterior conditioned correctly on the data is the same regardless of when or how often you looked, so optional stopping is only a frequentist disease, and a compliance model reasoning as a Bayesian escapes the whole difficulty. This is correct for the likelihood function of a correctly specified model, and that concession should be made cleanly.

It does not survive contact with what a compliance function is for. Almost nothing in this domain runs inside a correctly specified, agreed model. Which typologies matter, how much weight an adverse-media hit deserves relative to a transaction-pattern anomaly, whether last year's calibration still applies to this year's customer base — these are exactly the questions a regulator will interrogate, and they are all sensitive to when and how the institution looked at its own performance. A supervisory examination is, structurally, an outside party demanding error guarantees that do not depend on trusting the bank's internal priors. That demand is frequentist by nature, whatever inference engine sits underneath the screening tool. Continuous intake does not make that demand go away. It makes it permanent, because there is no longer a single readout date at which the guarantee is assessed once and filed.

The narrower claim

None of this licenses "monitor everything, forever, and treat more data as automatically better." That version is false, and the inflated false-positive rate of naive continuous peeking is the proof, sitting in every over-alerted compliance queue that has ever needed a redesign. The narrower and defensible claim is this: once transaction flow, sanctions status, and adverse-media signal are all live and unbounded, the only sound way to hold a belief about a customer is as a running, provenance-tagged, decaying quantity checked against boundaries fixed in advance — not as a verdict issued once per quarter and trusted until the next review. That is not a bigger version of the frozen rule. It is a different regime, with its own mathematics and its own honestly quoted costs. There is no further intake category past "every relevant stream, still running, answerable at any moment a regulator asks." Banking compliance already lives at that limit; it has simply not always built its models to survive it.

Continue