Home/Concepts/Batch versus stream processing in insurance underwriting
Batch versus stream processing in insurance underwriting
On the intake axis, the terminal position is structural, not aspirational. Batch and stream are not two points on a continuum of latency; stream is the general case and batch the…
The case against, stated at full strength
An underwriter who has priced catastrophe risk for twenty years will hear this argument and be unimpressed. Batch is how the discipline works, and it works because it is bounded. A property book is priced against a hazard curve refreshed once or twice a year, using models from AIR Worldwide or RMS or Verisk, run against the season's exposure snapshot, reconciled with the treaty terms in force at bind date. That snapshot has a name, a version number, and an audit trail. Everyone in the room, from the actuary to the reinsurer to the regulator, agrees on what was known when the price was set. Continuous intake sounds, to this reader, like continuous uncertainty — a book of business that never stops moving under the underwriter's feet, priced against beliefs that might have changed between quote and bind. Insurance is a business of contracts fixed at a point in time. A pricing model that will not commit to a point in time is not an improvement. It is a liability.
This objection deserves to be taken seriously before it is answered, because it is not wrong about the mechanics. It is wrong about what follows from them.
What batch actually buys the underwriter
Batch processing in underwriting is not laziness. It is a bounded set computed over as a whole: the exposure registry as of the valuation date, the catastrophe model version in force, the reinsurance programme terms as filed. This gives three things the industry genuinely needs — reproducibility (a regulator can rerun the same inputs and get the same price), reconciliation (finance can tie premium to a fixed model run), and contractual clarity (the policy wording references a rating basis, not a moving target). Nightly and quarterly batch cycles for exposure aggregation, PML rollups, and treaty renewal pricing are not primitive. They are the correct architecture for a large share of the work, for exactly the reason nightly ETL jobs still dominate data warehouses a decade after the streaming argument was settled in computing.
That is the concession, and it should be made without hedging. Batch will remain the cheap, auditable path for most rating work. The Large Universe Model argument is not that this should stop. It is that batch is a special case of something larger, not a rival to it — and that the something larger is where the failure mode that actually breaks books lives.
Where the snapshot breaks
Here is the characteristic failure, and it is not hypothetical: a book is priced on a hazard curve that the last two seasons have already broken. Florida wind books priced against a curve calibrated on pre-2017 storm behaviour, before Irma and Michael reset loss expectations. Wildfire books in California priced against exposure data that predates the fuel-load and drought conditions that made 2020 and 2021 the worst fire years on record. The catastrophe model vendor issues a new curve version eventually — RMS and AIR both cycle major model updates on multi-year schedules — but the claims flow, the actual loss experience arriving week by week from the field, already told the story before the vendor's release notes did. The underwriter priced correctly against the snapshot. The snapshot was stale before the ink dried.
This is not a modelling error in the actuarial sense. It is an intake error. The book was priced by a system that had, by construction, declared its inputs complete at a point that preceded the evidence.
The two streams worth naming
Two things are actually running continuously beneath every underwriting decision, whether or not the pricing architecture admits it.
The first is the claims flow itself. Every first notice of loss is a data point about the hazard curve, not just about the individual policy. A cluster of unexpected water-damage claims in a region priced as low-risk is a signal about the curve's error, available in real time, long before the next scheduled model refresh would surface it.
The second is the exposure registry and the reinsurance terms that sit against it — which also move continuously. Insureds add locations, change occupancy, shift limits mid-term. Reinsurance treaty terms get amended, reinstated, or exhausted as the season progresses. A batch snapshot freezes all of this at valuation date and then treats the frozen version as truth until the next cycle, even as the underlying registry keeps changing underneath it.
Neither of these streams stops so that underwriting can compute over a complete set. They are unbounded in the same sense a Kafka topic is unbounded: there is no final record, only a position in the log.
| assumption about input | typical cadence | who owns the gap | |
|---|---|---|---|
| Batch rating (current norm) | complete at valuation date | quarterly / annual model refresh | the underwriter, retroactively |
| Streaming intake (the argued alternative) | never complete | continuous, with lateness handled explicitly | the underwriter, at bind time |
The person who carries the gap
The failure does not land on an abstraction. It lands on an underwriter, personally, at the moment a loss develops that the priced curve did not anticipate. The underwriter signed a book against a hazard model that was the best available snapshot at the time, and is then accountable when the claims flow reveals the snapshot was already wrong. This is the underwriting equivalent of a late-arriving event in a stream processor: a fact that pertains to a period already closed, arriving after the window shut. Streaming infrastructure has a name for this and a mechanism for it — watermarks, lateness bounds, retractable results. Underwriting mostly does not. It has an annual model version bump and an underwriter who absorbs the interval.
What a continuous-intake architecture would actually require
None of this argues for real-time repricing of bound contracts, which would be both operationally absurd and legally meaningless — a policy is a fixed instrument once bound, and that should not change. The claim is narrower and structural: the belief that feeds pricing should never be treated as closed, even though the contract that results from it is.
Concretely, this means the claims flow, catastrophe model outputs, exposure registry and reinsurance terms should be held as a continuously updated belief state with provenance — this hazard estimate derives from model version X, adjusted by Y actual losses observed since, with a confidence that decays as the underlying model ages — queried at the moment of quote, rather than baked into an annual snapshot and treated as gospel until the next cycle. The rating engine's inputs become a stream with a position and a lateness policy, even though the contract it produces remains a discrete, dated, auditable artefact. Batch survives exactly where the earlier concession placed it: as the bounded special case used for reproducible year-end reserving and regulatory filing, where a frozen snapshot is the correct answer to a question that genuinely wants one.
Where the objection still wins
There is a version of the counterargument that survives entirely intact, and it should be stated plainly rather than argued around. Beliefs that revise under late-arriving claims data can oscillate, and an underwriter chasing a hazard curve that updates weekly risks pricing volatility that a stable annual cycle was specifically designed to prevent. Reinsurers price treaty terms against the assumption that primary carriers are not repricing their book weekly in response to noisy early loss signals. A continuous-intake system that lacks explicit hysteresis on how much a curve is allowed to move per unit of new evidence, and explicit provenance distinguishing a confirmed catastrophe-model revision from a short run of unlucky claims, is not an improvement on the snapshot. It is a snapshot that shakes.
The honest position is therefore narrower than "stream beats batch in underwriting." It is that the hazard belief underlying a rate should be held as a continuously revisable, provenance-carrying object rather than a frozen file, that batch remains the right mechanism for the contract itself and for regulatory reporting, and that the gap between a broken curve and its next scheduled refresh — the gap an underwriter currently absorbs alone, after the fact — is exactly the space a properly built continuous-intake system exists to close. Nothing beyond that gap needs claiming, and nothing less than that gap should be conceded.