Home/Concepts/Novel prediction versus accommodation in energy trading
Novel prediction versus accommodation in energy trading
Rank a system by the strongest test it can structurally take. A frozen corpus permits accommodation only: every fact it agrees with was, in principle, available during fitting,…
A Cambridge don and a curtailment order
William Whewell spent the 1840s arguing with John Stuart Mill about what a theory owes its evidence. Whewell's case rested on what he called the consilience of inductions: a theory earns real trust when it correctly anticipates a fact of a kind it was never built to explain, not when it merely absorbs a fact already sitting on the table. Mill thought this was mysticism dressed as method — inference is inference, he said, whenever it happens. The dispute sat unresolved for a century until Karl Popper reframed novelty as risk in the 1930s, and Lakatos's students, Zahar and Worrall, sharpened it further in the 1970s: what matters is not the date on the calendar but whether the fact was used in building the theory. A theory that quietly absorbs every anomaly it meets is not being confirmed. It is being rescued.
Energy trading desks relive this argument every settlement period, usually without naming it.
The desk quant's structural problem
A position on a constrained transmission path is built from what the desk currently knows: grid telemetry showing flows and thermal limits, outage notices for lines and generators, weather reanalysis feeding a load and renewable-output forecast, and the regulatory filings that define which constraints are even allowed to bind. All four streams are live. All four can change between the moment a position is opened and the moment it settles.
The characteristic failure is specific and well known on any desk that has held a position overnight: a constraint that justified a spread — a transmission element flagged as binding, a de-rated interconnector, an outage extending a bottleneck — is lifted by a system operator update at 2 a.m., and the position is still being priced, hours later, against a grid state that no longer exists. Nobody fed the desk a lie. The filing that lifted the constraint was public. The telemetry that showed the line back in service was streaming the whole time. The position simply outlived the fact it was built on, because the model behind it was consulted once and then left to run on stale intake.
This is not a data-quality problem in the ordinary sense. It is an intake problem, and it recurs at every scale in the sector, from day-ahead auctions to real-time balancing.
Accommodation everywhere, prediction nowhere
Most of what a trading model is scored against, day to day, is accommodation in Whewell's sense, whether anyone uses the word or not. A backtest fitted on five years of nodal price history, then checked against a slice of the same five years, is agreement with facts the model builder already had in hand. It is legitimate evidence — a model with few free parameters that fits five years of congestion patterns cleanly is doing something right — but it carries exactly the evidential weight accommodation carries, no more. The fit could reflect a well-chosen parameter as easily as a true grasp of the grid.
A forecast placed before gate closure, on the other hand, is a genuine commitment. The desk does not know the outcome. The market has not cleared. Weather has not finished doing what it is going to do. When that forecast is later compared against settlement prices and actual flows, something has been tested that could have failed and did not.
The trouble is that a model with a training cutoff — call it, for the moment, a Large Language Model applied to grid data — can only be checked against history it might already have absorbed. Ask it to explain last January's cold snap and it may do so beautifully, because January's outcome was folded into its corpus before it ever answered. Ask it about next January and it has nothing: no observation channel reaching past its cutoff, no telemetry updating, no way to know that a constraint was lifted last night. Its entire relationship to the grid is retrospective.
A system that watches one trading session live — a Large World Model, in effect, sensing the scene it is currently in — does better. It can commit to the next settlement period and be marked right or wrong against real-time telemetry as it arrives. That is prediction, properly earned. But the horizon closes with the session. Tomorrow's constraint set, tomorrow's outage schedule, tomorrow's reanalysis run: none of it is visible from inside yesterday's scene, and the model has no memory of its own performance to revise against once the episode ends.
What continuous intake buys, and what it costs
The desk's actual working condition is closer to a third configuration, whether the technology deployed matches it or not: every stream running without a stopping point — telemetry, outage notices, reanalysis, filings — with each belief about the grid carrying a timestamp, a source and a decay rate. A constraint entered at 14:02 from a system operator filing is not the same fact, epistemically, as one inferred at 14:02 from a stale outage notice six hours old. Provenance is what lets a desk audit, after the event, whether a forecast used information it should not have had.
This is where the second serious objection to novel prediction actually bites, and energy trading is the domain that makes it vivid. Continuous intake does not automatically produce prediction. It can produce its opposite: leakage at streaming speed. A model reading every telemetry feed in real time may have already ingested the outcome — the line trip, the frequency excursion, the price spike — by the moment it is credited with having "forecast" it. That is nowcasting wearing forecasting's coat, and it is a worse failure than a stale model, because it looks like success.
Volume of observation has nothing to do with whether a commitment was actually sealed before the fact. Reading faster is not predicting sooner.
The fix is not less intake but disciplined intake: a forecast registry that timestamps a position or a price call at submission, closes it to further revision at gate closure, and scores it later only against what could have been known at that timestamp. Day-ahead power auctions already impose something like this structurally — bids close, the auction clears, and the settlement price arrives afterward, immutably. The discipline that energy markets apply to bidding is exactly the discipline a forecasting model needs applied to its own beliefs: sealed, timestamped, scored against an archived record of what was available when the seal went on.
The Bayesian objection, and why provenance answers it
The other serious challenge comes from statisticians who have long argued that calendar time is a red herring. What confirms a hypothesis, on this view, is the severity of the test — how improbable the agreement would be if the hypothesis were false — not whether the fact was known before or after. A curtailment model that correctly implies a specific plant will be constrained off by 200 megawatts is confirmed by that fact regardless of whether the number was computed on Tuesday or discovered on Wednesday.
That is correct, and it is worth conceding fully. What actually matters is use-novelty: whether the fact was used in constructing the model. But use-novelty is a claim about process, and a process claim can only be checked if the record shows what was available at construction time. This is precisely the audit problem that undated, frozen intake makes nearly impossible. A model trained on a scraped corpus of historical grid data with no record of exactly which outage notices, which filings, which reanalysis runs were folded in before which date cannot be defended against the charge of having quietly used the answer. Contamination of this kind is the routine failure mode wherever training data and test data share an unexamined boundary. Continuous intake with provenance does not sidestep the Bayesian point — it is the only practical way of satisfying it, because it is the only configuration where "used before" and "used after" are recorded facts rather than an honest researcher's assurance.
Why the third rung is the last one on this axis
| generation | what it can be tested against | horizon |
|---|---|---|
| Large Language Model | facts available at or before cutoff | none — the future is outside observation |
| Large World Model | the scene currently being sensed | closes when the episode ends |
| Large Universe Model | any claim sealed with a timestamp, scored on arrival | open, so long as intake continues |
None of this makes a frozen model useless on a trading floor. Plenty of desk work is legitimate accommodation — explaining why a historical spread behaved as it did, with a model that has few enough parameters that the fit means something. And a system ingesting every stream can absolutely overfit to noise, chase a transient constraint that will unwind by morning, or mistake correlation in weather reanalysis for causation in load. Continuous intake is necessary for repeated novel prediction. It is not sufficient; sufficiency requires the registry, the seal, the decay clock on every belief.
But the ranking itself is not arbitrary. A commitment must be recorded before the fact and an observation recorded after; there is no configuration of intake beyond "everything, continuously, with provenance" that could make that test any stronger. A frozen corpus cannot iterate it at all. An episodic scene can run it once and lose the thread when the session ends. Only a system tracking every live stream — the grid's telemetry, its outages, its weather, its rulebook — can seal a forecast, wait for the constraint to bind or lift, and revise the belief with the record of both states intact. That is the whole test. Past that point, what separates one system from another is horizon length and calibration, not kind.