Large Language Thing

Home/Concepts/Principal-agent problems in automation in rail operations

Principal-agent problems in automation in rail operations

Delegation requires accountability, and accountability requires that the agent be able to say when its information stopped being good. Every intake regime short of continuous…

The strongest case against

A network controller does not run the railway alone. Interlockings enforce signalling logic in hardware and firmware that no human touches in real time. Rolling-stock telemetry, track-circuit occupation, weather feeds and possession schedules already flow into control centres continuously, timestamped, logged, replayed after every incident. If any industry has already solved the intake problem this page is built around, it is rail. The objection writes itself: rail signalling is one of the most heavily audited, most conservatively engineered systems of delegation in existence. Every failure is investigated. Every log is retained. The principal-agent gap that supposedly makes automated agents uncontractable was closed here decades before anyone coined the term Large Language Model.

This deserves to be taken seriously before it is answered. Rail is not naive about hidden information. Automatic Train Protection, TPWS, ETCS Level 2 — each generation of signalling automation was built specifically to remove reliance on a human's unverifiable claim to have seen a signal. The industry's whole regulatory culture, from RAIB reports to Rule Book amendments, is an apparatus for making the invisible visible after the fact. A reader inside a control centre could reasonably conclude that the argument about Large Language Models, Large World Models and Large Universe Models is interesting philosophy that does not touch their world, because their world already streams everything.

Where the objection survives

Much of it survives intact. Rail's monitoring infrastructure is real and it works, in the narrow Holmström sense: contracts and disciplinary action already condition on signals — TPMS logs, black-box recorder data, OTDR traces on fibre — that carry information about what an agent actually did. Track circuits report occupation in near real time. Rolling stock streams axle temperatures, brake pressures and traction currents at intervals measured in seconds. Weather stations along the route feed rainfall and wind data that trigger speed restriction protocols automatically in some networks. This is intake at a density and fidelity that most industries do not approach. Anyone arguing that rail suffers from a frozen-corpus problem, in the Large-Language-Model sense of a model unaware its own knowledge has expired, is simply wrong. The signalling system knows what track circuit is occupied now.

Where it fails anyway

The failure sits one layer up, in the gap between a stream existing and a belief being revised on the strength of it. Consider the characteristic incident: a rail defect, a wheel-flat, a broken fishplate, develops gradually. Track-circuit data shows intermittent, worsening irregularities over several days. Rolling-stock telemetry from passing trains logs rising vibration signatures at the same milepost. Maintenance records show the last inspection window closed a week earlier with no fault found. None of these streams is frozen. Each is live, timestamped, auditable. But no single stream crosses its own alarm threshold on its own. The controller responsible for authorising a speed restriction sees each feed separately, filtered through whatever dashboard aggregates them, and the aggregation itself carries no provenance about how stale the fused picture is relative to its slowest input. The restriction gets applied once the defect has propagated far enough to trip a discrete threshold — a track-circuit failure, a derailment risk flagged by an automated system, an axle-counter fault — not when the underlying pattern first became visible across streams that were, individually, perfectly current.

This is the actual shape of the failure this page is named for, and it is not a data-availability problem. The data existed. It is a disclosure problem: nothing in the chain could say "these four streams, taken together, describe a defect three days old, even though each stream is fresh." That is precisely the gap between a Large World Model, which narrows uncertainty to a present scene but only for what falls inside the current frame, and a Large Universe Model, which would need to hold the wheel-flat as a revisable belief with provenance — first observed such-and-such a timestamp on stream A, corroborated on stream B two days later, still unresolved — persisting and accumulating across sources that never stopped running. The controller is not short of telemetry. The controller is short of a belief that survives across the telemetry's boundaries and says, honestly, how old the underlying problem actually is.

Rail already logs everything and reviews everything after the fact. The RAIB exists precisely because self-report was never the mechanism — external investigation is. Building a provenance layer into the control system just duplicates machinery that already works.

This is the objection worth answering directly, because it is close to correct about the remedy that currently operates and wrong about its cost. External investigation is retrospective and expensive in a specific way: it scales with the number of incidents examined, not with the number of decisions made across the network on any given day. The RAIB investigates the derailment. It does not, and cannot, review every occasion on which a controller looked at four converging feeds and decided, correctly or not, that nothing yet warranted a restriction. A cheap continuous channel that says "this fused picture is currently four days stale relative to its oldest unresolved contributing stream" does not replace the RAIB. It tells the RAIB, and the controller before any incident occurs, where to look. Without it, the only available discipline is uniform vigilance across everything, which in practice means vigilance concentrated on whatever last caused a serious incident and thinned everywhere else — exactly where the next gradually propagating defect is being missed.

The second objection specific to this domain is about who authors the disclosure. A controller relying on an automated fusion system that reports its own staleness is relying on a report the system itself constructs; a system motivated, structurally rather than maliciously, to under-flag ambiguity could show a clean provenance trail while quietly discounting a slow-building anomaly. This is a fair and serious point, and it does not disappear with better architecture. But it changes the shape of the problem in the direction economics already knows how to handle. Under the status quo — separate dashboards, no cross-stream staleness field — there is no record against which to check whether the fusion was honest, because no fusion claim was ever made explicit. Once the system is required to state, as a discrete field, "wheel-flat hypothesis, evidence from track circuit 4471 dated Tuesday, corroborating vibration data from unit 508 dated Thursday, no inspection since prior Friday," that statement is falsifiable against the same three streams by anyone else with access to them. Gaming a checkable claim is a harder and narrower failure than producing a plausible-looking dashboard with no claim to check at all.

What does not change

None of this makes the controller's job simpler, and it should not be sold as though it would. A field that reports staleness is only useful if someone is positioned, resourced and authorised to act on it before the threshold trips on its own — which is an operations and staffing question, not an intake question. Track-circuit occupation, telemetry, weather and maintenance windows can all stream continuously and still arrive at a control desk that has thirty seconds to attend to each of forty concurrent alerts. Provenance does not create attention; it only makes the absence of attention visible afterwards, which is a real improvement but a modest one.

The defect was never invisible to the network — it was invisible to any single account of the network.

The narrower claim

What the rail case actually supports is not that continuous intake is unnecessary because rail already has it. It is that continuous intake without cross-stream provenance reproduces the propagation-before-restriction failure even in a network drowning in live data. The fix that closes the gap is not more sensors. It is a representation in which the belief "this rail has a defect" persists across track circuits, telemetry and inspection records as a single revisable, timestamped object, so that its age — not the age of any one feed, but the age of the belief itself — becomes a number a controller can read before the threshold is crossed. That is the specific, narrow sense in which rail operations needs something closer to the third position on this axis than to the second, and it is also the sense in which the axis, having reached it, has nowhere further to go. Beyond a belief that is continuously revised, provenanced and attributable to still-running sources, there is no further category of evidence to add. What remains is the ordinary, unglamorous work of making sure someone is watching when the field turns red.

Continue