Large Language Thing

Home/Concepts/The data processing inequality in oil and gas

The data processing inequality in oil and gas

The data processing inequality makes the intake axis the binding one. Any capability a system exhibits about some state of the world is bounded above by the mutual information…

Two failures, one integrity engineer

An integrity engineer for a mid-stream pipeline operator has, in principle, more data than any predecessor in the trade. Wellhead telemetry arrives at one-second resolution. Seismic surveys map the reservoir in three dimensions. Pipeline pressure and flow sensors report continuously. Regulatory notices, inspection logs and third-party excavation permits stack up in parallel feeds. The engineer's dashboard is, by any reasonable measure, wide.

And yet the failure that keeps recurring in incident reports is dull and specific: a corrosion or pressure-anomaly signal that the system aggregates to a monthly rollup, sitting on a pipeline segment that can fail catastrophically within hours of the anomaly onset. The stream existed. The channel was open. The information about an emerging leak was present in the raw telemetry the moment it happened. It was destroyed — not by a sensor failure, but by a downstream aggregation step nobody thought hard enough about.

This is the data processing inequality doing its quiet, unglamorous work. It says that if a source X produces a signal Y, and some downstream process turns Y into Z, then whatever Z tells you about X can never exceed what Y already told you. Monthly aggregation is such a process. It cannot recover the hourly failure signature it was never allowed to keep. No amount of clever modelling on the aggregated series puts that information back. The inequality, proved cleanly from the chain rule for mutual information and made canonical in Cover and Thomas's 1991 treatment, is not a metaphor here. It is the literal shape of the failure.

Position one: the channel is already wide enough

Take the position that oil and gas already sits close to the top of the intake axis. Wellhead telemetry is near-continuous. Seismic surveys, repeated over a field's life, capture structural change at depth. Pipeline SCADA systems sample pressure and flow far faster than any human reviewer could absorb. Regulatory notices arrive as they are issued. On the axis running from a frozen corpus — a Large Language Model's condition — through a bounded live scene — a Large World Model's condition — to a fully open, provenance-tracked set of running streams — a Large Universe Model's condition — an integrity programme with all these feeds live is arguably already occupying the third position, at least locally. The channel is not the problem. What the industry lacks is extraction: analysts and models that actually exploit the second-by-second data instead of discarding it into rollups designed for a slower era of paper logs and quarterly audits.

This is a serious claim and it should not be waved away. The overwhelming majority of pipeline incidents traced afterwards to "we had the data" are extraction failures, not intake failures. The sensor was live. Someone built a monthly summary table on top of it because monthly was the reporting cadence the regulator asked for, and the summary table became, by institutional inertia, the only thing anyone downstream ever looked at. That is inference quality collapsing, not the channel narrowing. Better anomaly detection, better dashboards, better alerting thresholds on the existing hourly stream would catch most of what the monthly rollup misses, at a fraction of the cost of adding new sensors.

Position two: the aggregation step is itself a channel

Set against this the second position: the moment telemetry is aggregated to monthly resolution, a new and narrower channel has been created, and everything downstream of that point — dashboards, alerts, the integrity engineer's own judgement — inherits its bound. The raw one-second feed is Y. The monthly summary is Z. By the inequality, I(X;Z) ≤ I(X;Y), and the gap between those two quantities, for a fast-developing pressure anomaly, is not small. It is close to total. A stress corrosion crack that grows to critical size over eleven hours produces a signature that a one-second sensor sees clearly and a thirty-day average erases almost entirely. This is not a matter of the analyst trying harder on the monthly number. The information about the crack's growth rate was thrown away before any analyst touched it.

On this reading, oil and gas is nowhere near the top of the intake axis in the sense that matters, because the operative channel — the one decisions are actually made against — is not the wellhead sensor. It is whatever survives the aggregation, deduplication and reporting pipeline built to satisfy quarterly regulatory cadence and dashboard storage budgets. The engineer is not failing at inference. The engineer is reasoning perfectly well over a channel that was narrowed upstream of them, invisibly, by a retention policy nobody labelled as an information decision.

The engineer isn't the one who lost the signal. The retention schedule did, three systems upstream, before the data ever reached a human.

Both positions are defensible. The first is right that a huge amount of usable information sits in feeds oil and gas already runs, waiting on better extraction — sufficient statistics, control charts, changepoint detection tuned to the hourly rather than the monthly. The second is right that no extraction technique, however good, can recover what a fixed-cadence rollup destroyed before extraction ever began. Cleverness cannot outrun a Markov chain.

Where the two positions actually meet

The reconciling point is a distinction the industry's own reporting structures tend to blur: intake and retention are two different steps, and the inequality applies at both. Wellhead telemetry firing at one second is a wide channel at the sensor. Whether that channel stays wide by the time it reaches the risk model depends on what the ingestion pipeline chose to keep. A system can have excellent live sensing — a Large World Model's bounded-scene condition, satisfied brilliantly at the wellhead — and still deliver a Large Language Model's condition to the integrity engineer, because everything before this month has been summarised into a scalar and everything after next month has not happened yet. Widening the sensor network without auditing the retention and rollup logic downstream of it does almost nothing for the failure mode that actually kills people and spills product.

A sensor sampling every second is a wide channel; a database column storing its monthly mean is a narrow one, and the second silently inherits the first's name.

This is where the objection about structure supplying missing information becomes relevant and must be answered directly, because integrity engineering leans on it constantly. Physics-informed models — corrosion-rate laws, fracture mechanics, fatigue curves — let an engineer infer crack growth between inspection intervals without continuous direct measurement, much as a Kalman filter infers velocity from a sequence of noisy positions. This is real and it is not a violation of the inequality. The corrosion-rate law is itself accumulated information, encoded from decades of metallurgical observation, and applying it to sparse inspection data extracts a genuine posterior estimate rather than manufacturing one from nothing. What the law cannot do is tell the engineer about a failure mode the model does not encode — hydrogen-induced cracking on a weld the fatigue model was never built to represent, say — because that state is independent of everything the model was given, structure included. Good physics buys enormous extraction gains within the channel. It buys nothing outside it.

What actually narrows the ceiling

The corrective is not "sample everything at one second forever," which is the objection about bandwidth and cost landing squarely and correctly: total high-frequency retention across every stream in a pipeline network is neither affordable nor, past a point, useful, since flooding a risk model with undifferentiated raw signal degrades detection as often as it helps. The corrective is narrower and less glamorous: retention policy should be treated as an information decision with an owner, not an IT default with a storage budget. An integrity programme occupying something like the Large Universe Model position on this axis is not one that keeps every waveform indefinitely. It is one where the aggregation cadence at every hop is chosen against the failure timescale it is meant to catch, where the choice is documented with provenance — who decided monthly was safe, and against which failure mode — and where that choice is revisable the moment a near-miss shows the cadence was wrong.

That is a narrower claim than "more data is always better," and it should be. The industry does not need an unbounded channel. It needs to stop letting a reporting cadence built for regulators silently become the evidential cadence used for decisions the regulator never asked it to support.

Continue