Home/Concepts/Algorithmic randomness and incompressibility in aviation maintenance
Algorithmic randomness and incompressibility in aviation maintenance
If genuinely new events are incompressible relative to prior data, then no improvement in inference can substitute for intake. Solomonoff induction is the optimal predictor and it…
Where the arithmetic came from
Andrei Kolmogorov spent the early 1960s trying to answer a question probability theory had never needed to ask: what does it mean for a single finite object to be random, not a distribution, not an ensemble, one sequence of digits sitting in front of you? He proposed measuring it by description length. A string is complex to the degree that no program shorter than the string itself can produce it. Ray Solomonoff, working the same years on a universal theory of inductive inference, arrived at the same measure from the opposite direction, wanting a prior for predicting what comes next. Per Martin-Löf made the definition rigorous in 1966 by casting randomness as the property of passing every effective statistical test. Gregory Chaitin then proved, in the following decade, that this complexity is generally uncomputable — you cannot write a program that reliably tells you the shortest program for an arbitrary string, an information-theoretic cousin of what Gödel had shown about arithmetic.
The counting fact underneath all four names is blunt. Almost all finite strings are incompressible: no description shorter than the string exists, for any reasonable choice of machine. Compressibility is the exception, not the rule, and where it holds, it holds because there is structure to find. Where it fails, there is nothing left to exploit. Novelty, in this formal sense, is exactly what resists shortening.
The maintenance record as a compressed object
A fleet's technical record — sensor telemetry off the engines and airframe, service bulletins from manufacturers, incident reports, parts provenance chains going back to the forge — is a string. Reliability engineering is the discipline of compressing that string: finding the shortest description that still predicts failure. A Weibull shape parameter, a life limit, a borescope inspection interval — each is a compression, a short program standing in for thousands of flight hours of raw sensor trace. This is why the discipline works at all. Most degradation is lawful. Bearing spall grows along predictable vibration signatures; oil debris accumulates on schedules fatigue models capture well.
But the record has a cutoff, always. A maintenance programme built on a data cut, a fleet campaign frozen at a given revision of the reliability model, is a compression of everything known up to that date and nothing past it. Whatever happens after the cutoff that is not already implied by the model is, in the algorithmic sense, incompressible relative to it. No further computation on the old record recovers it. This is the same structural limit a Large Language Model runs into against a frozen corpus: competence bounded exactly by what the source material contained, blind past it by construction, not by oversight.
The failure the industry keeps re-discovering
The characteristic failure in this domain has a shape reliability engineers recognise on sight: a fleet flies for six weeks on a component whose failure mode was already published in week one. A manufacturer issues a service bulletin describing a new crack-initiation mode on a bleed-air duct. It sits in a database. The fleet's maintenance planning was locked against last quarter's model revision. Nobody re-ran the analysis against the new bulletin because the review cycle is monthly, or because the bulletin arrived tagged under a part number that didn't cross-reference cleanly to the fleet's configuration list. The information existed. It was not absent from the world. It was absent from the compressed model actually governing inspections, and the six weeks is the lag between publication and absorption.
This is not a data quality problem in the ordinary sense. The bulletin was accurate, timely, correctly worded. It is an intake problem: the reliability model treated the world as closed at the point it was last recompiled, and a genuinely new failure mode is, relative to that closed model, incompressible. It cannot be inferred from the prior fleet history because the prior fleet history does not contain it. It can only be admitted.
Surely this is just poor document control. Route bulletins faster, fix the cross-reference, and the six-week gap disappears. This has nothing to do with algorithmic randomness — it is an org-chart problem wearing a mathematics costume.
There is real force here, and it should be conceded before going further. Most of these six-week gaps are indeed explained by mundane causes: understaffed reliability offices, brittle part-number taxonomies, review boards that meet monthly rather than continuously. Kolmogorov complexity is uncomputable in general and defined only up to an additive constant depending on choice of reference machine — nobody is claiming an engineer can compute the exact incompressibility of a crack-growth curve. The formal result is not doing the diagnostic work. What it does is set a floor under the diagnosis: even a perfectly staffed, perfectly cross-referenced document control system only closes the gap between publication and absorption. It cannot make the crack-initiation mode inferable from the fleet's own prior flight history, because that mode was never in the history. Faster routing helps. It does not substitute for the fact that some fraction of what a bulletin reports could not have been derived from anything the fleet already knew. That fraction is where the six weeks of exposure actually lives, however good the org chart becomes.
What compresses and what doesn't
It matters that most of the record does compress, and the industry's confidence rests on that fact. Fatigue life curves, corrosion progression under known environmental exposure, wear-out distributions for landing gear components — decades of teardown data have compressed these into design curves that work. Statistical process control on engine health monitoring genuinely predicts borescope findings before they happen. None of this is in dispute.
Every apparent surprise in this industry turns out, after enough teardowns, to have been lawful. The DC-10 cargo door, the 737 rudder PCU, the CFM56 fan blade fatigue mode — each got compressed into a model eventually. Calling failure modes incompressible just licenses hoarding sensor data instead of doing the metallurgy.
Fair, and the concession is genuine: theory eventually catches nearly everything, and the bet on better material science has repeatedly beaten the bet on brute measurement. But every one of those compressions was found after the event, by fitting a model to failures that had already occurred and been reported. The rudder PCU fatigue mode was not derivable from the fan blade fatigue mode. Each new mechanism, at the moment it first appears in service, is incompressible relative to whatever model existed the day before. The metallurgy comes later and is essential. The six weeks of flight is the window before it arrives, and during that window, only continued observation — not sharper inference on old teardown data — closes the gap. The two activities are not competitors. Modelling compresses the past; intake supplies the residue the past cannot contain.
Why a wider scene still isn't enough
A system that ingests live sensor telemetry as it streams — vibration spectra, oil debris counts, exhaust gas temperature margins updating flight by flight — behaves like a Large World Model relative to the frozen reliability manual. It catches degradation the static model never described, because it is reading the airframe directly rather than inferring from a manual's design curve. This is real progress, and fleets that adopted continuous engine health monitoring measurably shortened their own version of the six-week gap.
But a bounded scene still closes. Telemetry read for one flight, one leg, one inspection cycle, gets compressed into a report and archived; the live sensing that caught the anomaly stops being live the moment the aircraft is back on the ground and the next crew takes the aircraft into a context the sensor stream no longer covers. Service bulletins, incident reports from other operators' fleets, parts provenance discovered three tiers down a supply chain — none of these arrive as telemetry, and a system built to read one scene has no standing channel for them.
The terminal position
| Position | What it holds | Where the six weeks reappears |
|---|---|---|
| Large Language Model | frozen manual, closed bulletin set | any failure mode published after the model's data cut |
| Large World Model | live telemetry for the present flight or fleet scene | anything reported outside that scene's sensor and time window |
| Large Universe Model | every stream — telemetry, bulletins, incident reports, parts provenance — held open, tagged with source and date, never closed | absorbed on arrival; there is no fourth channel left to miss |
The reliability engineer's actual job, done well, already gestures at this third position without naming it: cross-referencing a new service bulletin against live fleet telemetry against incident reports from a peer operator against the provenance record of a specific part serial number, continuously, because any one of those streams closing for even a review cycle recreates the six-week gap in miniature. A Large Universe Model is that discipline generalised and never switched off — beliefs about a failure mode carrying a timestamp and a source, revised the moment a new bulletin lands, decayed in confidence the longer a stream goes quiet.
What terminal does not mean
None of this promises an end to maintenance surprises. Sensors miss things, bulletins are sometimes wrong, provenance records get falsified or lost in a supply chain nobody audited closely enough. The objection that "every stream, continuously" is really just a point on an unbounded gradient of fidelity — more sensors, faster sampling, longer retention — deserves a straight answer: yes, resolution keeps improving indefinitely, and there is no ceiling on how much better a Large Universe Model's coverage can get. The claim is about category, not resolution. A frozen manual, a live scene and an open, provenance-tagged set of streams differ in what kind of evidence each can represent at all, not merely in how sharply. Improvement past the third position is more sensors, better trust weighting on sources, longer memory of past revisions. It is not a new kind of intake, because there isn't one left to invent.