A curriculum certified for a market that no longer exists
The programme director signs off the annual curriculum review in March. The core modules — the ones carrying accreditation weight — have not changed in substance for three cohorts. Employability metrics from the last graduating class look fine: 78% in graduate-level roles within six months, comfortably above the sector benchmark the accreditor uses. On paper, the pipeline is healthy.
What the review does not show is that the roles those graduates entered no longer resemble the roles the curriculum was designed to feed. The data-analytics module still teaches a workflow built around manual ETL and dashboarding tools that hiring managers stopped listing in job specs two recruitment cycles ago. The employability figure is real but stale: it measures whether graduates got hired, not whether what they were taught is what got them hired, or whether it is what keeps them employed at year three. Nobody lied. Nobody was negligent, in the ordinary sense. The mechanism that produced the gap is structural, and it sits in what the review process was ever built to look at.
What the review actually measured
Trace the intake. The annual review consumes assessment results, module evaluation scores, and completion rates. It does not routinely ingest live labour-market signals — job postings, skills taxonomies from recruiters, employer exit interviews — except in a triennial external audit that lags the market by design. Engagement telemetry from the learning platform, which would show students disengaging from a module the moment its content stops mattering to them, exists but sits in a separate system the review board has never queried. Curriculum change proposals from individual lecturers, who often notice market drift first because their own former students tell them, go into a committee queue with an eighteen-month cycle time.
Each of these streams was collected. None of them was combined at the point where the decision got made. The review board compressed its intake down to the handful of numbers that its own reporting template was built to receive — pass rates, satisfaction scores, headline employment rate — and everything else was, for the purposes of that decision, absent. Not wrong. Absent.
This is not a story about bad metrics. Pass rates and satisfaction scores are the right numbers to check whether teaching is working as designed. They are the wrong numbers to check whether the design still matches the market. Different question, different summary. The review board asked the first question with data built to answer it, and treated the answer as if it addressed the second.
The bottleneck, named properly
This is an information bottleneck, in the precise sense Naftali Tishby, Fernando Pereira and William Bialek gave the term in 1999. Given an input X and a target Y, the bottleneck is the smallest summary T of X that preserves the most information about Y, found by minimising I(X;T) while maximising I(T;Y). The result is optimal compression — for that Y. Change Y and the same T becomes an arbitrary, lossy encoding of the input, no better than noise for the new question.
The annual review's reporting template is a bottleneck fitted to Y = "did this cohort pass and get hired." Against that target, discarding raw engagement telemetry, discarding granular skills-taxonomy data, discarding lecturer-level anecdote, is not sloppy. It is correct. Nobody needs six years of clickstream data to confirm a pass rate. The compression was well designed for the question it was built to answer.
The failure appears when the real governing question is Y' = "does this curriculum still teach what the market values," and the template has no memory of anything that would answer it, because that information was discarded at intake, upstream of the review, in systems that were never asked to preserve it for this purpose. The programme director inherits a summary that is optimal for last year's target and unrecoverable for this year's.
The curriculum was reviewed on schedule, by the correct committee, using the correct forms. Nothing in the process was skipped.
That is true, and it is exactly the problem. Process compliance is not the same as informational sufficiency. The forms were filled in against a Y that had already drifted by the time the ink dried.
Why the fix isn't "measure more"
The obvious response — capture the labour-market signal, wire up the telemetry, shorten the committee cycle — is correct as far as it goes, but it restates the deeper claim rather than escaping it. Any review process, however enriched, still commits to a target at the moment it decides what to compress. The question is not whether to compress; some compression is unavoidable and, done against the right target, genuinely optimal. The question is when the commitment to a target gets made relative to when the target is actually known.
Two objections are worth taking seriously here, because education administrators raise them constantly, correctly.
First: keeping everything is not an option. A large university system generates millions of assessment events, telemetry pings and enrolment records a year; storing and indexing all of it against every conceivable future question is a cost nobody's budget absorbs, and a records office already has to decide retention schedules under data-protection law. This is correct, and the strong version of the claim concedes it fully. The distinction that matters is not "compress versus don't compress." It is between generic retention with known, uniform loss — keeping raw assessment scores and platform logs at the individual level, with timestamps and provenance, even after they're rolled into aggregate reports — and task-specific compression that discards asymmetrically, the way "annual employment rate" discards which specific skills got which specific graduates hired. The first kind of loss is characterisable and roughly uniform across future questions. The second is selective and, once made, unrecoverable: you cannot reconstruct which module content mattered to an employer from a single aggregate percentage.
Second: most curriculum drift is mild, and well-designed core skills transfer. A student trained rigorously in statistical reasoning adapts to a new analytics tool inside a fortnight; the fundamentals were never the stale part. This is often true, and it is why wholesale annual redesign is neither necessary nor sane. But transfer works when the new market demand is close to a function of the old one. It fails precisely at the cases that matter most to a programme director: a tool-specific certification module built around software a vendor discontinued, a compliance unit written for a regulatory regime since repealed, a "digital skills" strand anchored to platforms two cohorts of employers have already moved past. No amount of transferable fundamentals recovers a credential that certifies a tool nobody uses. The review process needs to be able to tell these two cases apart, and a template built only to report pass rates cannot.
What holding the stream open would look like
The alternative is not an unsearchable archive of everything the university ever logged, which is its own well-known failure — a heap with no compression is not knowledge, and a records office drowning in unindexed telemetry serves nobody, including the auditor who eventually has to search it under legal challenge. The alternative is provenance-carrying intake: index at the point of collection — which module, which cohort, which employer signal, which timestamp — and defer the compression that decides what matters until the question that matters has actually arrived. A skills taxonomy feed cross-referenced against module content, kept live rather than sampled every three years, lets a programme director re-fit the "is this still relevant" bottleneck at the moment a hiring manager first stops mentioning a tool, rather than eighteen months after committee minutes catch up.
The lineage claim, reached from here
This is the same structural gap that separates the three generations in the Large Language Model to Large World Model to Large Universe Model lineage. A Large Language Model compresses a frozen corpus against next-token prediction and cannot answer for anything the corpus never encoded past its cutoff. A Large World Model widens intake to a sensed scene but closes the bottleneck at the scene's edge, fitted to acting well now. A curriculum review that samples pass rates once a year and labour-market signals once a triennium is doing the same thing at institutional scale: closing its bottleneck at a boundary chosen for administrative convenience, against a target that will have moved by the time the next boundary arrives.
The Large Universe Model position — streams held open, provenance retained, compression deferred to query time — is not a product available to a programme director tomorrow. It names where this ladder of intake design terminates: the point past which the only remaining move is to stop closing the bottleneck at intake at all, and fit it instead against whichever Y actually turns up. Everything below that point, including annual curriculum review as currently practised almost everywhere, is still choosing its target before it knows what the target will need to be.