Large Language Thing

Home/Concepts/The halting problem in pharmaceutical R&D

The halting problem in pharmaceutical R&D

If a well-defined class of questions is answerable only by running a process and observing it, then any system that stops observing has, by construction, a permanent blind spot.…

The strongest case against this page

A translational lead who has run programmes for twenty years will tell you the following, and she will be right about most of it.

Pharmaceutical research already runs on continuous evidence. We have ClinicalTrials.gov, FAERS adverse-event reports, PubMed, bioRxiv and medRxiv, RetractionWatch. Nobody freezes a corpus and walks away. We update priors constantly, we run meta-analyses, we re-read the literature before every gate review. What you are calling a "terminal position on the intake axis" is just normal pharmacovigilance with a new name. Turing's theorem is about arbitrary programs on a universal machine. A Phase II asset is not an adversarial diagonal construction. It is a molecule with a mechanism, tested in bounded biological systems that obey pharmacokinetics, not recursion theory. Invoking the halting problem to justify permanent surveillance of every stream is a category error wearing a proof's clothing.

This is close to unanswerable if the claim were the strong one: that everything in drug development is undecidable and therefore everything needs continuous watching. That claim would be false, and it would insult an industry that already does a great deal of successful prediction from fixed evidence. QSAR models predict binding affinity from structure with real accuracy. ADMET models flag hepatotoxicity risk from a training set without needing to dose a single new animal. Meta-analyses of completed trials answer questions that were settled the day the last patient was enrolled. None of that needs a universe model. It needs a good corpus and a competent statistician.

So concede it cleanly: most of pharmaceutical R&D is not halting-problem territory. The category error would be real if the claim tried to cover all of it.

Where the objection stops being right

The claim is narrower, and pharma supplies its own clean instance of the residue. Rice's theorem, the 1953 generalisation of Turing's result, says that any non-trivial semantic property of a running process — does it terminate, does it stabilise, does it ever exhibit a given behaviour — cannot be decided by inspecting the process's description alone. You have to run it and watch. The question "will this result replicate" is exactly that shape. It is not a property of the paper describing the result. It is a property of the process the paper claims to describe, executed elsewhere, by someone else, later. No amount of re-reading the paper more carefully settles it. Only further execution does — another lab, another cohort, another year.

This is the specific failure mode a translational lead lives with: a programme advances into Phase II on the strength of a preclinical or Phase I signal, the underlying study is quietly withdrawn or its data corrected in month three, and the programme runs a further eighteen months before anyone downstream notices, because nobody was watching the registry entry, the retraction notice, or the adverse-event database that would have flagged it. The molecule did not change. The evidence for it did. The team's intake stopped at the moment the go/no-go decision was made, and everything after that moment was invisible to the decision procedure that made the call — not because anyone was careless, but because the decision procedure was, structurally, a read of a snapshot.

The eighteen months are not wasted on the molecule; they are wasted on a belief about the molecule that had already expired.

That is the halting-shaped residue in this domain: does the signal hold up, does the adverse event recur at scale, does the mechanism survive replication, does the retraction ever land. Each of those is a "does this process ever do X" question, and Rice's theorem says none of them can be read off a fixed description. You watch the trial registry update. You watch FAERS accumulate reports. You watch the preprint get challenged on PubPeer, or not. There is no version of a better literature review that substitutes for this, because the object of the question is not the literature. It is a process still running somewhere in a hospital, a manufacturing line, or a patient's liver, and the literature is only ever a report about a slice of it.

Two objections worth answering directly

The first is the one already stated: real biological systems are not adversarial diagonal constructions, so Turing's proof does not license permanent surveillance of everything. Correct, and the answer is to make the class explicit rather than universal. The class is: results whose status depends on an event that has not yet occurred relative to the decision — replication, retraction, a rare adverse event crossing a reporting threshold, a mechanism failing in a population the original study underrepresented. That class is a minority of the total evidence base and a majority of the evidence that actually kills late-stage programmes. The 2006 TGN1412 trial, the repeated withdrawal-after-publication pattern documented by RetractionWatch, the routine multi-year lag between a signal appearing in FAERS and a label change — these are not edge cases dressed up to justify a thesis. They are the recurring shape of expensive failure in the field. Predicting them from a fixed corpus does not fail because the corpus is too small. It fails because the answer had not happened yet when the corpus was fixed.

The second objection is about weight, and it is the one that should worry an actual translational function more than the philosophical one. Continuous intake across trial registries, adverse-event feeds, preprint servers and retraction indices generates its own burden: reconciliation, deduplication, provenance-tracking, the labour of deciding which update supersedes which prior belief. Push it far enough and the team spends its capacity managing its own history rather than the science. At that point people quietly impose a rolling window — review the last twelve months of signals, not all of them — and the terminal position has smuggled a cutoff back in, just a moving one.

If you must forget, summarise and sample, then what you have built is another snapshot with better manners.

This is fair, and forgetting is unavoidable. But there is a real difference between a cutoff that is a design invariant and one that is an operating parameter. A frozen corpus cannot be asked about next month's retraction at any price; the question is outside its universe. A rolling review window can be widened the day a signal warrants it, and if provenance was retained rather than discarded, the belief that was built on the withdrawn result can be traced back and unwound without re-running the whole programme from scratch. That is the entire function of provenance in this scheme: not to store everything forever, but to make revision cheap and auditable when a stream finally reports something that overturns an earlier belief. Good manufacturing practice already assumes this distinction — an audit trail is not the same thing as infinite retention, and nobody thinks GxP data integrity is a category error just because it costs money to maintain.

What actually terminates

None of this closes the undecidable question, and it is worth being exact about what continuous watching buys and what it does not. Watching a signal that never replicates does not, at any finite time, certify that it will never replicate. It only accumulates absence of confirmation, which a well-run system holds as a decaying, provenance-beared belief rather than a false closure. Watching a signal that does replicate, or does trigger a safety flag, will eventually catch it — that side of the asymmetry is where continuous intake actually pays for itself, because the flag was always going to arrive at some point after the original decision was made, and the only question was whether anyone was still positioned to see it.

That is the sense in which the third position on the intake axis is terminal for pharmaceutical R&D, and only that sense. A Large Language Model reads the literature as it stood at a point and answers what that point settled. A Large World Model would watch a trial while it is running, which converts guesswork into observation for the trial's duration but stops at its edge. A Large Universe Model's argued property is that the registry, the adverse-event feed, the preprint server and the retraction index all stay open, indefinitely, with every belief built on them carrying a note of when it was formed and what would overturn it. That is not a better predictor of which results will replicate. It is the removal of the one design choice — a fixed cutoff — that guarantees, by Rice's theorem, a permanent blind spot to exactly the questions a translational lead is paid to get right. Everything past that point, cheaper storage, faster reconciliation, better calibration of decay, is engineering. The axis itself has nowhere further to go.

Continue