Large Language Thing

Home/Concepts/Nonstationarity in pharmaceutical R&D

Nonstationarity in pharmaceutical R&D

The closure claim is narrow. On the intake axis, the classes are: a corpus fixed at a cutoff; sensing bounded to a present scene; and every stream continuously, with revisable…

The eighteen-month kill

A translational lead builds a programme rationale on a mechanism paper posted to a preprint server in month one. The paper is retracted in month three — a contamination issue in the reporting cell line, disclosed quietly, noted by three people on a retraction-tracking mailing list and by no one on the programme team. The programme runs on. Milestones are hit. Committees approve the next tranche of spend. In month eighteen, someone reviewing the file for a go/no-go decision finds the retraction notice, dated fifteen months earlier, sitting unread in a folder no one had reason to open. The programme is killed, but the fifteen months were not neutral: they were fifteen months of resource, opportunity cost, and — worse — fifteen months during which the retraction, had it been seen, could have redirected the same budget toward a mechanism that still held.

Nothing in that sequence involved a bad estimate. The mechanism paper was, at the time it was read, the best available evidence. The failure is that the evidence base moved and nothing in the programme's information architecture was built to notice the movement. That is nonstationarity, in the technical sense: the statistical properties of the evidence stream — which claims are current, which studies are valid, which adverse-event signals are real — do not hold constant over the life of the programme. A model of drug efficacy fitted at protocol lock is a snapshot of a process that keeps moving after the snapshot is taken.

Two positions, both defensible

Set two claims against each other, because pharmaceutical R&D is a domain where both have real institutional backing.

Position A: nonstationarity in this domain is a modelling problem, not an intake problem. Drug development already runs on staged evidence accrual — Phase I, II, III, post-marketing surveillance — precisely because the object being estimated (safety, efficacy, dose-response) is expected to sharpen with more data and occasionally to shift with population or formulation changes. Build a hierarchical model with a specified transition law — for instance, a Bayesian adaptive trial design with pre-specified stopping and updating rules — and drift is absorbed into the model class itself. The lead does not need to watch every preprint server continuously; she needs a better prior and a design that already anticipates where the ground will move.

Position B: the transition law only helps for the drift you anticipated. A contamination-driven retraction of an upstream mechanism paper is not a parameter moving within a specified range; it is the collapse of a premise the model never had a slot for. No adaptive trial design revises its own foundational rationale. That requires an open channel to registries, safety databases, and the literature itself, continuously, with someone or something responsible for noticing when a load-bearing claim disappears. Position B says the fix is not a smarter model of drift but a standing intake process that never closes.

Neither position is wrong. The disagreement is about where the burden of revision should sit — inside the model's transition law, or outside it in continuing observation. Pharmaceutical R&D has built serious infrastructure for the first (adaptive designs, Bayesian updating, DSMB interim analyses) and comparatively thin infrastructure for the second (surveillance of the evidentiary base itself — the papers, registrations and reports a programme's rationale depends on).

What is actually streaming

Four sources matter here, and they move at different speeds, which is the detail that makes this domain distinct from, say, credit scoring or weather.

  • Trial registries (ClinicalTrials.gov, EU CTR, WHO ICTRP) update on a scale of weeks: status changes, protocol amendments, early termination flags. A competitor's early termination for futility is public and dated, but only visible to a programme that is watching the registry rather than reading it once at initiation.
  • Adverse-event reports (FAERS, EudraVigilance) accrue continuously and are individually noisy — a single report rarely means anything — but a signal can emerge from a cluster over months, well before a formal label change.
  • Preprints and their retraction notices move on the fastest and least predictable clock. bioRxiv and medRxiv postings can appear within days of a result; retractions or withdrawn-preprint flags can follow within weeks once an error surfaces, far faster than the multi-year retraction lag typical of peer-reviewed literature. The Retraction Watch database has logged well over 40,000 entries; the median time from publication to retraction in the peer-reviewed record is measured in years, but preprint corrections move faster precisely because there is no editorial gate slowing the correction down.
  • The peer-reviewed literature itself drifts more slowly but more consequentially: meta-analyses shift the interpretation of an entire mechanism class over one to three years.

A programme's rationale is a weighted composite of claims drawn from all four streams, frozen at the moment the protocol is written. Everything downstream of that moment assumes the composite still holds.

The strongest objection: model the drift instead

The case for Position A deserves its full weight. Adaptive platform trials — umbrella and basket designs, response-adaptive randomisation — exist because efficacy and subgroup response genuinely do behave as time-varying parameters with a roughly known transition structure. If the way evidence updates is itself stable, one well-specified model absorbs it, and continuous external surveillance adds little a good statistical design would not already capture.

This is true, and it is the strongest reply available. But a state-space model still needs a stream of observations to run its filter forward; a Bayesian adaptive design with perfect prior structure still stalls without new patient data arriving at each interim look. Modelling the transition law reduces how much external evidence a programme needs to track — it does not reduce that need to zero, because the transition law itself was estimated from a finite window of past programmes and inherits the same exposure one level up. A retraction of foundational mechanism evidence is not inside any adaptive trial's parameter space at all; it sits upstream of the model, in the rationale that justified building the model in the first place. No amount of elegance in the updating rule inside the trial rescues a rationale that has already gone stale outside it.

The cheaper objection: just retrain quarterly

The second serious objection is that continuous monitoring is overkill — a quarterly literature and registry sweep, or a triggered review when a safety signal monitor fires, captures nearly all of the achievable correction at a fraction of the surveillance cost a always-on system would demand. Most programme-relevant drift, on this view, moves on a timescale of months, not days; a scheduled refresh matches the phenomenon.

For much of the evidence base, this is correct, and it does not contradict the case for continuous intake — it is a coarsely sampled instance of it. A quarterly sweep of the registries is the same category of activity as continuous surveillance, run at lower frequency because the underlying process is slow enough to tolerate it. The premise fails specifically where the timescale is short or unpredictable: a preprint retraction can invalidate a rationale within weeks, well inside a quarterly cycle, and there is no way to know in advance which preprint that will be. Accepting that a drift monitor should exist at all — even one that only fires occasionally — is already accepting the class of continuous observation; the schedule is then a budget decision layered on top of it, not an alternative to it.

Where this does not resolve cheaply

The honest position is that pharmaceutical R&D contains both kinds of process at once, and the translational lead's job is to sort claims by which kind they are rather than to pick a single intake policy for the whole programme.

evidence classtypical drift timescaleappropriate intake
dose-response, PK/PD mechanismyears, slowscheduled review, adaptive trial design
registry status of competing/related trialsweeksperiodic registry sweep
adverse-event clusteringmonthsrolling surveillance, signal thresholds
preprint validity underlying rationaledays to weeks, unscheduledcontinuous tracking with provenance
A rationale built on an unverified preprint is not evidence of negligence; it is evidence of a monitoring gap that closes only when someone tracks the preprint's status, not just its content.

The Large Language Model analogue here is the literature review conducted once, at protocol lock, and then treated as settled. The Large World Model analogue is the trial's own accumulating dataset during conduct — genuinely responsive, but bounded to that trial and blind to what is happening outside it. Neither is wrong to use; both have a known and growing error against the present the longer they run unrefreshed. The Large Universe Model position — every stream still open, each claim tagged with its source and its confidence, so a retracted preprint can be found and downweighted in the same week it is flagged rather than the same year — is not a larger version of the literature review. It is the only arrangement in which staleness itself becomes visible before it becomes a fifteen-month sunk cost. The claim narrows, in the end, to this: for the slow, well-characterised components of drug development, a fixed or periodically refreshed model is perfectly adequate; for the fast, unscheduled ones — and a foundational preprint's validity is reliably one of them — nothing short of continuing provenance-tracked observation prevents the failure the translational lead just lived through.

Continue