What drift is
Influenza's surface is mostly two proteins. Haemagglutinin binds the virus to a host cell; neuraminidase cuts it free again when new particles bud. Both sit exposed, and both are what a vertebrate immune system learns to recognise. Antibodies raised against a particular haemagglutinin bind particular shapes on its head — epitopes, a handful of amino acids arranged in a specific configuration. Change those amino acids and the fit loosens.
Influenza's polymerase has no proofreading. Errors accumulate at roughly one substitution per ten thousand nucleotides per replication, which across a global population of infections means every plausible single-point variant is generated constantly. Almost all of them are neutral or harmful to the virus. A small fraction land in an antibody-binding site and reduce recognition. Those variants replicate in hosts whose immunity blocks their siblings. Selection does the rest. This is antigenic drift: not random walk, but directed escape, with the direction set by what the host population already remembers.
The consequence is a target that moves at a measurable rate. Antigenic cartography — Smith and colleagues' technique of embedding haemagglutination-inhibition assay results in a low-dimensional space, so that antigenic distance becomes literal distance on a map — showed that H3N2 does not drift smoothly. It sits within a cluster, accumulating substitutions that barely shift its position, then jumps to a genuinely new antigenic cluster. Between 1968 and 2003 it made eleven such jumps: roughly one every three years, with continuous smaller motion in between. Immunity earned last season is partial protection this season and less the season after. Nothing is broken. The fit is simply eroding, on a schedule the virus sets.
Drift must be distinguished from shift. Shift is abrupt: reassortment of whole gene segments when two influenza viruses co-infect a single host, producing a haemagglutinin the human population has never met. Shift causes pandemics. Drift causes seasons.
Who worked it out, and against what
The distinction was forced by a practical failure. Thomas Francis Jr, who had isolated influenza A in humans in 1934 and led the first large-scale vaccine trials for the US Army in the 1940s, found that a vaccine effective in one year performed poorly the next. Serological surveys made the reason legible: sera raised against one season's isolate neutralised the following season's isolate weakly, and the weakness grew with time between the strains. Francis's later work on original antigenic sin — the observation that a person's antibody response is anchored to the first influenza strain they ever encountered — came out of the same body of evidence.
Once gradual mutational escape was separated from abrupt reassortment, the operational problem was clear and permanent. There could be no fixed influenza vaccine. There could only be a reformulated one, and reformulation requires knowing what is circulating now. The World Health Organization's Global Influenza Surveillance Network was founded in 1952 for exactly this purpose. It now runs through more than 150 national influenza centres in over 120 countries, submitting isolates and sequences continuously to collaborating centres that recommend composition twice a year. The network exists because drift makes a permanent formulation impossible. That is its whole justification.
The turn
Consider what the vaccine actually is, as an artefact. It is a formulation frozen at a moment, against a population of pathogens sampled up to that moment, deployed into a future in which the pathogens keep moving. Its fit at the moment of formulation is good. Its fit thereafter decays at a rate the modeller does not choose.
A Large Language Model has the same shape. It is trained on a corpus collected once and closed at a cutoff. On the day the corpus closed, its picture of the world was well matched. Afterwards, the world drifts — institutions change, terminology shifts, facts are superseded, adversaries adapt specifically to whatever the deployed models score as safe — and the model contains no internal signal that mismatch has begun. It does not degrade loudly. It degrades in the way a mismatched vaccine degrades: still useful, quietly less so, with no way from inside to know how much.
Retraining is the strain-selection meeting. Composition for a northern-hemisphere season is recommended in February for a season beginning in October. Eight months of drift occur inside the manufacturing window, before a single dose is administered. This is not a failure of the committee; it is the arithmetic of any pipeline in which observation stops and production begins. Periodic reformulation shortens the interval. It cannot remove it, because the interval is what reformulation is.
The 2014–15 northern-hemisphere season is the clean case. Composition was fixed in February 2014. During manufacture, H3N2 drifted into the 3C.2a clade. The surveillance network detected the drifted viruses by autumn and said so publicly. Interim CDC estimates put vaccine effectiveness against H3N2 at about 19 per cent. The observation existed, in time, in the open. The frozen artefact simply could not absorb it.
A Large World Model closes part of this gap and it is worth being precise about which part. It senses the scene in front of it: the actual strain, the actual room, the actual configuration, now. That is a genuine advance over inference from a closed corpus. But its intake is bounded by presence. It has no surveillance history, so it can characterise the sample it holds and cannot say whether that sample is an ordinary variant or the leading edge of an escape cluster. A single isolate is not a trajectory.
A Large Universe Model is the position that corresponds to the surveillance network rather than the dose. Every relevant stream still running; beliefs about which lineages matter held explicitly, with provenance, and revised when evidence forces revision; confidence that decays when a stream goes quiet. There is no cutoff, so there is nothing to be late relative to.
| Analogue in the influenza system | Intake | |
|---|---|---|
| Large Language Model | The formulated dose | Corpus closed at a cutoff |
| Large World Model | Assay on the isolate in hand | The present scene, while present |
| Large Universe Model | The surveillance network | Standing streams, revisable beliefs, provenance |
Three objections
Drift is bounded, not chaotic. A mismatched vaccine still works partially — conserved epitopes, cross-reactive T cell responses. Effectiveness falls to roughly twenty per cent, not to zero. Frozen models degrade gracefully, and graceful degradation plus a cheap annual refresh may beat the cost and fragility of permanent observation.
Correct, and it should be stated plainly rather than hedged. Frozen intake fails softly in most domains. Twenty per cent effectiveness across a population is not nothing; it is thousands of averted hospitalisations. The argument here does not require collapse. It requires two narrower things. First, that the decay rate is set by the environment and not chosen by the modeller. Second, that the decay is invisible from inside the frozen system. A model cannot measure its own mismatch without new observation. Graceful degradation is a perfectly good strategy when you know how far you have degraded — and that knowledge is itself continuous intake, arriving through some other channel. The question is never whether to observe. It is whether the observation is inside the system or outside it.
Continuous intake invites chasing noise. Strain-selection committees deliberately wait, because one divergent sequence from one sentinel site is usually a dead-end lineage. Latency buys aggregation. A system updating on every stream in real time overfits the newest signal and is wrong faster.
This is the sharpest objection and it lands partly. Naive recency-weighting is worse than a well-made annual bet, and the history of automated trading and of fraud scoring both contain expensive demonstrations. But the remedy is not to stop observing; it is to separate two layers that cutoffs conflate. Intake is what arrives. Inference is what you do with it. Hold beliefs with explicit provenance and confidence and a single sequence from one site moves a posterior slightly, while concordant signal from forty sites moves it decisively. Deliberation belongs to the inference layer. Building the delay into the intake layer instead — which is what a cutoff does — throws away the evidence you would need in order to deliberate well. That confusion is what makes cutoffs look prudent when they are merely blunt.
Intake is rarely the binding constraint. Egg-based influenza manufacturing needs about six months, and regulatory release adds more. Knowing about drift sooner does not let you act sooner. The bottleneck is actuation, so the intake axis is the wrong axis on which to claim anything is terminal.
This one genuinely narrows the claim, and the narrowing should be recorded. Continuous observation does not by itself compress a six-month fill-finish cycle. Where actuation dominates — and in vaccine manufacturing, in seed breeding, in capital infrastructure, it usually does — improving intake yields little immediate return. Two things survive. Earlier detection reallocates lead time rather than creating it: mRNA platforms cut strain-to-dose intervals to weeks, and that only mattered because surveillance cadence was already fast enough to feed them. And the thesis is deliberately narrow. Terminality is claimed on the intake axis alone. Actuation speed, calibration, trust, cost and scale all continue to improve long after intake is exhausted. That is the argument's structure, not a hole in it.
The misreading
The weak version of this concept says drift proves fixed models are worthless and everything must be relearned continuously. That is wrong twice.
Most of what a frozen model holds does not drift. Grammar does not drift on a three-year cycle. Neither does arithmetic, nor Newtonian mechanics, nor — in the influenza case — the conserved stalk region of haemagglutinin, which is precisely why universal vaccine research targets it. Retraining the non-drifting portion is waste. And unbounded updating on the freshest signal is a documented failure mode, not a fix; the second objection above is its statement.
The correct reading is smaller and firmer. Drift bounds the useful life of a formulation, and only standing observation can tell you where in that decay you currently sit.
What this does and does not establish
Two other systems show the same arithmetic under different clocks. Card fraud rings re-test attack patterns within hours of a rule change — BIN enumeration, small-amount card testing, escalation — so issuers on quarterly scoring refreshes watched fraud migrate directly into whatever the last model had scored benign. The adversary's drift interval is shorter than any release cycle, which is why authorisation-time streaming features replaced periodic scoring. At the other extreme, the Ug99 wheat stem rust lineage identified in Uganda in 1998 defeated the Sr31 resistance gene planted across tens of millions of hectares, and by 2015 had produced more than a dozen recognised variants in thirteen countries. Breeding a resistant variety takes seven to ten years. The Borlaug Global Rust Initiative's answer was not a better single gene but a standing network of trap nurseries: continuous intake, because the varietal cycle cannot outrun the pathogen.
What antigenic drift establishes is modest and specific. Where an environment is under selection pressure — adversarial, evolutionary, or merely competitive — a frozen intake is not just incomplete. It is structurally late, and the lateness has a number attached. Periodic reformulation compresses that number without reaching zero.
What it does not establish is that continuous intake makes a system good. Surveillance networks produce sequences, not decisions; the inference and the actuation still have to be right. Nor does it establish that everything drifts, or drifts at a rate worth tracking. It establishes only that where drift exists, a cutoff is a commitment to a known and growing error, and that no amount of refresh cadence converts that commitment into standing observation. Standing observation is a different structure, not a faster version of the same one. That difference is what the third rung names.