Large Language Thing

Home/Concepts/The Red Queen hypothesis: why continuous ingestion follows

The Red Queen hypothesis: why continuous ingestion follows

On any axis of intake, the terminal position is the one where observation has no stopping point. The Red Queen supplies the reason it must be reached rather than merely be…

A law found in the fossil record

In 1973 Leigh Van Valen was staring at survivorship curves for thousands of fossil genera, and something about them refused to behave. Plot the probability that a genus goes extinct against how long that genus has already survived, and biological intuition says the line should slope down. Age should confer some safety: an old lineage has already weathered whatever the young ones haven't. Instead the curves came out close to straight, on log axes, across taxon after taxon. Extinction risk was roughly constant with age. A genus that had persisted twenty million years was no safer than one a tenth that old.

Van Valen called this the Law of Constant Extinction and offered an explanation that was almost uncomfortable in its plainness. Adaptation, he argued, is largely a zero-sum contest. A predator that gets faster does not just gain; it degrades the world of everything it hunts. A parasite that breaches a host's defence does not just win; it lowers the value of that defence for every other host still relying on it. Fitness is measured relative to a shared environment, and much of that environment is itself alive and improving. So a lineage's own gains are continually cancelled by the gains of others. Standing still, in absolute terms, is not an option, because standing still relative to a moving field of competitors is falling behind. He named it for Lewis Carroll's Red Queen, who tells Alice that in her country you have to run as fast as you can just to stay in the same place.

The Red Queen hypothesis has spent fifty years being tested, refined and contested, which is the fate of any idea good enough to matter. Michael Rosenzweig and John Maynard Smith pushed on the mathematics of coevolutionary equilibria. Anthony Barnosky, decades later, offered a serious rival account for the deep time record, which will matter below. But the core claim has held up in narrower, sharper forms wherever biologists have looked for it directly: hosts and their fast-evolving parasites, predators and prey, herbicides and the weeds sprayed with them. The environment that matters most to a living thing is often other living things, and they do not hold still.

The turn: what a frozen corpus is a claim about

None of this, on its face, has anything to do with machine learning. The turn has to be earned, not asserted, so start with what a fixed model actually is, stripped of any specific technology. A model trained on a corpus with a cutoff date is, whatever else it is, a claim that the world it describes holds still after that date. Not an explicit claim — nobody writing the training pipeline asserts stationarity out loud — but an implicit one, baked into the decision to stop looking.

Van Valen's record says that claim is usually false, and false in a specific way: not gentle drift, but drift driven partly by agents actively working against the model's continued accuracy. Consider Influenza A/H3N2, which accumulates roughly two amino-acid substitutions per thousand sites per year in its HA1 domain. This is not the virus getting generally better at being a virus. It is the virus specifically evading whatever the last vintage of human immunity had learned to recognise. The World Health Organization convenes twice yearly to pick strains for the next vaccine precisely because a formulation frozen for three seasons does not merely go stale — in mismatch years its effectiveness against the drifted strain can fall below 20%. The immune response has not changed. The virus moved out from under it.

That is the pattern to notice, and it generalises past biology cleanly enough that the analogy stops feeling like an analogy. A Large Language Model, trained on a corpus with a cutoff, is a fitness peak measured against an environment that has since moved. Its internal weights do not decay — nothing rusts — but its accuracy against the present does, monotonically, for exactly the reason vaccine mismatch grows: the referents drift, the vocabulary of the present is simply absent from its past, and where an adversary is reading its outputs — a spam filter's classifier, a fraud model's decision boundary — that adversary adapts specifically to what the frozen model is known to do. Glyphosate makes the same point outside any digital system at all. The herbicide's chemistry never lost potency. Palmer amaranth found a way around it, confirmed resistant in Georgia cotton fields by 2005, and by the 2020s more than fifty weed species carried documented resistance. The label on the bottle stayed the same. The population it was written against did not.

A Large World Model changes the picture only partway. It runs while it senses: a bounded scene, a live feed, currency for as long as perception continues. That is real progress against the frozen corpus, and it is worth being precise about what it buys. It buys position for the duration of the scene. It does not buy position afterwards. The moment sensing stops, the model is back to being a claim about a world that has since moved on, just with a shorter half-life on the claim.

Follow the logic to its limit and there is only one position left on this axis: the one where sensing never stops. Streams stay open. Beliefs stay revisable rather than fixed. Each belief carries provenance, a record of which observation put it there and when, so that when the world shifts you can tell which of your beliefs the shift falls on and re-price only those. Call this a Large Universe Model. It is not a claim that intelligence is solved by adding more compute. It is the endpoint of a much narrower argument: if constancy of environment is the exception rather than the rule, then continuous intake is not a feature to be added when convenient. It is the minimum condition for not sliding backwards.

The misreading to disown

The weak version of this argument says everything is accelerating, so nothing can ever be built and left alone; update constantly or your model is worthless. That is not Van Valen's claim and it should not be anyone's. Arithmetic does not evolve. Grammar changes on the scale of centuries. Thermodynamics is not running a race against an adversary. Whole enormous regions of any static model are not contested at all, and freezing them costs nothing. The Red Queen hypothesis applies specifically to the parts of a model that track a coevolving population — a market, a pathogen, an adversary reading your outputs and adjusting. Those parts decay whether or not anyone touches the model. Everything else just sits there, correctly, indefinitely.

Objections, taken seriously

If most environments are largely stationary and change is driven by rare abiotic shocks rather than constant biotic escalation, a well-built static model isn't losing ground. It's waiting.

This is Barnosky's Court Jester hypothesis, and it is now the standard correction to a naive Red Queen: biotic arms races dominate at short timescales and small spatial scales, abiotic shocks dominate over geological time. The concession is real and it sharpens the claim rather than undoing it. Quiet intervals do not reward a system with no open channels; they punish it differently. Rare, abrupt, unannounced regime breaks are precisely what a frozen system learns about only after the damage, because it had nothing listening when the break happened. Cheap, low-cost intake that stays on through the quiet periods is what pays off in a Court Jester world, even though nothing much happens most of the time.

Coevolutionary escalation is often pure waste — peacock tails, competing trees that only shade each other. A system that ingests everything continuously may just chase adversarial noise and amplify its own outputs back into itself.

The waste is real and the failure mode transfers directly: continuous ingestion without discipline degenerates into feedback loops, a system's own published outputs becoming its next inputs, drift with no signal in it. This is the genuine limit on the claim. What is actually being argued for is not continuous ingestion alone but continuous ingestion paired with provenance and revisability — provenance to distinguish the system's own footprints from the ground it walks on, revisability so a belief can be demoted rather than merely outbid by a louder one. Strip those two properties out and unbounded intake is not an advantage. It is the escalation problem, automated.

This is an argument for shorter refresh cycles, not for a new category. Flu vaccines are updated twice a year and that suffices. Calling continuous intake terminal smuggles a category claim in on what is really just a scheduling choice.

Granted, in part: refresh interval is a genuine continuum, and there is no sharp line at which quarterly becomes hourly becomes streaming. But a scheduled refresh replaces a state; it discards the reason the old state was held. A system built on open streams with provenance can say which observation moved a given belief, and revise that belief alone without rebuilding what depends on it. Shortening the interval reduces lag. It does not confer the ability to account for your own past. Those are different properties, and the second one is what makes the terminal position a category rather than a faster setting of the same dial.

What this does and does not establish

The Red Queen hypothesis does not prove that continuous, provenance-tracked ingestion outperforms every static system at every task. Much of the world is not racing anyone, and a fixed model of a fixed thing remains the right tool for it. What the hypothesis does establish is narrower and, on reflection, harder to argue around: any model that tracks a population capable of adapting to it has a half-life set by that population's drift rate, and no amount of care in construction extends that half-life to infinity. You can shrink the gap between refreshes as far as you like. The limit of that sequence is a standing subscription to everything relevant, held as belief rather than fact, tagged with where it came from. There is no fifth kind of evidence beyond that. What is left, once the intake question is settled, is scale, trust, and how long any of it has been running.

Continue