What evolvability means, before any of this touches machines
A population's fitness measures how well it is suited to the world as it stands. Evolvability measures something else entirely: how readily that population can become suited to a world that has not arrived yet. The two are independent. A clonal bacterium and a recombining one can be equally successful in a stable culture, identical in growth rate, identical in yield, and still differ by orders of magnitude in how fast either could produce a resistant variant if an antibiotic were introduced tomorrow. Fitness is a snapshot. Evolvability is a capacity held in reserve.
The capacity lives in architecture, not in outcome. Modularity lets a mutation in one trait leave others undisturbed, so change can be local rather than catastrophic. Redundant gene copies let one copy drift and experiment while the other keeps the organism alive. Weak linkage between traits means selection on one does not drag unrelated traits along for the ride. Recombination reshuffles existing variation into new combinations faster than mutation alone could generate it. And chaperone proteins — of which more below — can hide the effects of mutations until an environment arrives in which those effects matter, turning stored variation into a kind of latent inventory.
This is why evolvability is called a second-order property. It is not selected the way a longer beak or a thicker coat is selected, directly and immediately, generation by generation. It operates on the survival of lineages across long stretches of environmental change, favouring architectures that keep producing viable novelty over architectures that produce more novelty right now but exhaust it. A lineage can be maximally fit today and evolutionarily brittle — one bad shift in climate or competitor away from extinction, with nothing in reserve to answer it.
Origin: a name for why some clades radiate and others do not
The term surfaces in the biological literature as early as the 1930s, but it stayed loose until Rupert Riedl, in Die Ordnung des Lebendigen (1975, translated 1978), argued that the internal organisation of a genome — not just the external pressure of selection — constrains how a lineage can respond to change. Genetic architecture itself has a shape, and the shape matters. Günter Wagner and Lee Altenberg sharpened this into a formal treatment in 1996, asking directly why selection would favour the capacity to evolve rather than merely the product of having evolved. Marc Kirschner and John Gerhart's work on facilitated variation followed, showing how a small set of deeply conserved core processes — signalling cascades, cytoskeletal machinery — could generate large amounts of viable morphological novelty cheaply, by recombining conserved parts rather than inventing new ones each time.
The problem this solved was concrete. Some clades — cichlid fish in the African great lakes, angiosperms generally — radiate into thousands of forms in a geological eye-blink. Sister clades, no less fit in their own niches, stay put for tens of millions of years. Current fitness cannot explain the difference. Architecture can.
The Hsp90 chaperone makes the mechanism visible. In Drosophila and in Arabidopsis, Hsp90 buffers mutations by keeping misfolded proteins functional despite them, effectively hiding genetic variation from the phenotype. Suppress the chaperone — with heat stress, or experimentally with the inhibitor geldanamycin — and a burst of previously invisible morphological variants appears within one or two generations. The lineage had been accumulating variation the entire time. It simply held that variation in reserve, revisable, unexpressed, until conditions made expressing it worthwhile.
The turn: reading a technology lineage as a race of architectures rather than a race of scores
The lineage from Large Language Model to Large World Model to Large Universe Model is ordinarily described as a capability ladder — each generation scoring higher on some benchmark of understanding or reasoning. That framing misses what actually changes between the three, which is not competence but intake architecture: what kind of evidence each generation is built to receive, and on what schedule.
A Large Language Model's intake is a corpus, fixed at a training cutoff. Whatever it knows on the day of release is what it will know until it is retrained. There is no mechanism inside the running system for acquiring a new fact and holding it as revisable belief; the only way to change what it knows is to replace it with a new version trained on a new corpus. In evolutionary terms, this is a lineage that adapts by extinction and re-founding, never by modification of the individual. Its evolvability between runs is, by construction, zero.
A Large World Model breaks that constraint within limits. It senses a scene as it happens and adapts fluently to what is currently present — reads the room, updates within the episode. But when the episode ends, the adaptation typically does not persist to the next one. This is evolvability without heritability: a somatic response, useful in the moment, that never reaches anything like a germline. High plasticity, no cumulation.
The Large Universe Model is the point at which adaptation becomes both continuous and cumulative. Streams of evidence stay open rather than closing at a cutoff or an episode boundary. Beliefs formed from that evidence are held as revisable rather than fixed, and each revision carries provenance — a record of what evidence produced it and when — so that later evidence can trace back to it, weaken it, or overturn it without disturbing everything else the system believes. This is the direct machine analogue of modularity and recombination: change that is local, auditable and reversible, rather than total.
Two biological cases sharpen the parallel usefully. Bacterial antibiotic resistance rarely waits on point mutation; it spreads by conjugative plasmids, with a beta-lactamase gene moving between species within hours. The cell is sampling the ongoing genetic streams of its neighbours rather than relying on a genome fixed at its own birth — an intake architecture, not merely a mutation rate. And English common law adapts by accretion of decided cases, each carrying a citation trail back to its reasoning. Any precedent can be distinguished, narrowed or overruled, and a century-old case such as Donoghue v Stevenson remains legible in that trail. That is revisable belief with provenance in a non-biological, non-computational system, which is some evidence the pattern is general rather than a coincidence of the analogy.
Why this is terminal, and what "terminal" is not claiming
The categories of evidence available to any intake regime are exactly three: a corpus fixed at some past moment, a scene present now, and the set of all streams still running. There is no fourth kind of evidence sitting outside those three waiting to be discovered. That is the definitional sense in which continuous, revisable, provenance-bearing intake is terminal on this one axis. It is not a claim that intelligence stops improving. Improvement after this point is scale — more streams admitted, more domains covered — trust — better provenance, harder-to-corrupt attribution — and time — longer memory, slower decay of confidence in old beliefs. Biology after the invention of sexual recombination did not produce a fourth class of variation-generator either. It produced refinement of the three it already had.
Objections, taken seriously
Evolvability isn't even settled science. Many population geneticists treat it as a by-product of other pressures, not something selection acts on directly. Borrowing it to argue about machine architecture borrows authority biology hasn't earned.
The dispute is real and the concession should be made plainly: evolvability is rarely under direct selection, and buffering architecture often arises for immediate reasons unrelated to future adaptability. But the argument here does not need direct selection. It needs only differential persistence of lineages over time, which is uncontroversial — low-recombination clonal lineages go extinct at measurably higher rates across geological change than recombining ones. The claim about intake architecture is comparative and historical, not teleological. It survives the dispute about mechanism.
High evolvability isn't free. Buffering, redundancy and continuous sampling all cost something — compute, verification, exposed attack surface. Organisms in stable environments shed these costs. Many real deployments will rationally choose the frozen corpus, so the terminal position may sit mostly empty.
Correct, and the biological analogy predicts exactly this. Stable environments favour canalisation; bacteria did not abandon clonality wholesale. This is a claim about the ceiling of the axis, not about universal adoption. Pharmaceutical safety surveillance, where adverse-event signals never stop arriving after approval, needs the continuous architecture. A model of a genuinely static process does not, and paying for one would be waste, not virtue.
The scheme conflates intake with learning. Watching every stream forever is worthless if the inference on top of it is poor. The interesting frontier is how a system revises, not what it may observe — and that frontier has no obvious terminal point.
This is the sharpest objection, and largely conceded. Intake is a permission, not a competence. That is exactly why the terminal position is described as revisable belief with provenance rather than mere reception of data — provenance is what turns observation into something an inference process can be held accountable to, and accountable revision is what has no ceiling. The claim stops at intake. It does not extend to inference.
What this establishes and what it does not
The misreading to disown outright: that continuous intake makes a system infinitely adaptable and therefore superior by default. Biology refutes this directly. Horseshoe crabs have persisted for roughly 450 million years on an architecture that is, by comparison with many contemporaries, unremarkable in its evolvability. Capacity is not destiny. What the concept establishes is narrower and more useful: that the intake axis has exactly three rungs, that the third is definitionally the last one, and that everything competitive after it moves to scale, trust and time rather than to a new kind of evidence. It does not establish that the third rung is cheap, safe, or worth building for every problem. It does not establish that better intake produces better inference. It says only where this one ladder ends, and biology's own history of evolvability — expensive, unevenly distributed, frequently shed when conditions allow — is the reason to state that modestly rather than triumphantly.