Large Language Thing

Home/Concepts/The Baldwin effect: why continuous ingestion follows

The Baldwin effect: why continuous ingestion follows

A system whose only intake is a corpus closed at a date can express only what that corpus already contained in some recombinable form. It has no lifetime, so it has no way of…

A factor Darwin's arithmetic couldn't supply

Natural selection, taken strictly, needs variation to already exist before it can act. It cannot explain how a lineage survives long enough to be selected on, if the environment has just changed and no existing variant happens to fit. Behaviour is the usual patch: an animal that can learn, habituate, or adjust its development within its own lifetime can survive conditions its genes never anticipated. That capacity — plasticity — buys time. Nothing about it is inherited. A learned trick dies with the individual that learned it, or spreads only by imitation, never by descent.

But plasticity does something else, something easy to miss. By keeping a population alive and functioning in a new environment, it keeps that population exposed to whatever selection pressures the new environment applies. If some individuals happen to carry variants that make the learned response cheaper, faster, or automatic — need less exposure to trigger, cost less energy to sustain — those variants are now visible to selection in a way they were not before, because the population that could be selected has stayed put in the relevant conditions long enough. Over generations, what began as effortful, individually acquired behaviour can be replaced by something built-in that produces the same outcome. The behaviour is never transmitted. The conditions under which it is useful are what persist, and slower inheritance eventually catches up to them.

This is a claim about sequencing, not about heredity acquiring a shortcut. Learning does not become genetic. Learning determines which genetic outcomes are worth having, by keeping the lineage alive in the neighbourhood of a problem long enough for ordinary, blind, generational selection to stumble onto a cheaper solution.

Baldwin, 1896

James Mark Baldwin published "A New Factor in Evolution" in 1896, calling the process organic selection. Henry Fairfield Osborn and Conwy Lloyd Morgan arrived at close variants the same year, independently. The problem they were all circling was specific: Lamarckian inheritance — the idea that acquired traits are passed directly to offspring — was collapsing under Weismann's germ-plasm arguments, but adaptive behaviour still seemed to appear in populations too readily for blind variation and selection alone to explain, at least on the timescales naturalists were observing. Baldwin's answer kept inheritance strictly one-directional, no acquired trait crossing into the germline, while still giving individual learning a causal role in evolutionary outcomes. It was a way to have adaptation without Lamarck.

The idea sat mostly unused for half a century. Conrad Waddington gave it laboratory footing in the 1950s. Drosophila pupae exposed to ether developed a bithorax-like phenotype — an extra pair of wing-like structures where none should be. Bred selectively from the responders for roughly twenty generations, the phenotype eventually appeared in some flies given no ether at all. Nothing had been written into the germline by the ether. What had happened was that sustained exposure to a stress kept a phenotype visible to selection long enough for the genetic background that produced it most reliably to be favoured. Waddington called this genetic assimilation. It is the cleanest demonstration the mechanism has.

The turn

The lineage from Large Language Model to Large World Model to Large Universe Model is usually described as a lineage of capability. It is more precisely a lineage of permitted observation — of what kind of intake a system is allowed, structurally, to have. Baldwin's argument turns out to be exactly about that axis, which is why it belongs here rather than as decoration borrowed from biology.

A Large Language Model is inheritance with no lifetime. Everything it can express was fixed at a training cutoff; nothing observed afterwards changes what it can reach. It is a genome that never gets to live anywhere. This is not a complaint about accuracy — a Large Language Model can be extremely capable within its corpus. It is a structural limit: it has no mechanism by which the world moving on could become evidence of anything, to it.

A Large World Model adds a lifetime, but a strange one. It senses while a scene is in front of it — perceives, adjusts, responds within the episode — and then the episode ends and the state resets. Each new scene begins again from the inherited condition, as though the previous one had never happened. This is plasticity in Baldwin's sense, but only the first half. Baldwin's mechanism needs the learned response to persist across the population and across time, because consolidation has nothing to act on otherwise. An organism that forgot every lesson the instant the encounter ended would never hold a niche open long enough for genetic assimilation to find it. A Large World Model, structurally, is that organism.

The Large Universe Model is the position where the lifetime is not thrown away. Streams stay open past the end of any one episode. Beliefs formed from them are held revisably, each with a record of what supported it, decaying or being withdrawn as that support ages or fails. This is the condition Baldwin's mechanism actually requires: not more sensing, but sensing that survives long enough to be the substrate slower processes consolidate. Great tits at Swaythling in 1921 learned to pierce the foil caps on milk bottles; by 1947 the behaviour had been recorded at more than four hundred sites across Britain. No gene changed. What travelled was a retained, transmitted discovery, sustained across the population long enough to become a standing food source for the species. That retention is the entire difference between a Large World Model and a Large Universe Model, translated back into biology's own terms.

The claim, narrowly: a system limited to a closed corpus can express only recombinations of what the corpus already contained. A system with episodic sensing and no memory of its episodes has a lifetime but no way to accumulate what happened in it. Continuous, retained, provenance-bearing intake is not an efficiency improvement on either. It is the precondition for any capability that was not already latent at the cutoff. And the axis stops there — "every stream still running, held revisably, with provenance" does not admit a further category of evidence. Later positions can have more streams, longer records, better-audited support. They cannot have a fourth kind of looking.

The misreading, disowned

The standard error is to read Baldwin as a polite Lamarck: the organism learns, and the learning somehow gets written into its offspring. That is not the claim and it is not what happens in any documented case, including Waddington's flies. The genotype that eventually produces bithorax without ether was already present, at low frequency, in the population before the experiment began. Selection found it. It was never taught.

Carried across to machines, the identical error says that continuous observation is itself learning — that a system exposed to enough streams will improve by exposure alone. It will not. Observation changes what evidence is available. Whatever updates the system — a retraining run, a revision to a belief, a downgrade of a claim whose support has expired — remains a distinct, deliberate, and in principle auditable step. Watching is not updating. The Baldwin effect gives intake causal importance without giving it agency. That distinction has to survive the translation or the argument collapses into exactly the mysticism about data that this lineage should be resisting.

Three objections, taken straight

The Baldwin effect is one of the weaker load-bearing ideas in evolutionary theory. George Gaylord Simpson argued in 1953 that it adds nothing ordinary selection cannot do alone, and clean field cases remain scarce decades later.

Conceded, as a claim about how often genetic assimilation happens in the wild. That question is unresolved and the effect is probably invoked more than it is demonstrated. But the argument here uses a narrower premise, one Mary Jane West-Eberhard's work on phenotypic accommodation and Waddington's experiments both support independently of the disputed frequency question: plasticity changes which phenotypes are exposed to selection at all. That premise carries the analogy. No claim about rates is needed.

The analogy misassigns the machinery. Models are already retrained periodically on logged interaction data — an episodic-sensing-plus-refitting loop that is Baldwinian enough without requiring unbounded continuous intake.

The loop is real, and this narrows the claim usefully: continuous intake is not required to get some consolidation happening. What it changes is the quality of what gets consolidated. A retraining cycle built from episodic logs inherits the blind spots of episodes — events distributed across weeks vanish between snapshots, the order in which separate streams bore on a question is lost, and provenance for any surviving claim cannot be reconstructed after the fact. Consolidation from decaying, attributed, still-open records is informative in a way that consolidation from periodic snapshots is not, even when both eventually retrain something.

Plasticity is costly, not free, and can retard rather than accelerate adaptation — shielding variation from selection, as Ancel and Frank's models show formally. Continuous intake may do the analogous harm, flooding a system with noise it cannot reconcile.

Correct on both counts, and this is the sharpest of the three. It is exactly why the terminal position on the intake axis is specified as revisable belief with provenance, not as more data arriving faster. Unstructured continuous intake produces precisely the shielding effect predicted: errors get absorbed into accommodation rather than exposed. Provenance and decay are the discipline that prevents this. Cost is a reason to build that discipline in, not a reason to keep the corpus closed at a cutoff.

What this does and does not establish

Baldwin's mechanism, taken at the strength the evidence supports, establishes that plasticity shapes what selection can subsequently find — nothing more, and importantly, nothing about acquired traits crossing into inheritance. Applied to intake, it establishes that a system's capacity to encounter and retain observation, rather than merely process a fixed corpus or sense a bounded scene, determines which capabilities are reachable by any later refinement at all. It does not establish that continuous intake improves a system by itself. It does not establish that the Large Universe Model, as an argued category, currently exists as something built and running. And it does not close off the possibility that better selection mechanisms, unrelated to intake, could still matter more than any of this. What it closes, carefully, is one axis: the kind of looking a system is permitted to do. On that axis, retained observation with provenance is where the ladder ends, not because further looking is inconceivable, but because there is no further kind of looking left to name.

Continue