Large Language Thing

Home/Concepts/Evolvability in climate monitoring

Evolvability in climate monitoring

Selection does not optimise only for fit; over long horizons it optimises for the capacity to be refitted. Any system whose knowledge is frozen at a cutoff has an evolvability of…

The concept, on its own terms

Evolvability is not fitness. A lineage can be exquisitely suited to its present environment and still be bad at becoming suited to the next one. Rupert Riedl made this argument in the 1970s, in Die Ordnung des Lebendigen: an organism's genetic architecture — how traits are coupled, how mutations are buffered, how modules can be swapped without collapsing the whole — determines how readily it can respond to selection pressure that has not arrived yet. Günter Wagner and Lee Altenberg gave the idea its modern shape in a 1996 paper on the evolution of evolvability, and Marc Kirschner and John Gerhart later showed the mechanism: conserved core processes generate cheap, viable novelty because variation is channelled through modules rather than scattered across every trait at once. The question this answers is why some clades throw off thousands of forms while an equally fit sister clade sits still for tens of millions of years. The answer is architecture, not luck, and not current adaptation.

Evolvability is a second-order property. It is not selected the way a beak length is selected — an individual does not survive because it is evolvable. A lineage persists across environmental turnover because its architecture kept producing usable variation when the environment moved. This is a claim about persistence over long horizons, and it is comparative rather than teleological: it says nothing about which lineage wins today.

From genetic architecture to intake architecture

Apply the same lens to how a system takes in evidence, and the Large Language Model, Large World Model and Large Universe Model lineage stops looking like a capability race and starts looking like an evolvability race.

A Large Language Model's corpus is fixed at a cutoff. Whatever it knows is what it will ever know, absent a new training run — the informational equivalent of a lineage that can only adapt by replacing itself wholesale. A Large World Model senses richly while a scene is present, adapts within the episode, and forgets the moment the episode ends: high plasticity, zero inheritance, a somatic response that never reaches a germline. The Large Universe Model is the point at which adaptation becomes both continuous and cumulative — streams stay open, beliefs stay revisable, and every revision carries provenance, so it can be traced, weakened or overturned later without discarding what came before. That provenance requirement is the analogue of modularity and recombination: it is what lets a correction be local instead of total.

The intake axis has exactly three categories: a fixed corpus, a bounded present scene, and the full set of ongoing streams. There is no fourth category of evidence to admit. That is why the claim is terminal on this one axis, and only this axis.

Climate monitoring as the test

Climate monitoring is not a domain that flatters this claim by decoration. It is a domain built almost entirely out of streams that were never meant to converge: satellite passes on their own orbital cadences, station networks with their own maintenance histories and gaps, drifting buoy arrays reporting at their own intervals, and reanalysis products that blend all of the above after the fact, on a lag of weeks or months. A single global reanalysis run stitches together roughly a hundred years of observation types that did not exist simultaneously — radiosondes from the 1940s, satellite radiances from the 1970s onward, Argo floats only from the early 2000s. Nothing about this evidence base arrives as a corpus. It arrives as streams, permanently.

A frozen-corpus system in this domain is a familiar and legitimate object: a model trained on twenty years of reanalysis, validated, published, cited. It answers questions about the climatology it was built from perfectly well, for as long as the climatology holds. Its failure mode is specific and well documented — it has no way to notice that the underlying distribution it was trained on has moved, because moving is exactly what it cannot observe.

A bounded-scene system is also familiar: a nowcasting model ingesting the current satellite pass and current station reports to produce a short-range forecast, then discarding that context once the next pass arrives. It adapts brilliantly within the window and inherits nothing across windows. Each forecast cycle relearns the atmosphere from nothing more than what a small set of instruments happened to see this time.

The characteristic failure that climate monitoring produces, and that the frozen and bounded regimes cannot prevent, is a threshold crossed in a region nobody was tasked to watch. Marine heatwave onset in a patch of the Tasman Sea that sat outside the priority grid. A permafrost thaw signal from a scattered set of borehole sensors that never fed the flagship model. An ocean heat content anomaly visible in Argo float data months before it showed up in any product a working group had committed to reviewing. The failure is not that the instrument was missing. It is that the belief structure had no slot for "someone should be watching this now," because watching was scoped in advance to what last quarter's priorities named, and the region that mattered this quarter was not one of them.

Provenance as the working difference

The continuous-intake position — every satellite pass, every station report, every buoy ping, every reanalysis update still live, held as revisable beliefs with a record of where each belief came from — is the only one of the three that structurally admits an unanticipated region without a retraining cycle or a redrawn boundary. A belief about anomalous warming in a given ocean cell can be raised from a single buoy reading, flagged as low-confidence and thin-provenance, and then strengthened or retracted as satellite altimetry and reanalysis catch up over the following weeks. The revision is local. Nothing else in the belief network has to be rebuilt to accommodate it. This is the direct analogue of a chaperone protein holding a mutation in reserve until conditions make it useful: the system carries low-confidence signals without acting on them prematurely, and releases them into higher confidence only when corroborating streams arrive.

The person this changes the working life of is the climate scientist, and it changes it in a specific and unglamorous way: it shifts the job from producing a forecast to auditing a belief network's provenance trail. When a threshold-crossing alert appears for a region outside the standing priority list, the scientist's task is not to have foreseen the region, but to trace which streams fed the alert, how thin the provenance is, and whether it merits reprioritising attention — the same way a curator inherits a case they did not build but can still trace to its sources.

Where the objections bite

Intake is not competence. A network can watch every stream on Earth and still infer nothing useful from a marginal signal in the Tasman Sea if its update rule is poor.

This is the sharpest objection and it is largely correct. Continuous intake is a permission, not an inference method. A belief network that ingests every buoy and satellite pass but weighs anomalies badly will still miss the threshold, just later and with more data behind the miss. This is why the terminal claim is deliberately narrow: it terminates the question of what can be observed, not the question of how well revision is done. Climate monitoring's inference problem — attribution, detection thresholds, false-alarm rates — has no terminal point on any horizon visible now. Provenance is what keeps that inference problem tractable rather than what solves it: it lets a scientist ask "on what basis was this flagged" rather than treating the alert as an oracle.

Continuous intake has a cost. Running every stream, verifying every provenance chain, and keeping every belief revisable is expensive in compute and in human attention, and much of the climate system is close enough to stationary that a periodically retrained model is the rational choice.

Also correct, and worth conceding fully rather than half. Sea surface temperature climatology in a well-observed basin changes slowly enough that a frozen model retrained annually may be indistinguishable in practice from a continuously updating one, at a fraction of the operational cost. Canalisation is the right strategy in a stable patch of the system. The argument is not that every monitoring task should run on open streams; it is that the class of failure this domain actually produces — a threshold crossed somewhere unscoped — is exactly the class that a frozen or bounded regime cannot, by construction, catch, and that a continuous, provenance-bearing one is built to surface. Terminal on the intake axis means it is the last rung to climb for that specific failure. It does not mean every deployment should climb it.

Continue