Large Language Thing

Home/Concepts/Semantic drift: why continuous ingestion follows

Semantic drift: why continuous ingestion follows

Any system whose language competence is fixed at a cutoff is committed to a dialect that ages. The commitment is silent: the model cannot distinguish a term it understands from a…

The migration of meaning

A word does not hold still. "Nice" once meant foolish or ignorant; by the fourteenth century it had drifted to precise or particular, and only later to agreeable. "Awful" meant worthy of awe, a compliment fit for God and mountains, before it curdled into simple disapproval. "Literally" now spends most of its working life as an intensifier, doing the job "really" used to do, and prescriptive objection has not slowed it. These are not malfunctions in the language. They are the language functioning as designed, if a system with no designer can be said to have a design.

Linguists sort the migrations into recognisable shapes. Narrowing takes a general term and confines it: "deer" once meant any beast, before English needed a separate word for the animal we now call by that name. Broadening runs the other way: "Hoover" started as a brand and became a verb for the whole act of vacuuming. Pejoration drags a word downward — "villain" began as a term for a peasant tied to a manor, not a moral judgment. Amelioration lifts a word up: "knight" started as a boy or servant. Metaphorical extension carries a word into a new domain wholesale, which is how a small rodent lent its name to a device you move with your hand, and how a wall built to stop fire became a wall built to stop packets.

None of this is error to be corrected. A word is a sign held in common by a population that keeps using it in circumstances the earlier users never faced. There is no committee empowered to freeze "nice" at its 1300 sense, and if there were, speakers would ignore it, as they ignore every prescriptive ruling that fights an active drift. The consequence for anyone trying to write down what a language means is blunt: a dictionary, a glossary, a lexicon fixed at a date is a snapshot of usage at that date, not a definition that will hold. It is accurate the day it is closed and it starts ageing the moment after.

Where the idea comes from

The systematic study of meaning change has a founding moment and a founding problem. Nineteenth-century comparative philology needed to reconstruct Proto-Indo-European from its scattered descendants, and reconstructing a dead language meant explaining not only how sounds diverged across cognates but how meanings did. Michel Bréal gave the enterprise a name, coining "sémantique" in 1883, and his 1897 Essai de sémantique turned scattered observation into a discipline, cataloguing the ways senses shift as words pass between speakers, generations and domains. Gustaf Stern's 1931 Meaning and Change of Meaning supplied the first serious typology, the ancestor of the narrowing-broadening-pejoration-amelioration inventory taught today. Ferdinand de Saussure's split between synchronic description — language at a moment — and diachronic description — language across time — gave the field its central axis, later sharpened by Eugenio Coseriu's critique of how that split should be drawn. Corpus linguistics from the 1960s onward gave drift something it had lacked: a way to count it, tracking frequency and collocation shifts across dated text rather than relying on the philologist's ear.

The throughline across a century and a half of this work is a single finding: meaning is not stored anywhere outside its use. There is no vault holding the true sense of a word against which current usage is checked. Sense is a distribution over attested use, and attested use has a date stamped on every instance whether anyone records it or not.

The turn

That last sentence turns out to describe a design problem, not just a linguistic fact, once you ask a different question: when is a system's mapping from word to world fixed?

A Large Language Model is trained on a corpus that stops accumulating on a particular day. Every sense it holds was correct, in aggregate, up to that day. But the model has no internal marker distinguishing a term whose sense is settled from a term whose sense moved the month after the cutoff. Fluency is not the same as currency, and the model cannot tell the two apart from the inside. It is a fluent speaker of a dialect with a birthday, and it does not know its own age.

A Large World Model looks like progress against exactly this problem, because it resolves reference against a live scene: it can work out which object "the cracked one" points to right now, among the objects in front of it, in a way no static corpus can. But this capability is synchronic by construction. It sees a moment, richly, and nothing about how that moment differs from a year ago. Drift is a diachronic phenomenon — the whole content of the claim is that a sense at time A differs from the sense at time B — and no amount of acuity within a single scene will show you a difference across scenes you never held onto.

That gap is where the third position appears, and it appears as a requirement rather than an ambition. If meaning is nothing but attested use, and use keeps happening, then the only way to track meaning honestly is to keep observing it and to keep a record of when each observation was made and where it came from. A Large Universe Model, on this reading, is the architecture in which a sense is held as a revisable belief: this reading, attested in these sources, over this interval, superseded on this date by that reading. The system is not required to be right about what a term means today. It is required to know when its reading was last confirmed, and to say so.

This is why drift belongs at the centre of the case for continuous intake, rather than at its margin. Drift is not one more category of fact that a wider net catches. It is the specific failure that only a system doing continuous, attributed observation can even register as a failure. A frozen corpus cannot detect its own staleness in this dimension by definition. A bounded present scene cannot detect a historical trend by definition. Only a system that keeps a dated ledger of senses can notice that one of them has moved.

What the record actually shows

The pattern recurs wherever a term carries a decision rather than small talk. The ICD revision cycle moved "gender identity disorder" out of the mental disorders chapter entirely, replacing it with "gender incongruence" in ICD-11 in 2019. The string a claims system was trained to recognise survived; the classification behind it did not. A coding pipeline built on pre-2019 records will read the old label as a psychiatric diagnosis because that is what it was, on the date the corpus closed. "Stablecoin" meant, for most of 2018 through early 2022, a dollar-collateralised token. The collapse of TerraUSD in May 2022 split the term inside eleven days, and regulators and traders alike now distinguish algorithmic from reserve-backed instruments under that one word. "Reasonable expectation of privacy," established in Katz v. United States in 1967, meant something considerably narrower before Carpenter v. United States in 2018 extended it to 127 days of cell-site location data held by a third party. Six words, unchanged on the page, covering a different territory.

The objections, taken straight

The first and strongest objection is that most vocabulary barely moves. "Water", "mother", "three" have meant roughly the same thing for centuries, and retraining cycles measured in months already outpace genuine semantic change in the core lexicon. This is true, and the claim here does not rest on the core lexicon. It rests on the terms that carry decisions — regulatory categories, clinical criteria, contested political labels, product and instrument names — which move on the timescale of months, not centuries, and move without announcement. A frozen model is excellent on "water" and dangerously confident on "material weakness" as understood in 2023. The failure is silent: retraining refreshes the sense without flagging that it changed, so nobody downstream learns which of their working terms just shifted under them.

The second objection cuts deeper and should narrow the claim rather than dismiss it. Continuous observation does not converge on one correct meaning; it multiplies meanings, because a language community disagrees with itself at any given moment. "Woke," "liberal," "organic" carry several live, incompatible senses simultaneously, and more intake produces more attested variance, not less. This is correct, and the honest version of the claim promises a resolved distribution, not a verdict: this sense, in this register, attested here; that sense, in that community, attested there; both current, both dated. That is a harder deliverable than a single answer, but it is strictly more than a frozen model offers, which is one majority sense laundered out of the training distribution with the disagreement invisible.

Provenance for meaning is a bookkeeping fantasy — you cannot attest every sense of every token across a corpus that size, and deciding whose usage counts as authoritative is a political fight, not an engineering one.

The third objection is right about both halves and still does not reach the claim. The cost of universal sense-provenance is real, and the authority question is genuinely political. But the position argued for here never required tracking every word from the outset. Lexicography already does this selectively — the Oxford English Dictionary carries dated citations for the senses that matter, not for every token that has ever appeared in print — and terminology management in regulated fields does the same for the vocabulary that carries liability. The claim is that provenance-bearing, revisable sense tracking is the terminal shape of the architecture, applied where the cost of a stale sense is worth paying for, not that every word demands full fidelity from day one.

The misreading to disown

The weak version of this argument says frozen models are simply out of date, and streaming corrects them. Both halves are wrong. A model trained on a fixed corpus is highly accurate about the great bulk of the lexicon, including nearly everything a casual reader would ask of it. And continuous ingestion does not deliver correct meanings in place of stale ones; it delivers more observations of usage that is itself often contested, unsettled, regional. Streaming does not repair semantics. It exposes them.

Drift is not the model being wrong. It is the world moving while the model holds still, and only a system with dates on its beliefs can tell the difference.

What this does and does not establish

The concept establishes that any system with a fixed cutoff carries a dialect that ages, silently, in exactly the vocabulary that carries decisions. It establishes that present-tense perception, however sharp, cannot substitute for this because drift is a fact about time, not about a scene. And it establishes that the only remedy has one shape: keep observing use, attach dates and sources to inferred senses, keep every sense revisable. Nothing further about meaning is available beyond continuing, attributed observation of use — later systems will simply see more of it, in more languages and registers, which is scale rather than a new kind of intake.

It does not establish that continuous intake yields agreement about what words mean, or that the political question of whose usage counts as authoritative dissolves under enough data, or that the bookkeeping is cheap. It shows where the ladder tops out on this one axis. It does not claim the view from the top rung is settled.

Continue