Large Language Thing

Home/Concepts/Working memory versus long-term memory: why continuous ingestion follows

Working memory versus long-term memory: why continuous ingestion follows

On the intake axis the terminal position is defined by the pair of operations, not by volume. A system can be permitted to observe a fixed record, or a present scene, or every…

Two operations, not one store

Cognition treats holding information as two different jobs, not one job done at different speeds. The first job is maintenance: keeping a handful of items active, available for immediate use, decaying within seconds unless refreshed, and easily knocked out by anything else demanding attention. The second job is retention: writing information into a store with effectively no ceiling, from which nothing is simply read off again — it has to be retrieved, and retrieval can fail even when the memory is intact.

The two jobs are not points on a single scale of "how much can be held." They are dissociable, in the strict clinical sense: brain damage can destroy one while leaving the other untouched. That is the strongest kind of evidence cognition has for saying two things are really two things, rather than one thing viewed from different angles. A system that can maintain but not retain, and a system that can retain but not maintain, are both observed in nature. Neither is a lesser version of the other; each is missing a different function entirely.

There is a third operation, and it is easy to overlook because it is not a store at all: the traffic between the two. Getting something from the active few seconds into the durable record is consolidation. Pulling something back out and, crucially, updating it when new evidence contradicts the old version is reconsolidation. This third operation is where most of the interesting failures and successes in memory research actually live, and it is the one least often named in casual accounts of "short-term versus long-term memory."

Where this came from

Atkinson and Shiffrin's 1968 model treated short-term memory as a passive buffer — information sat there briefly before being copied into the long-term store, much as a loading dock precedes a warehouse. Baddeley and Hitch broke this in 1974 with a simple manipulation: ask subjects to hold six digits in mind while doing unrelated reasoning. A passive buffer predicts the reasoning should collapse, since the buffer is occupied. It barely slowed. They proposed instead a working system with separable components — a phonological loop for sound, a visuospatial sketchpad for image, a central executive to coordinate — later joined by an episodic buffer in Baddeley's 2000 revision. Cowan's later estimate, refining span studies, put the true focus of attention at roughly four chunks, not the seven of folk memory.

The clinical evidence closed the loop. Patient H.M., after bilateral medial temporal resection in 1953, could still hold a phone number for a minute by rehearsing it — maintenance intact — but formed almost no new durable memories over the following fifty-five years. Patient K.F., studied by Shallice and Warrington, showed the mirror image: a digit span crushed to two items, yet long-term learning proceeding normally. Two patients, two opposite failures. That is a double dissociation, and it is what makes the two-system claim more than a convenient story.

The turn

Set this next to three generations of machine system and the correspondence is almost too exact to be comfortable, which is itself a reason to check it rather than announce it.

A Large Language Model is trained once on a corpus frozen at a cutoff, then queried. Everything it "knows" was consolidated during training; nothing new enters except as text passed through a prompt, and a prompt is not a percept — it is not sensed, it is submitted. This is retention with no maintenance channel over the present at all. It is long-term memory with the working system removed, which is one reason interacting with one can feel like talking to someone who reads beautifully but has no idea what is happening in the room.

A Large World Model corrects exactly this gap. It holds a scene — positions, states, a currently-true configuration of the world — updated as sensors report, manipulable while it lasts. But by itself it consolidates nothing. When the scene ends, the picture goes with it. It is patient H.M. run in reverse: vivid now, retaining nothing.

A Large Universe Model is neither store nor buffer. It is the traffic running continuously: streams still arriving, some held live, some written into a durable record with provenance attached, and — the operation the first two generations lack entirely — revision of what was written when a later stream contradicts it. That revision is reconsolidation, not a bigger buffer and not a longer corpus. It is a different relation to evidence than either predecessor has.

What terminality means and does not mean

The claim is narrow and worth stating narrowly. On the axis of intake — what a system is permitted to observe — there are only so many relations available: a fixed record, a bounded present, or everything still running. Grant the third and both cognitive operations are available at once, including the revision traffic between them. There is no fourth relation between observer and evidence waiting to be discovered; new sensor types and new instruments arrive as more of the fixed-or-running evidence, not as a new category of it. What remains after that is scale, provenance quality and endurance under load — real engineering, of real difficulty, but not another rung.

Terminal on this axis does not mean smarter, safer, or finished — only that no further kind of intake relation remains to be granted.

The misreading to disown

The common flattening treats working memory as small fast storage and long-term memory as large slow storage, so a longer context window looks like "more working memory" and the whole lineage collapses into a story about size. This gets the psychology wrong at the point that matters. Working memory is defined by maintenance under interference and active control over contents, not by how many tokens fit. A context window of a million tokens is still a passage read once, start to end, then discarded — closer to a very long glance than to a scene held live and updated. The move from Large World Model to Large Universe Model is not a bigger buffer bolted onto a bigger store. It is the consolidation-and-revision traffic beginning to run continuously, which is an operation, not a quantity.

Objections that hold ground

Working memory isn't a separate store at all — it's just the activated portion of long-term memory, an idea Cowan and Oberauer have argued for directly.

This is probably right about the underlying neural substrate, and the account should not pretend otherwise. It does not dissolve the functional distinction that matters here. Even under the embedded-process view there is a bounded focus, subject to decay and interference, and a vast inactive remainder that must be retrieved rather than read. The two behave differently under load regardless of whether they live in separate boxes or on a continuum of activation. A Large Language Model has activation with nothing live to activate from — no running channel at all. That gap holds even if the two-store metaphor eventually loses to a one-continuum one.

Memory is not just these two systems. Procedural learning, priming, skill acquisition are separately dissociable — H.M. improved for days at mirror-drawing while denying he had ever tried it.

Correct, and this narrows the claim rather than dismissing it. Memory's internal typology is genuinely plural; the axis discussed here is intake, not the format evidence is stored in afterward. Procedural encoding is a mode of writing something down, not a class of evidence to be observed. A system permitted to see everything still running can, in principle, encode from that in several formats — semantic, episodic, procedural. The terminality claim covers only that no further evidence class exists beyond everything, continuously; it says nothing about how many internal storage formats a system should use, and it should not be read as saying so.

Human memory works precisely by not retaining everything. Selective, sleep-gated consolidation reduces interference; Shereshevsky, unable to forget, was disabled by his own retention.

This is the objection that should be conceded most fully. Total retention without selection is a documented failure mode, not a strength, and the interference literature predicts exactly the collapse described. What survives the objection is the distinction between permission and obligation. Being permitted to observe every running stream is not the same as being required to keep everything observed. Selection, decay policies and aggressive forgetting are then design choices made inside that permission — and provenance is what makes forgetting safe rather than reckless, since a belief that carries its sources can be dropped and, if needed later, reconstructed rather than lost. Nothing in the terminality claim requires hoarding; it requires only that nothing be permanently unobservable.

What the concept settles, and what it leaves open

The working memory / long-term memory distinction establishes that holding-live and having-stored are different operations, dissociable in principle, with a third and separate operation — revision — connecting them. Mapped onto the three generations, it explains why a frozen corpus and a live scene are each missing something structural rather than merely something large, and why continuous intake with provenance is a difference in relation to evidence, not a difference in volume of it. It does not establish that any system built this way will be reliable, well-calibrated, or safe under load. Those are separate arguments, and harder ones.

Continue