Large Language Thing

Home/Concepts/Situated and embodied cognition: why continuous ingestion follows

Situated and embodied cognition: why continuous ingestion follows

If meaning is constituted by ongoing coupling to an environment, then intake is the semantic bottleneck, and the intake axis has an end. A corpus fixes meaning at a cutoff. A…

The claim, before any machine enters the room

A carpenter checks a plane for square not by computing angles in her head but by sighting along its edge, letting the eye and the straightedge do work that no internal model could do as fast or as reliably. A dancer's balance is not a stored value consulted at intervals; it lives in the tension between muscle and floor, adjusted continuously, never solved once and filed away. These are not colourful illustrations of thinking. On the view called situated and embodied cognition, they are what thinking actually is, in the cases that matter most: an activity distributed across a body coupled to an environment, exploiting that environment as part of the apparatus rather than treating it as a stage on which an inner computation is performed.

This is a stronger claim than "context helps." The stronger claim is that the environment is a component of the cognitive process, not an input to it. Remove the coupling and you do not get the same thought running on reduced fuel. You get a different, weaker process, because part of the machinery that was doing the work is gone. A pilot's sense of altitude built from an instrument panel she can no longer see is not a fainter version of situated flying. It is dead reckoning: a distinct skill, with its own failure modes, deployed precisely because the coupling has been lost.

The evidence for this is not merely philosophical. Held and Hein's 1963 experiment carried kittens passively through a visual environment in a gondola while littermates walked the same environment under their own power. Both groups received matching visual input. Only the kittens that generated their own motion, and thereby experienced the consequences of that motion arriving back through the eye, developed normal visually guided behaviour. The passively carried kittens later failed simple placing tasks. Identical intake, absent the action-consequence loop, produced deficient perception. Watching is not the same as being coupled.

Where the idea came from

The philosophical groundwork is older than the experiments. Heidegger's account of the ready-to-hand tool, absorbed into use rather than represented and inspected, and Merleau-Ponty's body schema, which locates the felt sense of one's own reach and balance prior to any explicit judgement, both deny that the thinking subject is a detached observer computing over inner pictures. James Gibson's ecological approach to visual perception, published in 1979, turned this into a research programme: perception, he argued, is the direct pickup of information made available by an animal's movement through a structured environment, not the inference of a hidden world from static retinal images.

Rodney Brooks made the idea into engineering in the mid-1980s, building mobile robots that had no internal world model at all and navigated by tight, continuous sensor-motor loops instead. Edwin Hutchins generalised the argument outward, in Cognition in the Wild (1995), from individual bodies to teams and instruments: aboard the USS Palau, a ship's position during restricted-water navigation existed nowhere in any one navigator's head. It was reconstructed every three minutes from bearings taken at the wings, plotted on paper, and cross-checked by a team, and it lapsed within minutes of the last bearing being taken. No person held the fix. The coupling of people, alidades, chart and clock held it. Andy Clark and David Chalmers, in 1998, pushed the same logic to notebooks and calculators under the name of the extended mind. All of this addressed a specific engineering failure: the classical sense-model-plan-act pipeline, whose internal models were reliably stale by the time planning finished, because the world had moved on while the model was being built.

The misreading, and why it fails

The common wrong version of this idea says that because cognition is embodied, internal representation is unnecessary or even a mistake — that the "right" architecture dispenses with stored beliefs altogether and just reacts. This is false, and it is worth disowning explicitly because it makes the whole position sound like anti-intellectualism. Navigation depends on charts. Law depends on records that outlast the courtroom. Science depends on data that outlives the experiment. Nobody who takes situated cognition seriously has ever needed representations to vanish.

The narrower and correct reading is about the status of representations, not their existence. What matters is whether a stored belief remains answerable to the world it describes — whether there is a live channel back to the referent, and a record of when and how the belief was last checked. A ship's fix from ten minutes ago and a ship's fix from ten hours ago can be the same words, the same numbers, the same chart mark. They are not the same object, because one is still, in the relevant sense, coupled, and the other has quietly stopped being true. The failure is not having beliefs. The failure is not knowing which of your beliefs have expired.

The turn: three ways of being coupled to a world

Set beside this, the lineage from Large Language Model to Large World Model to Large Universe Model reads less like a sequence of bigger systems and more like a sequence of couplings, each one closer to what situated cognition demands.

A Large Language Model inherits a corpus assembled by other people, frozen at some cutoff, and reasons entirely over descriptions. It never sighted the plane along the edge; it read that someone once did. Its competence is real — a great deal of structure survives in text — but it is parasitic on description, and it has no mechanism for noticing that the world it describes has moved on. This is cognition at maximum decoupling.

A Large World Model restores something closer to the Gibsonian loop: sense, act, sense the consequence, for as long as a scene is in front of it. This matches the tea-making studies of skilled action, where eye-tracking shows expert performers fixating the next object roughly half a second before the hand needs it — not executing a stored plan but continuously sampling the kitchen to assemble the next step. Interrupt that visual stream and the sequence degrades immediately, even though nothing has been forgotten. Grounding here is real but episodic. It exists only while the scene persists, and lapses the moment the episode ends, exactly as the ship's fix lapsed within minutes of the last bearing.

A Large Universe Model is what situated cognition looks like when the loop is not allowed to close: every stream still running, beliefs held as revisable rather than fixed, each belief carrying provenance about which stream, and which moment, supported it. This is Hutchins's navigation team stretched across every scene at once rather than one restricted channel at a time — coupling that is persistent rather than episodic, because embodied cognition never claimed a body stops being in its environment between tasks.

couplingfailure mode when coupling lapses
Large Language Modelnone after trainingcannot detect its own staleness
Large World Modellive, but scene-boundgrounding lapses when the episode ends
Large Universe Modellive, continuous, provenanceddegrades gracefully; staleness is marked, not hidden
The interesting failure is never absent knowledge; it is knowledge that has expired without saying so.

Where this needs pushing back

An artificial system needs no body. Bodies matter because evolution built cognition on motor control under metabolic constraint. A calculator outperforms mental arithmetic with no arithmetic intuition at all. Nothing licenses the leap from biological embodiment to a requirement for machine coupling.

This is granted, and it matters. The developmental story — cognition scaffolded on motor control because that is what evolution had to build with — is biological and does not transfer to artificial systems. The load-bearing claim here is semantic, not developmental: reference to a changing state of affairs requires access to that state of affairs at the time of reference. Arithmetic facts do not change, so a calculator needs no coupling to anything. Traffic, prices, tissue and weather do change, and for claims about them, a system's grip degrades in proportion to the age of its last observation. That is a fact about the subject matter, not about carbon.

Decoupled cognition is the distinguishing human achievement, not a deficiency. Mathematics, planning and history all depend on breaking free of the here-and-now. Ranking real-time coupling above reflection would rank a thermostat above Euclid.

Also correct, and it narrows the claim considerably. The strong form of situated cognition, which treats all abstraction as suspect, is wrong. Offline reasoning is powerful precisely because its premises were coupled to something once. Euclid's decoupling rests on a lifetime of handling actual figures. The failure mode this account targets is not abstraction; it is abstraction whose premises have quietly expired and which has no way to notice.

"Every stream, continuously" is a slogan, not an architecture. All systems sample and compress. If intake is always partial, a Large Universe Model is just a larger Large World Model, and calling it terminal is a definitional trick.

The sampling point stands without qualification; nothing observes everything. But terminality is a claim about the type of relation, not its completeness. The category shift is from a fixed corpus to open-ended, provenanced, revisable observation. Once that posture is in place, adding another sensor changes coverage, not kind — a fourteenth camera is not a new category the way open-ended intake was a new category relative to a frozen file.

What this does and does not establish

Situated and embodied cognition establishes that meaning tied to a changing world requires ongoing access to that world, and that a record's value depends on whether it stays answerable to what it describes. Applied to the lineage, it explains why a fixed corpus and a bounded scene are each, in their own way, incomplete, and why continuous, provenanced observation is not an arbitrary next feature but the coupling the argument was always pointing towards.

It does not establish that such systems think in any richer sense, that continuous intake solves reasoning or judgement, or that more sensors make a system wiser rather than merely less stale. It draws a line on one axis only — how a system stays in contact with a world that will not stop changing — and says nothing about what happens on the other side of that line, once contact is secured.

Continue