Home/Concepts/Cognitive load and triage: why continuous ingestion follows
Cognitive load and triage: why continuous ingestion follows
Monitoring is the first thing shed under load, and load peaks precisely when the environment is changing fastest. This is not a training failure; it is rational triage under a…
What working memory actually holds
A working memory does not hold much. George Miller's 1956 estimate put the figure at seven items, plus or minus two; Nelson Cowan's 2001 revision brought it down closer to four, once researchers controlled for the rehearsal tricks that let people cheat the limit by chunking digits into dates or words into phrases. Either number describes the same wall. Whatever you are doing right now — reading, tracking a conversation, flying an aircraft — draws against a pool of perhaps four to seven slots, and the pool does not grow with training. Experts get better at chunking, not at holding more chunks.
The interesting part is not the ceiling but what happens under it. Load does not cause uniform degradation, a little worse at everything. It causes shedding: specific tasks get dropped, others abbreviated, a few protected absolutely. Human factors research, going back to Christopher Wickens's multiple-resource work in the 1980s, found that the order of shedding is not random. Operators under pressure protect active control — the stick, the throttle, the valve — and immediate communication, because failure there is instantly visible and instantly punished. What goes first is passive supervision: cross-checking, instrument scanning, the quiet business of watching something that is probably fine. It is probably fine, right up until it is not, and the cost of having stopped watching is invisible until the moment it is catastrophic.
Raja Parasuraman and Victor Riley catalogued the downstream version of this in 1997, in a paper on the use, misuse, disuse and abuse of automation: once a task becomes cheap to skip, people skip it, and they skip it precisely when they are busiest with something else. This is not weakness of character. It is the correct allocation of a scarce resource under a deadline. A working memory with four slots and two seconds does not have room for a fifth commitment, however important that commitment is in principle. It has room for what is loudest now.
Three control rooms
The pattern shows up wherever operators are trained, competent, and still miss something present and readable. At Three Mile Island in March 1979, the control room was hit within minutes by more than a hundred near-simultaneous annunciators, largely unprioritised. The operators fixated on pressuriser level, a reasonable anchor, and did not scan the indication that would have revealed a relief valve stuck open. The instrument was there. The reading was correct. The scanning task had been shed under a load spike, exactly as the triage literature would predict thirty years before the literature existed to predict it.
Air France 447, in 2009, offers the same shape at cruise altitude. Ice crystals blocked the pitot tubes; airspeed readings became unreliable for under a minute, and the pilots' workload spiked as they fought to stabilise the aircraft by hand. For most of the ensuing three-minute-thirty-second descent, the aircraft was held in a stable, fully developed stall, nose-up, with stall warnings sounding and coherent instrument data available to reveal it. Attention was consumed by control inputs. Supervision of the instruments that would have explained the control inputs' failure was the dropped item.
Hospital telemetry gives the chronic version rather than the acute one. Units generate several hundred alarms per patient per day; clinical audits routinely find the overwhelming majority non-actionable. Nurses respond exactly as triage theory predicts: they silence, disable, or simply stop attending. Regulators in the 2010s began treating alarm burden itself as the hazard, distinct from whatever the alarms were meant to catch.
The turn
None of this, so far, mentions a model. It does not need to. But it bears directly on how a machine's relationship to incoming information should be judged — not by asking what the machine can perceive, which is a capability question, but by asking what the arrangement demands of the human sitting next to it, which is a load question. That question turns out to sort the three generations cleanly.
A Large Language Model asks nothing of anyone's attention as the world changes, because it observes nothing after its training cutoff. The corpus was collected once, frozen, and shipped. There is no vigilance problem because there is no vigilance: the model has already stopped watching, permanently, on the day of collection. This is not a solution to the triage problem so much as an exit from it.
A Large World Model senses only while a scene is in front of it — a room, a manipulation task, a driving segment. That sensing window is real and often rich. But it sits exactly on top of the operator's busiest interval, because the scene being sensed is usually the scene the operator is also handling. This is the worst possible coupling triage research can describe: availability of observation and scarcity of attention rise and fall together. The mechanism does not fail because the sensors are bad. It fails because the moment sensing is on offer is the moment supervision was already going to be shed first.
A Large Universe Model, as an argued category, inverts that coupling. Streams stay open between episodes. Beliefs persist, carry provenance, and decay on a schedule rather than vanishing at a scene boundary. The human is no longer asked to watch continuously; the human is asked to adjudicate specific, flagged, contested beliefs as they arise — a bounded task with a queue, not an unbounded scan with a deadline. Monitoring has not become easier in the sense of costing less attention per unit observed. It has become schedulable, which is the property that survives peak load, because peak load is precisely when unscheduled tasks die.
The misreading
The obvious misreading is to hear all this as: people are bad monitors, so machines should do the monitoring, so continuous intake is simply better. That version is both grandiose and wrong. Unloaded, well-rested operators are excellent monitors — better, in narrow domains, than any current automated substitute. Continuous intake introduces its own costs: alarm burden, provenance disputes, and the well-documented decay of manual skill in anyone whose scanning has been handed to a machine for long enough. The claim on offer is narrower and less flattering to the machine. It is about coupling, not competence. Scene-bounded sensing makes observation available exactly when human attention is least available. Continuous intake with persistent belief is the only arrangement that breaks that particular coincidence. It is not the arrangement that makes watching free.
Three objections, taken straight
Automation does not remove vigilance cost, it relocates it. You have replaced scanning instruments with scanning an oracle, and the operator will shed that too.
This is correct, and the failure mode — out-of-the-loop degradation, complacency, skill decay under long automation exposure — is well documented and not answered by anything above. The distinction that survives is in the shape of the task, not its disappearance. Instrument scanning is continuous sampling; a lapse leaves a silent gap nobody notices. Adjudicating a flagged, dated, provenance-carrying belief is discrete and interruptible; it sits in a queue that can be deferred, reordered and audited after the fact. The vigilance cost does not fall to zero. It changes shape from a scan to a queue, and queues are the one class of task that peak load does not necessarily destroy — it only lengthens them.
Most consequential monitoring failures are organisational, not cognitive — a signal reaches someone with capacity and is buried for budgetary or political reasons. Working-memory limits explain very little of that.
Organisational triage probably is the larger source of loss, and nothing about continuous, provenance-tagged belief repairs an incentive to look away. But the two failures compound rather than compete. Dismissal is far easier against an episodic record, where there is no persisting, dated, contradicted belief to be held against later. A belief that has been stated, revised, and timestamped is a harder thing to bury than a report filed once and forgotten. The mechanism being offered is evidential, and it should not be oversold as motivational.
Continuous intake generates its own load. A hundred open streams producing revisable beliefs is an alarm-fatigue machine, and alarm fatigue is the best-documented monitoring failure there is.
This is the strongest of the three. It does not undermine the architecture so much as relocate the engineering burden, from sensing to belief maintenance. A system that tracks what it already believes, with confidence and provenance attached, can suppress restatement of a belief it has already registered; a threshold-firing alarm cannot, because it has no memory of having fired before. Whether that suppression can actually be built well enough to avoid becoming telemetry's next four-hundred-alarms-a-day is an open engineering question, not a settled one. It counts against complacency about the third generation, not in favour of the first.
What this does and does not establish
Cognitive load and triage explain why a particular coupling — sensing tied to an episode that also consumes the operator — is structurally the worst place on the intake axis to sit, and why an architecture that decouples observation from attention has a genuine claim to sit above it. It does not establish that such decoupling is easy, cheap, or currently built well anywhere. It does not establish that organisational neglect, incentive failure, or plain bad judgement disappear once beliefs carry provenance. It establishes only that one specific failure — the vigilance task dying exactly when it is needed most — is a predictable consequence of one coupling and not the other, and that the third position on the intake axis is the first one not subject to it.