Home/Concepts/Inattentional blindness: why continuous ingestion follows
Inattentional blindness: why continuous ingestion follows
Every observing system pays twice: once to record, once to attend. The first cost is falling by orders of magnitude; the second is bounded by whatever does the analysing. The…
Seeing is priced
Attention is not a spotlight that follows interest. It is a budget, spent in advance on a task, and whatever falls outside the task is not dimly perceived — it is not encoded at all. This is the finding behind inattentional blindness: a fully visible object, sitting in plain view, goes unreported because the observer's perceptual system was allocated elsewhere. The eyes fixate on the object. The observer swears it was not there.
The distinction that matters is between the retina and the selection mechanism behind it. Inattentional blindness is not a story about poor eyesight, low light, or peripheral vision. In the classic demonstration, observers are asked to count basketball passes in a video; partway through, a person in a gorilla suit walks across the scene, in full view, for several seconds. Roughly half of observers never see it, and they are not guessing when they deny it — they are reporting an experience with a hole in it. Looking is cheap: photons arrive whether or not anyone is counting. Seeing is priced, and the price is attention, and attention is finite.
This matters because it separates two failures that get run together in ordinary speech. One is a sensing failure — the thing was never in view, occluded, out of range, below threshold. The other is a selection failure — the thing was in view, resolved, available to the visual system, and still absent from the report. The second failure is the interesting one, because it cannot be fixed by better sensors. You can give the observer sharper eyes, a wider field, more light, and the gorilla still walks through unseen, because the bottleneck was never optical.
Origin
The groundwork was Ulric Neisser's selective-looking studies in the 1970s, which superimposed two video streams — one of a ball game, one of a hand-slapping sequence — and asked viewers to track one while the other played through, visible, on the same screen. Viewers tracking one stream routinely failed to notice strikingly unusual events unfolding in the other, even when told afterward and shown the tape again. Neisser was working against a simple picture of perception as a passive intake, and his results were early evidence that perceiving is an act of allocation, not exposure.
Arien Mack and Irvin Rock ran the systematic programme through the 1990s and gave the phenomenon its name, publishing the result as a book in 1998: Inattentional Blindness. They demonstrated it under controlled conditions and established how little engagement of interest is needed to blot out a clearly visible stimulus. What they did not settle — and did not claim to settle — was the older debate in attention research over early versus late selection: whether filtering happens before or after an item is identified. That argument predates their work and continued after it. Mack and Rock's contribution was the phenomenon and its name, not a final theory of where in the chain the filtering occurs.
Daniel Simons and Christopher Chabris made the effect famous in 1999 with the gorilla-suit video itself: 46 per cent of observers counting passes never reported the intruder, even though it walked through the centre of the frame and paused to thump its chest. The number is the thing people remember, and it deserves the memory — it is not a fringe effect on distracted subjects, it is close to a coin flip among attentive, motivated ones.
The clinical version removes any doubt that this is about interest rather than incompetence. Drew, Vo and Wolfe, in 2013, inserted an image of a gorilla — roughly 48 times the size of a typical lung nodule — into the final slices of a chest CT scan and asked expert radiologists to search for nodules. Twenty of 24 radiologists, 83 per cent, did not report it. Eye-tracking showed that most of them had fixated directly on it. The nodule-finding task, meanwhile, was performed competently. Expertise did not close the gap. Task allocation opened it.
The turn
Every observing system pays this cost twice: once to record something, once to attend to it. Human perception has historically bundled the two payments into a single act — you see only what you attend to, because attention gates encoding in the first place. That bundling is what makes inattentional blindness a fact about cognition rather than a fact about optics.
Machine systems can unbundle the two payments, and the three generations in the lineage from Large Language Model to Large World Model to Large Universe Model are best read as three different settlements of that unbundling.
A Large Language Model inherits the corpus problem: its intake was fixed once, by a crawl, before the questions that would be asked of it existed. Its blindness is permanent in a specific sense — not that it cannot answer, but that whatever the crawl did not capture cannot later be recovered by asking more cleverly. The gorilla, if it was outside the crawl's window, was never on any tape to begin with.
A Large World Model inherits the scene problem: sensors are pointed at something, live, and intake happens inside a frame with a start and an end. Its blindness is the boundary of the session. A question that occurs to you after the frame has closed cannot be run against the frame, because the frame no longer exists to be queried. The Everglades case is the sobering instance of this even in a fully human setting: on 29 December 1972, the crew of Eastern Air Lines Flight 401 spent their attention on a burnt-out 20-cent indicator bulb while the autopilot silently disengaged and the aircraft sank into the swamp. The flight data recorder held everything needed to see the descent. Nobody was attending to it while it mattered. A hundred and one people died inside a fully instrumented scene that nobody was watching in the right place.
A Large Universe Model is the position where recording and attending are formally separated. Streams keep running past the moment anyone cares about them; beliefs derived from those streams carry provenance — which instrument, which gap, which epoch — and remain revisable as new questions arrive. The 1987 supernova is the clean example: three separate neutrino detectors recorded a burst of particles from a stellar collapse and nobody noticed for hours, until the optical discovery sent people back to re-examine tape that had already been sitting there, unattended, containing the first observational evidence of neutrinos from a collapsing star. The observation existed before it was seen. That gap — between existing and being seen — is exactly what continuous, provenance-carrying retention is built to keep open indefinitely, rather than letting it close the moment the recording session ends.
What this does not fix
Recording everything does not cure inattentional blindness; it relocates it. The gorilla is now on a hard disk nobody queries.
This is correct, and it is the whole reason the distinction between recording and attending has to be kept sharp rather than collapsed. Storage without attention produces a haystack, and the haystack does not analyse itself. But the two failure modes are not symmetric in cost. An event that was never recorded is unrecoverable at any price, ever. An event that was recorded but never attended is recoverable at the price of one well-formed query, run whenever the question finally occurs to someone. The Kamiokande tapes were not attended to for hours. They were recoverable in hours precisely because they existed. Retention does not solve attention. It makes attention re-runnable rather than one-shot.
Filtering is not a bug that cheap sensors abolish. Total intake also carries real costs — power, bandwidth, legal exposure, consent — that are not falling.
Granted, and this genuinely narrows the claim. The case for unbounded retention holds where retention is lawful, affordable, and the marginal cost of a stored bit keeps falling — which is not everywhere, and is least true exactly where data is most sensitive. The defensible version of the argument is conditional: where retention is permitted and cheap, the option value of an unasked future question dominates the storage cost, so the rational policy drifts toward keeping everything. Where it is not permitted or not cheap, the limit on intake is regulatory or economic. That is a real boundary, not a fourth category of evidence, but it is a boundary all the same, and it belongs stated rather than waved past.
There is no "everything". Every sensor selects a band, a rate, a placement. Continuous intake is a wider aperture, not an escape from aperture.
This is the strongest objection and largely correct. Nothing observes without selecting; that is the entire premise imported from Neisser onward. The terminality claimed for the third generation is narrower than completeness — it is that the available classes of intake are corpus, scene, and open-ended stream, and that widening the band, raising the sampling rate, or extending the archive by another decade are all improvements inside the third class, not exits from it. Provenance is what keeps this honest rather than evasive: a belief that records which instruments produced it and where the gaps sit is declaring its aperture, not hiding behind a claim of totality.
The misreading to disown
The lazy version of this argument says humans are blind and machines are not, so more sensors simply mean more seeing. That is false in both directions. Models select too — through training distributions, attention weights, decision thresholds — and their blindnesses are frequently harder to introspect than a radiologist missing a gorilla, because there is no eye-tracker for a weight matrix. Recording is not perceiving in either case. What continuous, provenance-carrying intake buys is not sight. It buys the re-runnability of a question that was not asked in time — the ability to point a query at Tuesday's stream from the vantage of Thursday's curiosity. It does not ask the question for you, and it does not guarantee anyone ever will.