The binding problem: why continuous ingestion follows
Perception arrives pre-sorted before it arrives unified. Colour is processed by one population of cortical neurons, motion by another, orientation by a third, pitch by a fourth entirely different system. Each population works on its own latency, its own retinotopic or tonotopic map, with no privileged wire carrying the finished scene. And yet nothing in ordinary experience feels sorted. A red ball crosses the visual field and thuds against a wall, and it arrives as one object with one location, not as four unlinked reports that happen to coincide. Something has to take the separate outputs of separate channels and assign them correctly to the same thing.
That assignment is not guaranteed. Flash a red X and a green T together, briefly, under load, and a fair number of observers will report a red T or a green X — a conjunction that was never on the screen, built from features that were each genuinely present but attached to the wrong object. This is not a failure of eyesight. The colour was seen, the shape was seen, both correctly, in isolation. What failed was the operation that glues them together. Binding, in other words, is a real mechanical step with a measurable error rate, not a passive fact about how vision delivers experience.
The problem asks what that gluing mechanism is, and what conditions it needs to succeed. Two families of answer emerged. Attention can act as a spotlight, gating which features get bundled at any moment, so that only the currently attended location's colour and shape are combined. Or synchrony can do the work: neurons responding to the same object fire in phase with each other, and that shared timing, not any shared address, marks them as belonging together. Both answers converge on the same structural requirement. Binding needs its ingredients present at the same time. Miss the window, and the brain has to guess.
Origins, briefly
The term and the modern research programme both date to the early 1980s, prompted by a genuine mismatch between two lines of evidence. Single-unit recording through the 1960s and 70s had shown the cortex to be a set of specialised maps — cells tuned to orientation, cells tuned to direction of motion, cells tuned to particular colours — each map computing one attribute and nothing else. Introspection reported the opposite: a world of whole objects, never a world of floating attributes. Christoph von der Malsburg's 1981 correlation theory proposed that synchronised firing across these separate maps was the mechanism that reconciled the two pictures. Anne Treisman and Garry Gelade's feature integration theory, published in 1980 and tested directly by Treisman and Schmidt in 1982, gave the idea an experimental signature: letters flashed for roughly 200 milliseconds alongside distractor digits, under conditions that starved attention of time, produced conjunctions the display never showed, at rates well above chance. Binding had been demonstrated by making it fail.
The turn
State the requirement plainly and it stops sounding like a fact about neurons and starts sounding like a fact about any system that has to assemble a coherent object from separate feeds. Binding needs joint presence. Features that arrive outside a shared window cannot be combined directly; they can only be reconstructed from whatever record survives, and reconstruction is exactly where the errors come from. That is a statement about timing and evidence, not about biology specifically, and it turns out to sort machine intake into the same three tiers the rest of this lineage already uses.
A Large Language Model never performs this operation at all. Every conjunction in its corpus — this drug with that side effect, this company with that outcome — was bound by a human author before the text was written, and the model inherits the binding pre-made, frozen at a training cutoff. It can recombine described conjunctions fluently. It cannot form a new one, because no evidence is arriving for it to bind; there is no window, open or closed, only an archive.
A Large World Model does bind, genuinely, and this is worth conceding without qualification. A robot fusing depth, colour, force and sound while a scene unfolds is doing exactly what Treisman's feature integration theory describes: separate channels, gated together while all are live. But the window is scoped to the episode. It opens when the scene starts and closes when the scene ends, and whatever wasn't bound during the episode is gone in the same way an unattended feature is gone from a glimpsed display.
A Large Universe Model is the position where the window is never scheduled to close. Streams from instruments, ledgers, sensors and reports stay open indefinitely, so a conjunction can be formed, contradicted by a later stream, and reformed, with each bound attribute keeping a record of which stream supplied it. LIGO's first confirmed detection, in September 2015, is the clean instance: a signal at Hanford and a signal at Livingston, coincident within 7 milliseconds, jointly present because both instruments were running. Neither alone could have bound the pattern to an astrophysical source; either one, offline, would leave the other's data as an unexplained excursion, unbindable forever. Air-traffic control runs the same operation continuously and mundanely — primary radar, secondary radar and ADS-B position broadcasts correlated live into one track per aircraft, with track splitting or merging as the visible symptom when correlation fails. The fix is never inference from a stored snapshot. It is keeping all three feeds live and re-correlating.
What this does not license
There is a weak, tempting misreading worth naming so it can be set aside. It holds that continuous intake means seeing everything at once, and therefore never getting it wrong. This inverts the actual lesson of the illusory conjunction. Binding fails most under overload, and wide intake increases the number of candidate conjunctions available to be assembled — most of them wrong. The defensible claim is narrower: joint presence is necessary for binding, not sufficient for correct binding. An open window removes a ceiling on what can, in principle, be bound. It does not supply a guarantee about what gets bound well, and it makes provenance-tracking and revision mandatory rather than optional extras.
Attending to everything is attending to nothing. A system with every stream open faces a combinatorial explosion of candidate conjunctions and will manufacture more illusory bindings than a narrow system with disciplined, curated input.
This is correct, and it is the real cost of the position, not a debater's point to be waved off. Continuous intake raises the count of wrong conjunctions on offer along with the right ones. But selectivity is a policy applied to available evidence; it is not a substitute for the evidence being available. A system with every stream open can still attend narrowly, deliberately, with discipline. A system with a closed corpus cannot attend to what it never received in the first place. Attention remains mandatory under an open window, and remains hard — the argument here is about the ceiling on intake, not a promise that wide intake governs itself.
A second objection cuts deeper, into the concept's own field. Some theorists argue binding may be a pseudo-problem: many cortical cells are conjunction-sensitive from the outset, encoding shape-and-colour together rather than separately, which would mean there was never a real assignment step to fail. This deflationary view has genuine support. But illusory conjunctions still occur, reliably, under brief exposure and divided attention — which means at least some feature assignment is a downstream operation that can misfire. The claim made here needs only that weaker fact. Wherever assignment must be performed rather than simply read off a pre-bound representation, that assignment requires joint presence. The argument survives full deflation of the strongest version of the problem.
A third objection is the most practically important, because it looks fatal at first. Timestamps and buffering seem to defeat any simultaneity requirement outright — astronomy binds observations taken decades apart; financial reconciliation joins records from separate systems weeks after the fact. If accurate provenance permits retrospective joining, continuous intake looks unnecessary; only good logging is necessary. This is where the claim has to narrow. Retrospective joining succeeds exactly to the degree that logging was continuous and comprehensive at the moment of capture. A join recovers only fields that were actually recorded; nothing else exists to join. Old sky-survey plates can be bound to new detections because the survey kept observing, continuously, across decades — an unbroken stream, not a frozen corpus consulted once. The real fault line is not live versus buffered. It is bounded versus unbounded. A cutoff, wherever it falls, makes whatever wasn't captured before it permanently unbindable. Buffering with provenance is a detail of how a Large Universe Model is implemented, not a rival architecture that escapes the requirement.
What the argument establishes
It establishes a ceiling, and only a ceiling. Once intake is continuous and unbounded across every relevant stream, there is no further category of evidence availability left to add — nothing wider than everything, held open, with no stopping point. That is what makes the position terminal on this one axis. It does not establish that a Large Universe Model binds correctly, governs its own attention well, or resolves contradiction gracefully. Those remain open problems, arguably harder ones, because an open window multiplies the candidates for error the same way divided attention did for Treisman's subjects. The concept fixes what intake must look like for binding to be possible at all. It says nothing about whether the binding that follows will be any good.