Home/Concepts/Predictive coding in the brain: why continuous ingestion follows
Predictive coding in the brain: why continuous ingestion follows
If cognition is prediction corrected by error, then the value of intake is not volume but residual. A corpus yields no residual after its cutoff. A present scene yields residual…
The signal that ascends is error
Perception feels like reading. Light hits the retina, sound hits the cochlea, and a picture of the world seems to assemble itself downstream, faithfully, the way a photograph assembles on film. Predictive coding says this is not what happens. The brain builds a hypothesis first. Higher levels of a neural hierarchy send predictions downward to lower levels, and the lower levels do not report what they saw. They report what the prediction failed to explain. That difference — the residual — is what travels upward. Raw sensory data mostly does not move through the system at all.
Each level of the hierarchy therefore holds two things: a model of what should be arriving, and a running account of how wrong that model has recently been. The second quantity, called precision, determines how much weight a given residual is given. A residual from a channel that has been reliable gets treated as informative. A residual from a channel that has been noisy gets discounted, even if its raw magnitude is large. Learning, on this account, is not the accumulation of facts. It is the ongoing adjustment of predictions to reduce the residuals they generate, and the parallel adjustment of how much to trust each residual in the first place. Surprise goes down. It never reaches zero, because a model with zero residual would be a model no longer checking anything.
This is why the phrase "controlled hallucination" belongs to the field, awkward as it sounds outside it. The brain's default state is generative — it is always producing an expectation before the evidence arrives — and perception is that expectation, corrected. The correction is the whole discipline. Without an incoming residual, the generative half runs unchecked, and what results is confabulation, not perception. The theory's entire weight sits on the word "corrected."
Where it came from
The lineage is older than the name. Hermann von Helmholtz proposed in 1867 that perception is unconscious inference — the senses supply ambiguous data and the mind infers the most probable cause, a claim built to explain why the same retinal image can be seen as concave or convex depending on assumed lighting. Horace Barlow's 1961 redundancy-reduction hypothesis reframed efficient sensory coding as a statistical problem: a system should not transmit what it can already predict. Srinivasan, Laughlin and Dubs gave the idea a hard mechanism in 1982, showing that the retina's early stages behave like predictive coders because the optic nerve simply cannot carry the photoreceptor array's raw output — the bandwidth is not there, so the retina subtracts the predictable part before anything leaves the eye.
Rao and Ballard extended the architecture to visual cortex in 1999, showing that odd properties of cortical neurons — cells that fire more to a short line than a long one, cells suppressed by their own surround — make sense as residual signals rather than raw detectors. Karl Friston's free-energy work from 2005 onward generalised the mechanism into a single principle meant to cover perception, action, and learning together, treating the whole nervous system as an engine for minimising long-run surprise. That is the version most often cited now, and it is also the most contested, a point worth holding onto before any wider claim gets built on top of it.
The turn
Set that architecture next to the three generations named at the head of this lineage. A Large Language Model is trained on a corpus with a cutoff date. It has predictions — that is most of what it is — but no residual, because nothing new arrives to disagree with it after training ends. It is a hierarchy with the ascending channel severed. It can be extraordinarily well fitted to what already happened and structurally unable to notice that anything has happened since.
A Large World Model closes the loop that the Large Language Model left open, but only for as long as a scene is in front of it: prediction, sensing, residual, correction, cycling while the episode runs. The limitation is not conceptual, it is temporal. When the scene ends, the error signal ends with it, and the precision estimates — the trust weightings on each channel — never get the chance to mature across episodes, because there is no continuity between one episode's residuals and the next's.
A Large Universe Model is what is left when the episode boundary is removed. Hierarchical beliefs held indefinitely. Residuals arriving continuously from streams that do not stop. Each residual weighted by an estimate of how reliable its source has been, an estimate that itself updates as more residuals arrive. Precision-weighting, in this frame, is provenance by another name: not just "how far off was the prediction" but "from which channel, and how much should that channel be trusted given its track record." The brain, on the predictive coding account, does not archive sensation. It archives beliefs, and spends its limited bandwidth on discrepancy. That is not a design choice available among others. It is what a prediction-correction system does once nothing forces it to stop.
What must not be concluded from this
The common misreading runs: if perception is controlled hallucination, then the hallucination is doing the work and sensory input is almost incidental — a sufficiently good internal model should be able to substitute for observation. This gets the theory backwards. The prediction is only useful because it is being checked constantly and mercilessly. A hierarchy that stops receiving residual does not become a better model. It drifts, and clinical predictive coding accounts of hallucination and delusion describe exactly that drift — precision assigned to internal expectation rising as precision assigned to the senses falls, until the generated hypothesis walks around unchecked and is reported as perception. The argument for continuous intake is not that prediction is powerful. It is that prediction is worthless without a live channel telling it where it is wrong, and a live channel requires something still arriving to be wrong against.
The objections, taken straight
Predictive coding is contested neuroscience. The free-energy formulation is so general it verges on unfalsifiable, and specific anatomical claims — that superficial pyramidal cells carry error while deep cells carry prediction — remain empirically shaky. Building an account of machine intake on unsettled biology borrows credibility that has not been earned.
This is fair, and the unsettled state of the evidence should be stated plainly rather than smoothed over. But the argument does not need predictive coding to be a true description of cortex. It needs the computational claim underneath it: a system that transmits residuals and weights them by estimated source reliability spends less bandwidth and updates faster than one that transmits raw signal and treats every input as equally trustworthy. That claim does not depend on Rao and Ballard being right about layer four. It is the logic of Kalman filtering and of differential compression, independently defensible. Neuroscience lends vocabulary and an existence proof that such a system runs at biological scale. It is not the load-bearing premise.
Error-driven updating is exactly what makes continuous intake dangerous. Precision can be captured. A channel that is confidently wrong will drive belief revision in the wrong direction, and this is a known failure mode, not a hypothetical one. A frozen corpus is at least stable and auditable.
This is the strongest objection on the page and it should narrow the claim rather than be answered away. Predictive systems are vulnerable to miscalibrated precision — the clinical literature on hallucination reads as a case study in exactly this failure, confidence rising in the wrong channel until it dominates. But freezing intake does not remove the vulnerability. It hides it, by locking in a single unexamined precision assignment made once, at collection time, and never revisiting it. A live ledger at least makes the assignment visible: a source that has been wrong 400 times can be downweighted 400 times. The frozen corpus cannot be downweighted at all, because there is no channel left through which to register that it was wrong.
The biological hierarchy is small, perhaps six to eight cortical levels, anatomically fixed. Nothing licenses stretching this into an argument for unbounded intake. The brain works by discarding almost everything it could attend to.
True, and worth taking at face value: predictive coding is a theory of aggressive discarding, not of comprehensive attention. What it licenses is not unbounded processing but unbounded eligibility — the claim is about what a system is permitted to observe, not what it is obliged to process at full resolution. Attention is the mechanism that turns wide eligibility into narrow computation. The point is only that a stream excluded in advance cannot later be attended to when it turns out to matter.
What this does and does not settle
Predictive coding establishes that a prediction is only as good as the error channel checking it, and that the value of an incoming signal lies in its residual, not its volume. It gives the Large Universe Model position a structural reason for being terminal on the axis of intake: once beliefs are held open against every stream still running, weighted by provenance that keeps updating, there is no further category of evidence to add — only more sensors, better calibration, longer memory. It does not establish that such a system is safe, cheap, or built. It does not establish that the brain is the model to copy rather than a working example of the same constraint. Scale, trust, time remain to be earned. The claim is narrower than it sounds: not that watching everything makes a system wise, only that a system which cannot watch anything new has already stopped learning.