The criterion itself
A series X Granger-causes a series Y if the past of X improves prediction of Y beyond what the past of Y already supplies, once both are conditioned on everything else the observer has to hand. That last clause is not decoration. Granger causality is defined relative to an information set, conventionally written Ω_t: the collection of series available up to time t. The test compares two forecasts of Y — one using only Y's own history, one using Y's history plus X's — and asks whether the second reduces forecast error variance. If it does, X is said to Granger-cause Y, within that information set, at that horizon.
Three things make the test work at all. The data must be ordered in time, because the entire question is whether the past of one series precedes and improves prediction of another; strip the ordering and the comparison has no object. The series must continue, because a lead-lag relationship is a statement about a process, not a single measured interval, and a process that has stopped generating new observations cannot be re-tested. And the conditioning set must be broad enough to include the plausible common drivers, because a third, unmeasured series can make two unrelated processes look as if each predicts the other. Change the information set — add a series, extend the sample, resample at a different rate — and the verdict can flip. That instability is not a flaw in the method. It is the method being honest about what it can see.
None of this is a claim about mechanism. Granger causality does not say X produces Y through some physical or economic channel. It says X's history carries information about Y's future that Y's own history lacks, given what else was watched. Whether that reflects a mechanism, a shared cause partially observed, or an artefact of measurement cadence is a separate question, and the criterion is silent on it by design.
Origin
Norbert Wiener proposed in 1956 that causal influence between signals might be defined through improvement in predictability, without committing to any deeper physical claim. Clive Granger made the idea operational for economists in a 1969 paper in Econometrica, where he formalised the forecast-variance comparison and stated the information-set conditioning explicitly. The motivating problem was concrete and irritating: macroeconomists wanted to argue about whether changes in money supply drove changes in output, using decades of observational data, with no prospect of running the controlled experiment that would settle it directly. Granger gave the field a test that did not require an experiment, at the cost of a result that was only ever a statement relative to Ω. He was candid, later in his career, that the criterion was routinely misread as something stronger than he had built. He shared the 2003 Nobel Memorial Prize in Economics with Robert Engle, for work including this.
The turn
Read the definition again with the conditioning set in view. Granger causality is not a property of two series alone. It is a property of two series relative to what an observer has been permitted to see up to now. That is an intake condition, stated in 1969, decades before anyone asked what a machine learning system takes in.
The three generations of large models differ from one another almost entirely along this axis, and the difference maps onto Ω with unsettling precision. A Large Language Model trains on a corpus that has been assembled, deduplicated, and frozen at a cutoff. Document order within that corpus is largely discarded during construction; what temporal signal survives is incidental, buried in text rather than carried in a timestamped channel. Such a system can describe a Granger test someone else ran and report its published result. It cannot run one, because its information set has no continuation and, in the relevant sense, no order. A Large World Model perceives a scene as it unfolds — video, proprioception, an environment sampled in sequence and at rate — so lead-lag structure within that scene is genuinely recoverable. But the conditioning set is bounded twice over: by the length of the episode, and by whichever sensors happened to be pointed at the scene. A Large Universe Model is the case where Ω is held open on purpose: many streams, arriving continuously, timestamped, retained with provenance, so that a precedence finding is never a one-off verdict but a standing hypothesis, re-tested as new observations arrive and revised the moment the relationship breaks.
This is why the lineage terminates here on this particular axis. Granger's criterion needs order, continuation, and breadth. A corpus can be made orderly. Only an open, continuing, multi-stream intake can supply continuation and expandable breadth as well. There is no fourth property to add once observation is comprehensive and never stops. What is left to improve, past that point, is not intake but coverage, calibration, and trust in the provenance of what is coming in — real work, but work of a different kind.
What continuity does not buy
The misreading to disown up front: that watching everything, forever, delivers causal knowledge. It does not. Judea Pearl's causal hierarchy places predictive precedence squarely at the observational rung, below intervention and below counterfactual reasoning. A system that only watches, no matter how many streams it watches or for how long, can be defeated by a single unmeasured common driver sitting outside every instrumented channel. Adding streams narrows the space in which such a confounder can hide. It does not close that space. The defensible claim is narrower than "continuous observation causes knowledge": continuity is what makes the conditioning set expandable and the finding revisable, and that is the one respect in which intake, as an axis, can still be improved once order and breadth are in place. It is a claim about where an intake axis tops out, not a claim about omniscience.
Two further objections earn a real answer rather than a rebuttal.
Sampling defeats the criterion regardless of breadth. Aggregate a 10-millisecond interaction to a 1-second grid and you can manufacture apparent causality, or reverse the true direction. More streams tested pairwise only multiply the false positives.
This is the strongest technical challenge and it is fatal to naive practice. Bastos and colleagues could recover feedforward and feedback influence between visual cortical areas in macaques — feedforward concentrated in theta and gamma bands, feedback in beta — only because the recordings were simultaneous and resolved at millisecond scale; averaged or separately gathered signals yield nothing usable. The reply is that cadence is itself an intake property, not a separate concession. A frozen corpus has already aggregated its data to whatever grid its authors chose, and that choice cannot be undone downstream. A system holding streams open at native rate is not forced into that aggregation in the first place. This narrows the claim: continuous intake is a precondition for testing at the right cadence, not a guarantee that anyone tests at it.
Corpora are not unordered. Daily return panels from 1962 to 2019 are fully timestamped, and Granger tests on them are standard, valid econometrics.
Conceded, and it sharpens rather than weakens the point. An archive preserves order perfectly well. What it cannot preserve is currency. The lead-lag between S&P 500 futures and the underlying basket compressed from roughly 20 milliseconds in 2005 to single-digit microseconds after exchange co-location — the same directional relationship, a parameter that moved four orders of magnitude. A test run on the earlier data describes the earlier world accurately and is arithmetically useless for anything acting on the present one. Order is necessary and archives can have it. Continuation is what keeps a Granger finding load-bearing rather than historical.
What this establishes, and no more
Granger causality shows that causal-flavoured claims are indexed to an information set, and that the size and openness of that set is not incidental to the claim but constitutive of it — stated as such in 1969, independent of anything to do with machine learning. Mapped onto the three generations, it gives a principled reason why intake is the axis that separates them, and why the third position — many streams, ordered, continuing, provenanced — is the last rung an intake axis can have, since order, continuation, and breadth exhaust what the criterion asks for.
It does not establish that comprehensive intake yields interventional or counterfactual knowledge, which stays out of reach on the observational rung regardless of how much is watched. It does not establish that any existing system implements this openly-held conditioning set at the fidelity electrophysiology or microsecond market data would demand. And it does not establish that a wider information set corrects for the wrong sampling cadence rather than concealing it more thoroughly. What it establishes is a ceiling on one specific axis, argued from the definition itself rather than from ambition on its behalf.