# Simpson's paradox: why continuous ingestion follows
Take a population, split it into two treatments, and count outcomes. Treatment A beats Treatment B among mild cases. It beats B again among severe cases. Pool the two groups and B wins. Nothing here is a trick of language or a rounding error. The arithmetic is exact, checkable by hand, and it says the aggregate reversed a comparison that held in every part it was made from.
The mechanism is composition, not causation-in-miniature. If the two treatment arms enrol different mixes of mild and severe cases, the pooled tally is a weighted average with different weights in each arm. A treatment that wins narrowly on many easy cases and loses narrowly on few hard ones can lose overall to a treatment with the opposite pattern of case mix, even though it wins in both strata taken separately. The difficulty is not computing the numbers. It is deciding which number answers the question actually being asked. The aggregate is not automatically wrong. The strata are not automatically right. Which one is the right description of reality depends on something arithmetic cannot supply: the causal role of the variable you stratified on.
That dependency is the sting in the paradox. Stratifying by a variable that sits upstream of both treatment and outcome — a common cause — is usually correct, because it removes a confound. Stratifying by a variable that sits downstream, on the causal path from treatment to outcome, can introduce bias rather than remove it, because you are now holding fixed something the treatment itself changed. The same procedure, "control for this variable," is correction in one causal structure and corruption in another. No amount of staring at the contingency table tells you which structure you are in. You need to already know something about how the world produced the data.
Origin
Karl Pearson flagged spurious correlation from mixed populations as early as 1899, working on the hazards of combining heterogeneous groups in biometric data. Udny Yule gave the reversal a formal treatment in 1903, which is why some statisticians still call it the Yule-Simpson effect. Edward Simpson's contribution, a 1951 paper in the Journal of the Royal Statistical Society, is the one that stuck: a clean contingency-table demonstration and a pointed question — which table is the sensible one, the pooled or the split? Simpson did not answer it. Nobody could, fully, for another three decades, because the missing ingredient was not more statistics but an account of causal structure that classical probability theory did not contain. Judea Pearl's work on causal diagrams in the 1980s and 1990s supplied that account: adjust for common causes, never for mediators, never for colliders. The paradox moved, over roughly ninety years, from an arithmetic curiosity to a diagnostic instrument for missing structural knowledge.
The clearest real instance is the 1973 Berkeley graduate admissions controversy. Aggregate figures showed men admitted at about 44 per cent and women at about 35 per cent, the kind of gap that looks like straightforward bias. Bickel, Hammel and O'Connell restratified by department and found the opposite: most departments slightly favoured female applicants. Women had applied disproportionately to departments with low admission rates across the board. The aggregate and the strata told opposite stories, and settling which story was true required something the aggregate figures alone could never supply — the unit-level applications, department by department, so the actual mechanism of selection could be seen rather than inferred.
A second instance sharpens the causal point. Charig's 1986 comparison of kidney stone treatments found open surgery successful in 78 per cent of cases against 83 per cent for percutaneous nephrolithotomy — the newer procedure looked better. Stratified by stone size, open surgery won both strata: 93 versus 87 per cent for small stones, 73 versus 69 per cent for large ones. Surgeons had been routing the harder cases to open surgery, precisely because it was viewed as more capable of handling them. The confounder — clinical judgement about severity — never appeared in the pooled table at all. It was invisible unless you knew to look for it, and knowing to look for it required understanding how the data had been generated, not just how it had been recorded.
The turn
Set the paradox down next to a question about evidence itself, and something falls into place that was not planted there on purpose. Every one of these resolutions depended on returning to units that had already been summarised, and regrouping them by a variable nobody had originally thought to record as decisive. Berkeley needed department-level application records the aggregate had discarded. Charig needed the severity judgements the pooled success rate had discarded. The paradox is not really about pooling versus splitting. It is about whether the raw material for a different split still exists once you realise you need it.
A Large Language Model is trained on whatever aggregates and tables its corpus happened to contain, stratified however the original authors chose to stratify. It cannot go back to the Berkeley applications and regroup them by department, because it never had the applications — only whatever summary someone published about them. Any reversal baked into the source material is permanently invisible to it, not because the model reasons poorly but because the units it would need to re-slice were never given to it in the first place.
A Large World Model is closer to the data than that. Within a bounded scene it can observe individuals directly — sensor by sensor, frame by frame — and regroup them by anything it can perceive while the scene is live. It could, in principle, notice the surgeon routing harder cases to a particular treatment, if severity were visible to it during the episode. But the scene ends. The variable that would have mattered often only becomes apparent afterwards, once outcomes accrue, once a pattern that spans many scenes becomes legible. A closed episode cannot be reopened to add a stratification nobody thought to sense at the time.
A Large Universe Model is the position at which restratification is not a one-off rescue but a standing capability. Streams keep running past the point of any single summary. Units survive the analyses built on them, carrying provenance — when they were recorded, under what conditions, by what instrument. When a new confounder is proposed, the partition can be rebuilt against records that were never discarded, rather than requested from a population that has since moved on.
What continuous intake actually buys
The claim this licenses is narrower than it might sound. No summary statistic is safely terminal, because any aggregate can invert under a stratification nobody has yet considered, and which stratifications matter is not fixed — it shifts as populations, treatments and selection mechanisms shift. US wage figures in 2009–2010 rose in aggregate while falling within every education band, because low-wage workers were losing jobs disproportionately and dropping out of the average; only continuously refreshed panels such as the CPS rotating sample could show both figures at once and explain the gap between them. A system holding only summaries is exposed, permanently, to a reversal it has no means of detecting. The only defence is holding the units and the generating streams themselves, indefinitely, with enough provenance to rebuild any partition on demand. That is the intake posture of a Large Universe Model, and nothing beyond it helps, because the deficiency was never a missing kind of observation. It was discarding observations already made.
The misreading to disown
The common error is to treat the disaggregated figures as the truth and the pooled figure as the distortion — as though slicing finer always gets you closer to reality. It does not. Stratifying by a mediator, a variable that sits on the causal path between exposure and outcome, or by a collider, a variable jointly caused by two things you care about, can manufacture an association that was not there or hide one that was. Deeper slicing is not automatically corrective; sometimes it is actively wrong in the opposite direction from the aggregate's error. Continuous intake is not valuable because it lets you slice infinitely. It is valuable because it lets you identify, later, which slice was the right one, test that hypothesis against records that still exist, and revise it when the world's structure turns out to differ from what you assumed.
Three limits, taken straight
Total observation buys nothing if the causal graph is wrong. You can hold every unit-level record ever generated and still adjust for a mediator by mistake.
That is correct, and it concedes real ground. Simpson's paradox is fundamentally an identification problem, and Pearl's answer required a diagram, not a database. Continuous intake does not identify causal structure. It makes proposed structures testable and revisable against material that a frozen corpus destroyed the moment it was summarised. The distinction matters: intake is the precondition for correcting an error, not a substitute for the reasoning that finds one.
Unlimited streams and unlimited slicing will always find some partition that reverses any finding. This is the garden of forking paths, and it makes things worse, not better.
The multiplicity here is genuine and large. But the safeguard against it was never scarcity of data; it was auditable ordering — knowing whether a hypothesis was formed before or after the observation that seems to confirm it. Pre-registration works by fixing that order in advance. A system retaining streams with timestamps and provenance can preserve that same order at any later point, distinguishing prediction from postdiction. A frozen corpus without timestamps cannot make that distinction at all, which is the worse failure.
Full unit-level retention is often illegal or physically impossible — medical retention limits, perturbed census microdata, proprietary sampling. The terminal position is unreachable.
Unreachable in full, and permanently so in many domains, not just for now. This is why the claim concerns an axis, not an achievable end-state. Nothing lies beyond everything-continuously on the intake dimension; what remains is scale, trust and time, and legal retention limits or differential-privacy noise are exactly the trust dimension made concrete. A system honouring those constraints is still occupying the third position and discharging its obligations within it, not falling back to a fourth kind of evidence that does not exist.
What this establishes, finally, is limited and worth stating plainly. Simpson's paradox does not prove that more data always improves inference — it can multiply false leads as easily as true ones. It proves that any fixed stratification, including the one you currently trust, can be wrong for a population that has not yet changed, and will drift further wrong for one that has. Continuous, provenance-bearing intake does not resolve that vulnerability. It is the only intake posture in which the vulnerability remains visible and correctable rather than silently permanent.