Large Language Thing

Home/Concepts/Granger causality in elections and polling

Granger causality in elections and polling

Granger causality cannot be computed over an unordered archive, and cannot be kept true by a frozen one. The criterion needs three things: order, continuation, and breadth of the…

The loop, run on an election

An analyst opens four dashboards on a Tuesday morning: a tracking poll, a voter-file registration feed, early-vote and mail-ballot return counts by county, and a media-monitoring stream tagging candidate coverage by tone. Each updates on its own clock. The tracking poll refreshes nightly on a rolling three-day sample. Registration data lands weekly from the state. Ballot returns update daily during the early-vote window and hourly near the deadline. Media coverage arrives continuously, scraped and tagged within minutes of publication.

None of these streams is "the truth about the electorate." Each is a lagged, noisy read on a thing that is still moving. Granger's criterion asks a narrow question of exactly this situation: does the past of series A improve the forecast of series B beyond what B's own past already gives you? Does yesterday's coverage tone predict tomorrow's polling movement, once you've already accounted for the poll's own momentum? The question is only well posed if you specify what else you're allowed to look at while asking it — Granger's Ω_t, the information set. That specification is where campaigns succeed or fail long before any test statistic is computed.

What arrives, and in what order

Order is the first thing that has to survive intake, and in campaign data it usually does — this is the domain's good fortune relative to, say, a scraped text corpus. Ballot returns are timestamped by the county clerk. Poll fieldwork dates are reported, if sometimes loosely, alongside topline numbers. Media coverage carries a publication timestamp to the second. The raw materials for a Granger test are, unusually, already ordered when they arrive.

What is not guaranteed is that the streams are held on a common, honest clock once inside the campaign's own systems. A tracking poll reported as "October 14" may reflect interviews conducted October 11–13, averaged, smoothed with a rolling window that blends five nights of data into one number. Feed that blended number into a test against same-day media coverage and you have already destroyed the fine-grained lead-lag structure you're trying to recover. The aggregation happened before the analyst ever saw the number, inside the pollster's own methodology, and it cannot be undone downstream. This is the domain-specific form of the sampling objection: a debate moment or a damaging clip can move opinion within hours, but if the only instrument reading opinion is a three-night rolling average, the true fast interaction gets smeared into something that looks like a slow one, or looks like nothing.

What is held

The analyst's working memory — the belief that gets acted on — is not the raw feed. It is a fitted state: current vote share by region and demographic cell, a turnout model, a set of "movable" segments identified as responsive to particular messages. This state is built by conditioning recent poll movement on registration trends and on coverage tone, over some fixed window, say the trailing 21 days. That 21-day window is Ω_t made concrete. It is a decision about what counts as "recent past" for the purposes of forecasting what happens next, and it is revisited far less often than the data itself updates.

The characteristic failure of the domain lives exactly here. A campaign fixes its message strategy on a snapshot — a memo written from a state built three weeks earlier, when suburban women aged 35–54 were the identified swing segment responsive to a healthcare message. The strategy hardens into ad buys, canvassing scripts, surrogate schedules. Meanwhile the streams kept running: a late-entering ballot measure changed turnout composition, a national news cycle shifted coverage tone toward an unrelated issue, and the swing segment moved on without anyone re-running the test. The analyst is now Granger-testing a population that no longer exists in the form the model assumes. The memo is not wrong about the past. It is stale about the present, and nobody re-dated the belief when the world moved past it.

What triggers revision

A defensible operation does not wait for the next scheduled poll to notice this. It watches for structural breaks across the streams it already holds: a discontinuity in early-vote composition relative to the prior cycle's pace, a sudden divergence between registration trend and poll-implied turnout, a spike in coverage volume uncorrelated with any campaign action. Each of these is a trigger to re-run the conditioning, not a trigger to panic-message.

This is also where the confounder objection bites hardest, and campaigns feel it viscerally every cycle. Coverage tone and poll movement often move together not because one Granger-causes the other but because both are driven by an unmeasured third thing: a debate performance, a court ruling, an economic data release that hits the same week. Widening the conditioning set — adding the economic release calendar, the court docket, the debate schedule as explicit series — does not give the analyst interventional knowledge of what a single ad would do. It does let the analyst rule out the specific spurious finding that "coverage moved the poll" when in fact a Fed announcement moved both. The confounder either enters the information set or it stays invisible; there is no third option, and no amount of watching without including the right series fixes it.

The polling average already accounts for everything relevant. Why maintain separate streams at all?

Because the average is a compression, and compressions are exactly what destroy Granger structure. A single blended number cannot tell you whether registration surges are leading poll movement by four days or trailing it, because the blending already averaged across that lag. Precedence requires keeping the components separate long enough to ask which moved first.

What the operator sees

In practice the analyst does not see a causality verdict. They see a dashboard flag: "early-vote composition in three target counties has diverged from the cycle's baseline pace by more than two standard deviations, coincident with a coverage spike about a ballot measure not on the campaign's own message calendar." That is the output of a Granger-style test run continuously rather than once: an alert that the previously stable lead-lag relationship between registration, coverage, and poll movement has itself shifted, which is different information from "the poll moved."

The cost of this is real and worth naming plainly. Continuous multi-stream testing multiplies comparisons — dozens of counties, several media categories, multiple demographic cells — and naive pairwise Granger tests across all of them will manufacture false positives at a rate that swamps genuine signal. A campaign that reacts to every flagged pair is worse off than one that ignores the dashboard entirely. The discipline required is statistical: correction for multiple comparisons, a higher bar for triggering a strategy change than for triggering a closer look, and an explicit acknowledgment that most flags will be noise. This is the fatal version of the sampling objection, and there is no way around it except rigour in the testing procedure itself — widening the conditioning set responsibly rather than promiscuously.

What it costs, and what is bought

The frozen alternative — a memo built once and defended until election day — is cheaper to produce and easier to brief to a candidate. It is also a bet that nothing structural moves between the memo's fixation date and the vote. In cycles with a late-breaking issue, an unexpected court decision, or a shift in early-vote rules, that bet loses, and it loses in a specific, diagnosable way: the strategy was correct about an electorate that stopped existing.

A polling average tells you where opinion was; a Granger test on live streams tells you whether the thing that used to move it still does.

Held-open intake buys something narrower than omniscience. It buys the ability to notice when the lead-lag structure between registration, returns, coverage, and stated preference has itself changed, and to re-condition the strategy on that change rather than on last month's fit. It does not buy proof that a given ad caused a given swing — that remains an experimental question, answerable only by actually running the ad against a held-out control, which is the interventional rung Granger's criterion never claimed to reach. What continuous, ordered, provenanced intake buys is the difference between a campaign that is wrong about the present and a campaign that at least knows, in near real time, that it has become wrong.

Continue