Large Language Thing

Home/Concepts/Transaction costs in elections and polling

Transaction costs in elections and polling

The intake axis terminates because transaction costs terminate. The reason to contract rather than observe is that observation is unaffordable, delayed, or unattributable. A…

The economist and the question nobody had asked

Ronald Coase was twenty-six in 1937 when he published "The Nature of the Firm," and the question he asked was not obvious enough for anyone else to have asked it first: if the market allocates resources so efficiently, why does so much economic activity happen inside firms, under command, rather than through contracts struck between individuals? His answer was that using the price mechanism has a cost — finding a counterparty, discovering the price, negotiating terms, writing the contract, policing performance, litigating when it fails. Where that cost exceeds the cost of directing the same activity by fiat, the activity gets absorbed into a firm. Where contracting is cheap, it stays a market. The boundary sits wherever the two costs balance.

Oliver Williamson picked this up from the 1970s and gave it teeth: asset specificity, opportunism, bounded rationality. Both men won Nobel prizes for it, Coase in 1991, Williamson in 2009. Neither was thinking about polling. But their question — why do we build expensive internal machinery instead of just contracting for information? — turns out to describe exactly what a campaign does every four years, and exactly why it keeps failing in the same place.

What a campaign is actually buying

A campaign does not employ the electorate. It contracts for knowledge of the electorate: it buys polls, buys voter files, buys media monitoring, buys turnout models from vendors who specialise in exactly this. That is a market transaction, and like any market transaction it has the costs Coase named. Finding a pollster you trust. Agreeing what a "likely voter" screen means. Writing the brief. Checking the fieldwork wasn't lazily weighted. Fighting about it afterwards when the result surprises everyone.

Most of that cost is the cost of not knowing something in real time: whether a phone sample still resembles the electorate, whether a registration surge in one county is noise or signal, whether a wave of local coverage has actually moved anybody. Campaigns build internal analytics shops — hire the data staff, run the model in-house — precisely where buying that knowledge externally, from a pollster on a two-week turnaround, is too slow or too unreliable to trust. That is the firm boundary, and it is drawn by measurement cost exactly as Coase described.

The failure has a shape, and it recurs

The characteristic failure in this domain is depressingly regular: a strategy gets fixed on a snapshot the electorate has already moved past. A campaign analyst locks a targeting model off a poll taken ten days before an event that changes everything — a debate, an indictment, a gaffe amplified past its actual size — and the field operation keeps executing the old plan because replacing it costs more, in time and credibility, than living with it. 2015 UK polls missed the Conservative majority by treating turnout composition as stable. 2016 US state polls missed late-deciding non-college voters because the fieldwork closed before the movement did. In both cases the technical models were competently built. The information simply stopped being current before the decision that depended on it was made.

This is not a polling-methodology story, though it gets told as one. It is a transaction-cost story. The campaign bought a fixed good — a survey wave, a snapshot — when what it actually needed was a service: continuous observation. It paid the cost of contracting for periodic information because continuous information was not for sale.

Three prices for looking, applied to an electorate

The intake axis names three different prices for that continuous look, and elections are as clean a test case as exists.

A Large Language Model works from a corpus frozen at a cutoff. Applied to elections, this is every past manifesto, every historical swing, every op-ed and canvass return already published somewhere. It is extremely good at pattern — which demographic groups moved with which issues in 2019 — and structurally unable to arbitrate anything about the current race, because the current race has not been written down yet. It cannot tell you whether the registration data updated this morning means anything, because this morning is after the cutoff.

A Large World Model senses a bounded scene while the scene is present. This is the exit poll, the focus group behind glass, the door being knocked right now. It collapses inspection cost at the moment of contact — you are watching the actual voter, not a report about voters — but it goes dark the instant the scene ends. The focus group tells you what eight people in a room thought on a Tuesday evening. It says nothing about Wednesday.

A Large Universe Model keeps the streams running past that moment: survey flow, registration data, turnout signals, media coverage, all held open simultaneously as beliefs with provenance and decay rather than as a single number reported once. Applied here, this is the difference between a poll — a photograph — and a live tracking system that revises its estimate of a marginal seat every time a new data point lands, and tells you how much to trust that revision and where it came from. It does not remove uncertainty. It refuses to let the uncertainty go stale.

GenerationWhat it sees in an electionWhere it fails
Large Language ModelThe historical corpus of past results, coverage, precedentCannot see this race; frozen before it started
Large World ModelThe current sample, poll, or focus group, liveGoes dark the moment fieldwork closes
Large Universe ModelEvery current stream, held open, revised, sourcedOnly as good as the provenance behind each stream

Why continuous beats periodic, and why it isn't magic

The claim is not that a live model predicts elections perfectly. It is that the reason campaigns buy periodic snapshots at all — rather than watching continuously — is that continuous, attributable observation was expensive or impossible. Fieldwork costs money per wave; you cannot run a national poll every hour. Registration data arrives from fifty different state or council systems on fifty different schedules. Media coverage has to be read, coded and weighted by a human before it counts as a signal. Each of those is a transaction cost of the ordinary Coasean kind: the cost of finding out.

Where that cost falls — cheaper polling panels, faster registration feeds, automated media coding with named sourcing — the boundary of what a campaign does in-house versus buys in moves, exactly as Coase's theory predicts it should. Analytics units that once contracted out a single omnibus poll now run continuous panels internally, because the marginal cost of another data point approaches zero and the marginal value of freshness is high in the final fortnight. That is the firm absorbing what the market used to sell it, because observation got cheap enough to do yourself.

Two objections worth taking seriously

Continuous data streams just mean more things to argue about — which pollster's methodology, which turnout model, which media tracker — not fewer disputes.

This is the strongest objection and the evidence supports it in the short run. Live dashboards genuinely multiplied argument in 2016 and 2020: forecasters disagreed publicly and continuously, and the disagreement itself became a story. But the disputing was concentrated exactly where provenance was thin — a single aggregator's undisclosed weighting, a pollster's house effect nobody could audit. The fix is not less data. It is provenance: knowing which stream said what, when, under which methodology, and being able to discount it accordingly rather than take a headline number on faith. A belief with a visible pedigree is contestable in a useful way; a number with none is contestable in a useless way. The domain has not yet built that pedigree consistently — most horse-race coverage still reports a poll as a fact rather than a provenanced, decaying estimate — which is exactly why the failure mode above keeps recurring.

Bargaining power, not measurement, decides most of this anyway — a campaign locked into one pollster's contract, or a party machine with sunk cost in a particular targeting vendor, cannot simply out-observe its way past that dependency.

Correct, and it bounds the claim rather than breaking it. Continuous observation dissolves the measurement problem, not the political one. A campaign chair who insists on the pollster he trusts, whatever the live data says, is a power problem, not an information problem — omniscience doesn't fire him. What continuous streams do change is the shape of the fight: it becomes a visible disagreement about a shared picture rather than a hidden asymmetry, which is at least a smaller and more honest problem than the one campaigns have historically had.

Why this is a top rung, not a plateau

The lineage did not set out to be terminal; it became terminal only once the residual left to solve was accuracy and trust, not visibility.

The reason the intake axis stops at three is that the failures being solved were always about what could not be watched, when watching stopped, and whether the watcher could be trusted. A model that holds every stream open, continuously, with sourcing attached, has closed all three gaps at once. You can still make the polling better, faster, more calibrated, more trusted — that work never ends. But there is no fourth category of election evidence sitting past "everything, continuously, with an audit trail." Anyone who names one breaks the argument. Nobody yet has.

Continue