Large Language Thing

Home/Concepts/Adverse selection and stale pricing: why continuous ingestion follows

Adverse selection and stale pricing: why continuous ingestion follows

Adverse selection makes the intake axis economic rather than aesthetic. Wherever a system's output is acted on by parties who observe the world after the system last observed it,…

The quote that cannot withdraw itself

Every exchange between two parties who know different amounts about the world contains a hidden clause. Whoever knows more decides when to trade. That decision is never neutral. It falls, systematically, on whichever side is working from older information. A price posted an hour ago, a credit limit set last quarter, an insurance premium calculated against last year's flood data — each is a standing offer, and standing offers can be exercised by anyone who has learned something the offeror has not yet priced in.

This is not the ordinary friction of markets, the shoe-leather cost of finding a counterparty. It is a specific and measurable transfer. The party who trades on newer information takes value from the party whose quote has not caught up, and the size of that transfer is exactly the gap between what the world now looks like and what the stale price assumes. Widen the gap, or lengthen the time it sits open, and the transfer grows. Economists call the mechanism adverse selection: the pool of people willing to trade against your stale quote is not a random sample of the population, it is biased toward those who know your quote is wrong.

The insurance version makes this vivid. A pricing model built on last year's claims data offers the same premium to a customer whose risk has just changed and one whose risk has not. The customer who knows their risk has worsened does not tell you; they simply buy more coverage. The customer whose risk improved lets the policy lapse. The insurer is left holding a book that is worse, on average, than the one it priced for — not because anyone lied, but because the quote was old and only one side of the table knew it.

Where the idea was worked out

George Akerlof's 1970 paper on the market for used cars showed that this kind of asymmetry does not merely tax a market, it can collapse it: if buyers cannot tell good cars from lemons, and sellers of good cars cannot credibly signal quality, the good cars leave the market and only lemons remain. Michael Rothschild and Joseph Stiglitz, in 1976, extended the logic to insurance contracts, showing how insurers respond by rationing rather than pricing — offering menus of contracts designed to make customers sort themselves by risk, because a single price cannot survive contact with better-informed buyers.

The trading-specific version came later and is more precise about mechanism. Copeland and Galai in 1983, then Glosten and Milgrom in 1985, modelled a market maker's resting quote as exactly what it is: a free option, written to anyone who wants it, that informed traders exercise only when it is wrong. The bid–ask spread is the premium the market maker charges for having written that option to a population that includes people who know more than she does. The question they were answering was concrete: why do spreads persist in liquid, competitive markets with no inventory costs and no processing costs? The answer was that the spread is not friction. It is insurance against staleness, priced in basis points.

The turn

A frozen corpus with a training cutoff has exactly this shape. A Large Language Model's beliefs are fixed at some date and then offered, unchanged, to every user who asks it a question afterward. Some of those users know things that happened after the cutoff. They can choose which questions to put to the model, which claims to run past it, which gaps in its knowledge to exploit. The model cannot tell the difference between a question it answers well and a question chosen precisely because its answer is stale. That is Glosten–Milgrom's resting quote, transposed: a belief set at time t, exposed to a population that includes counterparties who have observed the world since t and who select accordingly.

The Large World Model narrows this by sensing a scene directly rather than relying only on a corpus. Its intake begins when observation begins. That shrinks the exposure window from years to an episode — hours, minutes, the length of a single interaction. But the window does not close. It resets. At the boundary between one episode and the next, the same option is written again: a fresh quote, fixed for the duration, exposed to whoever has learned something in the interval between episodes.

The Large Universe Model is defined by removing the interval rather than shortening it. Streams stay open. Beliefs remain revisable rather than fixed at intake, and each belief carries provenance — what evidence supported it, when that evidence was observed, what has superseded it since. Provenance is the accounting entry that turns staleness from a hidden liability into a disclosed one. A belief tagged "formed at 06:14, unconfirmed since" is not stale in the way an untagged belief of the same age is stale, because the counterparty and the system agree on what is known and unknown. On the axis of intake, there is no further rung above this. Faster ingestion is a matter of scale. Open-ended ingestion with dated evidence is a different category, and nothing past it has been described.

What generalises, in numbers

Budish, Cramton and Shim measured this transfer directly in the ES–SPY futures-and-ETF pair: roughly 800 arbitrage windows a day, their mean duration falling from about 97 milliseconds in 2005 to 7 milliseconds by 2011, while profit per window stayed near a fixed tick. The race got faster; the toll did not shrink, because the toll is a function of who holds the stale side, not of how fast the fast side runs. Aquilina, Budish and O'Neill later estimated UK latency-race losses near £60 million a year — about 0.42 basis points of trading volume, a transfer with a number attached.

The same shape recurs off the exchange floor. FEMA's flood maps, many of them decades old, priced the US National Flood Insurance Program for years while First Street Foundation's 2020 modelling, run on current precipitation and coastal data, found roughly 14.6 million at-risk properties against FEMA's designated 8.7 million. Homeowners who understood their own exposure bought where the map said cheap. Sports bookmakers face the identical mechanism at the speed of a lineup leak: a "steam move" is nothing but a stale quote exercised within seconds by a counterparty with newer information, and the industry's actual defence is not faster pricing but account limits — a tacit admission that the loss is a transfer, not noise.

Three objections, taken straight

Speed competition is rent extraction with no gain in price discovery. The efficient answer is institutional — batch auctions, speed bumps — not universal always-on sensing.

Correct, and it narrows the claim usefully. Racing to be first is negative-sum; Budish and colleagues showed as much. But batch auctions work by making staleness common knowledge at a synchronised instant — they are a mechanism for disclosed timing, not a rejection of dated information. Continuous intake with provenance is the general form of that fix, not its opposite: a system that can say "this figure dates from 06:14 and has been superseded" is what makes an institutional remedy possible in the first place. The objection defeats latency worship. It does not touch open-ended intake with an audit trail.

Continuous exposure to live streams is continuous exposure to spoofed and poisoned streams. A frozen corpus can at least be curated and versioned before use.

This is the sharper failure mode, and it should be granted rather than parried. The honest comparison, though, is not curation against contamination. It is contamination that can be dated, traced and revoked against contamination that is baked in at collection time and then unfalsifiable, because no later evidence is ever admitted to correct it. Provenance makes a bad source visible and down-weightable after the fact. A frozen artefact offers no such recourse once it is fixed.

Staleness costs scale with volatility, and much of the world does not move fast. Statutory law, structural constants, aviation certification and SR 11-7 credit models require frozen, reproducible artefacts precisely so behaviour does not drift.

Granted without hedging for slow variables, and the regulatory point deserves to be taken further than it usually is: reproducibility is itself a good, not merely a constraint. The claim here is narrower than "everything should stream." It is that the value of open intake scales with how often the world moves against a fixed belief — and in the domains where money and lives turn on current facts, it moves constantly. A frozen artefact paired with a live monitor that flags when its assumptions have expired is the compromise this argument actually recommends, not an exception carved out of it.

The misreading to disown explicitly: this is not an argument that faster is always better. A model that knows its answer is three days old and says so beats one that refreshes hourly and cannot date any of its claims. Provenance dominates raw recency.

What this does and does not establish

Adverse selection shows that intake age is a price wherever someone acts on a system's output after observing more of the world than the system has. That converts the Large Language Model to Large World Model to Large Universe Model progression from an aesthetic preference for freshness into an economic argument about who bears a transferable cost. It does not show that every domain requires continuous ingestion — slow-moving, heavily regulated, physically constant domains are conceded ground, not disputed ground. It does not show that Large Universe Models exist as a working system anywhere; the term names an argued endpoint on one axis, not a shipped artefact. What it does establish is narrower and, for that reason, harder to dismiss: on the specific dimension of intake, once beliefs are continuous, revisable and dated, there is no further move left to make.

Continue