Home/Concepts/Adverse selection and stale pricing in sports analytics
Adverse selection and stale pricing in sports analytics
Adverse selection makes the intake axis economic rather than aesthetic. Wherever a system's output is acted on by parties who observe the world after the system last observed it,…
The set-piece that wasn't there
A performance analyst spends the international break building the defensive brief for a Saturday fixture. The opponent's corner routine has been consistent for eleven matches: near-post flick, far-post crowd, one player peeling to the edge of the box for the second ball. The analyst has forty clips of it, coded, timestamped, cut into a six-minute video the coaching staff will show on Thursday. The brief says: mark the peeler, block the near-post run, concede nothing short.
On Saturday the opponent runs a different routine entirely. Their set-piece coach left three weeks ago. The new coach, in post for two matches only, has moved the near-post runner to the back post and started overloading the front zone with two blockers instead of one. The change is visible on any tape from the last fortnight. It is not visible on the tape the analyst used, because that tape stopped updating when the report was finalised. The team concedes from the first corner of the match. The post-match review will call it a marking error. It is not a marking error. It is a pricing error: the club acted on a belief formed weeks earlier as though it were still current, against an opponent who had already moved.
What actually failed
Nothing in the analyst's method was wrong. The video coding was accurate. The sample size — eleven matches — was defensible by the standards of the trade. The failure sits entirely in the interval between when the belief was formed and when it was used. Somewhere in that interval, the fact the report rested on stopped being true, and no mechanism in the workflow was watching for that. The report had no expiry date written into it, no flag that said this tendency is derived from matches 1 through 11 of a coaching regime that ended on the 14th. It was treated as a fact about the opponent rather than a fact about the opponent as observed up to a certain point.
The opponent, meanwhile, was the better-informed party in the exchange, not through espionage but through simple continuity: their own coaching staff obviously knew what they had changed, when. The analyst's club did not know that they didn't know. That asymmetry is the entire story.
The concept: who holds the option
This is adverse selection, and it is worth being precise about the mechanism rather than waving at "the opponent adapted." In any exchange where one party knows more than the other, the informed party chooses when and how to act, and that choice is not neutral. It lands, specifically, on whichever counterparty is working from the older information. A scouting report, like a quote posted on an exchange or a premium set on an insurance policy, is a standing position taken against the world at a point in time. Anyone who has observed the world since that point holds a free option against it. They can choose to exploit exactly the part of your belief that is now wrong, and you have already committed resources — training-ground hours, personnel decisions, match-day shape — against the stale version.
Betting markets show the mechanism in its cleanest form. A bookmaker prices a match; a syndicate monitoring team-news feeds sees a confirmed late injury to a first-choice centre-back before the market re-prices, and backs the opposite side within seconds — the "steam move." The bookmaker's stale quote is not a mistake of judgement, it is a mistake of timing, and it costs money on every occurrence, priced in the shortened odds the syndicate captured. A performance analyst's dead-ball report works the same way, except the currency is goals conceded rather than basis points, and the counterparty is not a syndicate but a set-piece coach who changed a routine three weeks ago and simply waited for someone to still be marking the old one.
Why the fix isn't just faster scouting
The instinctive response is to scout more, later, closer to kick-off. That helps, up to the point where the interval between observation and use shrinks to the length of the preparation cycle itself — and then it stops helping, because the report is frozen again the moment it's finalised, and the opponent's next change lands in the new gap. This is the structural limit that separates three ways of building the model an analyst works from.
| Generation | Intake | Failure mode in this domain |
|---|---|---|
| Large Language Model | frozen corpus, fixed cutoff | brief built once from historical tape, used unchanged all season |
| Large World Model | bounded scene, refreshed per episode | brief rebuilt each week from the latest matches, still frozen for the week it governs |
| Large Universe Model | every stream running, dated and revisable | tendency held as "peeler run, confirmed through match 11, coaching change flagged match 12" — a belief with a timestamp attached to the claim itself |
The first two are both real workflows in professional analytics departments, and the second is a genuine improvement on the first — a weekly refresh narrows the exposure window from a season to a week. But it does not close the window; it relocates it. Every Thursday brief is still a quote posted against Saturday, and everything the opponent does between Thursday and kick-off — a training-ground tweak, a late injury, a return from suspension that changes the press — is a free option handed to them, again. Only a structure that treats tendencies as beliefs with provenance, continuously revised as tracking data, injury reports and transfer activity arrive rather than collated once per cycle, removes the interval rather than shrinking it. That structure does not exist as a product on the touchline. It exists as the direction the problem points, and nothing past it is a different kind of solution — it is the same solution, sustained.
Two objections worth taking seriously
The first: competing to be faster is often just rent extraction. Recruitment departments already burn budget on video analysts working through the night to catch changes half a day before rivals do, and none of that racing produces a better game, only a better-informed dugout at the other club's expense. This is correct, and it is the strongest case against treating this as a speed problem.
Every club just needs a faster video department than the opponent's, and the arms race cancels out.
It cancels out for the clubs, not for the fixture. What doesn't cancel out is disclosure. The useful reform is not "watch tape faster" but "mark every tendency with the date of its last confirmation and the event that might have superseded it." A brief that says the near-post peeler was last confirmed eight days before a coaching change is a different object from a brief that states the tendency as fact. The first can be reopened by anyone reading it on Thursday; the second cannot be reopened at all, because it doesn't admit that it has an age.
The second objection cuts the other way: widening intake widens what can be fed to you. Opponents already stage decoy patterns in open training sessions precisely because they know they're being filmed, and a club that streams every available signal — tracking feeds, leaked shape from a friendly, a rumoured transfer that never completes — is exposed to manipulation in a way a curated, vetted report is not. This is the more serious problem, and any analytics department that has been burned by a deliberately staged session in a public-facing training ground will recognise it immediately.
Live feeds are exactly where an opponent plants what they want you to see. A report finalised from vetted match footage can't be gamed after the fact — it's already closed.
Granted that the exposure is real. But a closed report is also unfalsifiable in the other direction: if the corpus was contaminated at collection — a scout misreading a friendly, a coding error in the tracking vendor's feed — there is no later evidence admitted that could correct it, because the report doesn't accept later evidence at all. A stream that carries its source and its date can be down-weighted the moment it's identified as staged; a tendency flagged "training-ground only, unconfirmed in competitive fixtures" is not treated the same as one confirmed across eleven league matches. The vulnerability to manipulation is a property of any open channel. The recovery from manipulation is a property only of one that keeps a record of where each belief came from and when.
The consequence
None of this makes the analyst's job different in kind. It makes the object they hand to the coaching staff different in kind: not a tendency, but a tendency with an age and a source attached, revisable the moment a newer stream contradicts it. That is the entire distance between a scouting report and a market maker's book — and it is the only distance left to close on this particular axis. Faster scouting buys time. Dated, provenanced belief buys the thing time was standing in for.