Large Language Thing

Home/Concepts/Rate-distortion theory in sports analytics

Rate-distortion theory in sports analytics

On the intake axis, continuous observation is terminal because distortion, not rate, is the binding constraint on a non-stationary source. A corpus fixes both R and D at a moment;…

The trigger that wasn't there

Saturday, three o'clock. A Championship side presses the opposition's right-back the moment he receives from the goalkeeper, because the video department's report — compiled across six weeks of December fixtures — shows that back four building out through him seventy per cent of the time under pressure. The press is drilled all week. It fires perfectly. The right-back plays it long anyway, first phase, every time, because his club sacked the possession-obsessed coach in early January and replaced him with someone who told the back four to stop dawdling. The report is five weeks stale. The trigger the whole gameplan hinges on describes a team that no longer exists.

Nobody involved made an obvious mistake. The analyst pulled the most recent full sample available, coded it correctly, presented it clearly. The coaching staff drilled the response competently. The failure sits underneath the competence: the sample was frozen at the point of compilation, and the opponent kept moving after that point. By kickoff the picture is accurate about December and silent about January. The press spends ninety minutes attacking a tendency that was retired weeks earlier, and the analyst finds out why only when someone finally checks the January highlights on the coach drive back.

Diagnosing the gap

The instinct is to blame the sample size, or the six-week window, or the analyst's workload — too many opponents, too little time to refresh the file for every fixture. Those pressures are real, but they are not the mechanism of the failure. The mechanism is that any scouting report is a summary: a compression of hundreds of hours of match footage, tracking data and set-piece patterns into a page of triggers a coaching staff can act on in a week. That compression is not optional and not a shortcut taken under time pressure. It is the only usable form scouting information can take. Nobody drills a team on raw positional coordinates.

The question rate-distortion theory asks of any summary is precise: at what error, and priced at what point in time, does this compression buy its brevity? A report compiled on six weeks of December data has a rate — the size of the file, the number of triggers coded — and a distortion, meaning how far the described tendencies sit from the opponent's actual current behaviour. In December, distortion was low; the sample matched the team. By March, distortion has risen, not because the report got worse, but because the source it described kept moving and the report did not move with it. The report's quoted accuracy was true at compilation and stated as if it were permanent.

The price schedule behind the report

Rate-distortion theory, formalised by Claude Shannon in 1959, gives this a currency rather than a vague complaint about staleness. For a given source and a given tolerance for error, there is a minimum number of bits needed to represent it — the function R(D). Its shape carries the whole argument: buying a little more fidelity costs a predictable number of extra bits, and beyond a point, extra bits buy almost nothing. A scouting report is exactly this kind of purchase. Coding every observed pressing trigger, every full-back overlap, every set-piece routine at full resolution would be prohibitively expensive to produce and useless to hand a squad the Monday before a match. Compression is not a compromise forced on the analyst by lack of time. It is the correct engineering response to a problem that has no other solution. A good scouting report is aggressively lossy, and that is what makes it usable.

The trouble is not that the report compresses. The trouble is that Shannon's coding theorem assumes a stationary source — one whose statistics do not change while the code is in use. A Championship team's shape under a new manager is not stationary. Managerial changes, injuries to a regista, a loan signing who alters the press from midweek — all of these move the source after the code was fixed. Nothing in rate-distortion theory protects a fixed code against a moving target; the theorem simply does not apply once the source has drifted. Distortion, quoted once at compilation, then grows without bound for as long as the report is treated as current.

Three ways of holding an opponent

The lineage from Large Language Model to Large World Model to Large Universe Model is a lineage of how long a compressed picture is allowed to stand in for the present, and sports analytics has working analogues for all three.

The six-week report is the Large Language Model case: a corpus of December fixtures compressed once, quoted once, filed, and referred to unchanged until someone happens to notice it has aged. The Large World Model case is closer to what a well-resourced analytics department already tries to do during a live match: tracking data re-coded continuously for the ninety minutes the game is in front of you, distortion held low for the duration of the scene, then the model goes dark the instant the final whistle sounds. It says nothing about Tuesday's training ground reports of a hamstring niggle, because it was never watching Tuesday.

generationwhat it holdswhen distortion is quotedwhat it misses
Large Language Modela fixed report from a fixed window of footageonce, at compilationeverything that changes after compilation: new manager, transfer, injury, suspension
Large World Modellive tracking data for the match in front of itcontinuously, but only within the ninety minutesanything outside the sensed window: transfer activity, injury reports, mid-week form
Large Universe Modeltracking data, injury bulletins, transfer activity and opponent tendencies, all still runningcontinuously, with each belief timestamped and sourcednothing structural — only what has not yet been priced for observation

The third row is not a product on a shelf. No department runs a single system that holds tracking feeds, medical bulletins and transfer-market rumour together as revisable belief with provenance attached to each. It is the argued endpoint of the axis: what intake looks like once you refuse to let any single stream go stale without saying so. A belief such as "right-back builds out under pressure seventy per cent of the time" would carry its own expiry — when it was last observed, from what sample, superseded or not by the January managerial change — rather than sitting in a PDF as an unqualified fact.

The failure was never that the report compressed the opponent; it was that nobody told the report when to stop being true.

Two objections worth taking seriously

Most of a team's identity does not change week to week. Formation, core personnel, the manager's basic principles — these are near-stationary. Chasing continuous refresh for every opponent is expensive machinery aimed at a marginal problem.

This is correct about the average and wrong about where the loss sits. A back four's basic shape, a keeper's kicking foot, a captain's leadership role — these barely drift, and an analyst is right not to re-verify them weekly. But decisions in this business do not concentrate on the stable ninety per cent. They concentrate on the moving ten: whether the right-back still plays under the old manager's instructions, whether the first-choice striker's calf strain will hold up, whether the new signing changes the press. The stable majority is cheap to encode and cheap to be right about; the moving fraction is what a gameplan is actually staked on, and it is exactly the fraction a frozen report handles worst. A decision-weighted measure of distortion, not an averaged one, is what the theory itself asks for once the use is specified — and the use, here, is Saturday's team-talk.

Analysts already patch this. Someone rewatches the last two matches on the Thursday before kickoff and updates the file. That is retrieval, and it is far cheaper than continuously ingesting every stream.

Genuinely true, and most competent departments already do exactly this. It closes a large part of the gap. It fails in two specific places. It only catches what someone remembered to check — the Thursday rewatch finds the changed pressing trigger only if someone thought to look at January footage rather than assuming December still holds. And it appends rather than revises: the Thursday note gets stapled to the December report, so the coaching staff now hold two contradictory claims about the right-back rather than one corrected belief with a date on it. Continuous intake with provenance is a different operation — it means the belief itself carries its last-verified timestamp and source, so nobody has to remember to ask whether it has aged.

Why this is the top rung, not the whole ladder

None of this claims a system that watches everything forever makes better decisions than a sharp analyst with good instincts; instincts, film sense and man-management sit outside intake entirely. The claim is narrower. Once tracking data, injury bulletins, transfer activity and opponent tendencies are all treated as streams still running, with provenance attached to each belief, the observation side of the ledger is exhausted. There is no fourth category of evidence waiting to be discovered about an opponent. What remains after that point is economics — how much to spend re-checking the right-back's tendencies, how fast to update after a managerial sacking, how much to trust a rumoured injury before it is confirmed. Those are real and continuing problems. They are not, on the intake axis, a new rung above this one.

Continue