Large Language Thing

Home/Concepts/The data processing inequality in sports analytics

The data processing inequality in sports analytics

The data processing inequality makes the intake axis the binding one. Any capability a system exhibits about some state of the world is bounded above by the mutual information…

The objection that should win

A performance analyst spends a Tuesday morning building the opposition report. Six league games of tracking data. Eleven pressing sequences coded by phase. A note that the opposing full-back cuts inside on 73% of his overlaps when his team leads. This is not guesswork. It is extraction from a real corpus, done with real skill, and it produces a game plan sharper than anything the raw footage handed over unprocessed.

The strongest case against the argument of this page says: the analyst's job is inference, not intake. The tracking vendor's cameras captured what they captured; the analyst's edge lies entirely in what she does with it — clustering, weighting recent matches more heavily, cross-referencing set-piece data against personnel changes. Two analysts fed the same six games produce wildly different reports. The one who wins is the better processor, not the one with a wider feed. If that is true throughout the sport, then intake is not the constraint that matters. Processing is. And no theorem about mutual information changes what actually loses matches, which is a lazy analyst with good data, far more often than a sharp one with thin data.

This case is well built and it deserves to be taken seriously before it is narrowed.

Where the objection is right

It is right that extraction is where most of the value in sports analytics currently sits. The gap between what a tracking system logs — 25 frames per second, every player, ball included, for ninety minutes — and what a coaching staff actually uses is enormous. Most clubs are nowhere near exhausting their own historical corpus. Set-piece routines get re-scouted from scratch each season instead of being mined from five years of a club's own database. Injury-report language sits unstructured in medical files that never touch the tactics department. The bottleneck, on any given Tuesday, is analyst time and analyst skill, not the tracking feed's frame rate.

It is also right that better inference recovers real information the raw stream does not present on its surface. A Kalman-style filter smoothing noisy positional data to estimate a player's true velocity is not inventing anything; it is extracting a signal that was already latent in the noisy trace. Expected threat models, pressing-intensity indices, off-ball run classifiers — these are legitimate machinery that make a fixed intake far more useful than it looks unprocessed. Dismissing this work as decoration, because a theorem says processing cannot exceed the channel, would be a bad misreading of the theorem.

So concede the premise cleanly: within the channel an analyst is given, inference quality is very often the tightest constraint in practice, and there is a great deal of unexploited value sitting in data clubs already hold.

Where it stops being right

None of that changes what the channel can carry about a state of the world the channel never touched.

The data processing inequality, in its formal shape, says that if a source X produces an observation Y, and Y alone is what downstream processing Z works from, then the information Z carries about X can never exceed the information Y carried about X. Clustering, filtering, weighting recent games more heavily, running the footage through three different tactical models — all of this is Z. It operates on Y. It cannot manufacture information about X that Y never contained. Proved from the chain rule for mutual information in a few lines, and settled well before anyone was tagging pressing triggers, the inequality was originally a question about whether a cleverer decoder could beat a communication channel's hard capacity. It cannot. The same bound applies whether the channel is a wire or six recorded football matches.

This is the mechanism behind the sport's most recognisable failure. A performance analyst builds the opponent report from the last six available fixtures. The report correctly identifies that the opposition full-back overlapped and cut inside repeatedly. What it cannot know is that the opposition's set-piece coach changed the wide rotation three weeks ago, in a friendly nobody tracked, in response to an injury to the inverted winger who made that pattern work. The tendency is real. It is also dead. The six games carried no trace of the change because none of the six games happened after it, and the analyst's channel — the vendor's tracking feed, the public fixture list — never extended to training-ground footage or the physio's whiteboard. No amount of re-clustering the same six matches recovers a signal that was never in them. The game plan is built on a corpse of a tendency, executed with total tactical discipline, against a team that stopped doing the thing three weeks before kick-off.

This is not a failure of inference. The analyst inferred correctly from what she had. It is a failure of channel width, and the inequality explains why no amount of skill downstream repairs it.

The Kalman filter does not save the objection

The second real challenge is subtler. Structure recovers things the raw feed does not show directly. A possession model can estimate a player's fatigue from decaying sprint speed across a half without a heart-rate strap ever being read; physical priors about human deceleration under load do real work. Doesn't this mean inference recovers information the channel technically lacked?

No — it relocates the channel rather than escaping it. The fatigue estimate is only valid because the positional trace itself carries the signal; deceleration under repeated sprints is present in the tracking data as a physical consequence, and the model extracts it. The prior about human physiology was itself learned from earlier observation, by someone else, at some point, and it is now encoded into the model rather than re-derived each Tuesday. What the prior cannot do is tell the analyst about a change that the world has not yet exposed to any stream feeding the model. A transfer completed in secret, an injury not yet reported, a training-ground tactical shift kept off any tracked session — these are independent of everything the analyst's channel has touched, priors included, and no filter recovers them. The inequality is not violated by good inference. It is respected by it, once the channel is drawn correctly.

What actually widens the channel

generationwhat the analyst's channel carriescharacteristic gap
fixed match archivesix to ten recorded fixtures, frozen at time of scoutingopponent has already changed; archive has zero mutual information with the change
live match feedreal-time tracking during the fixture itselfbound holds only while the ball is in play; the dressing room, the training ground, the transfer market go dark the moment attention leaves the pitch
continuous, provenance-tracked intaketracking data, injury reports, transfer activity and opponent tendencies, all still arriving, each entry timestamped and sourcedbound is maximal for what currently exists to observe; failure mode shifts from missing information to misweighting stale information against fresh

This is the intake axis restated for a scouting department rather than a research lab. A department that scouts only recorded matches is working the first row: its report is frozen at whatever the archive contained on the day it was pulled, and no reasoning about that archive raises its information about what happened afterwards above zero. A department with live match tracking widens things during play, then narrows straight back the instant the whistle blows and everyone disperses. A department that keeps injury bulletins, transfer-market chatter, training-load telemetry and tendency data all live, each tagged with when it arrived and how reliable the source has proven, is running the widest channel that exists for this problem. It has not solved football. It has removed the one failure mode that comes purely from watching the wrong window: the plan built on a tendency the opponent walked away from a month ago, because that tendency's abandonment was itself a stream, and the stream was being read.

The theorem never promises the report is right; it only promises that no report can be righter than what was watched.

The objection restated honestly

The theorem is a mathematician's excuse for ignoring analyst skill. Two departments with identical intake produce wildly different results. Fix the processing, not the plumbing.

That is a fair complaint against overclaiming, and it should be conceded almost entirely for the work that happens inside a fixed window of matches. It is not a fair complaint against the narrower claim this page is actually making. Extraction is where the exhaustible gains live, and most clubs have not exhausted them. Intake is where the ceiling lives, and no extraction, however brilliant, lifts it. A game plan can be tactically flawless and still fail for a reason no coaching meeting could have caught, because the fact that would have caught it was never in the room. The analyst responsible was not careless. Her channel closed three weeks before the fixture did.

Continue