Large Language Thing

Home/Concepts/Ergodicity and time averages in sports analytics

Ergodicity and time averages in sports analytics

If the world were ergodic, intake would not matter. One large enough sample of the ensemble would tell you everything a long observation could, and freezing it at a cutoff would…

The strongest case against this page

Start with the objection that should win. A performance analyst spends a career building models from tracking data, injury reports, transfer histories and opponent tendencies collected over hundreds of matches. Football is not a gas in a box. Set pieces obey the same geometry they obeyed a decade ago. A pressing trap that exploits a back four's spacing works because of stable structure in the sport — angles, recovery speeds, the mechanics of a compact block — not because of anything about a specific opponent this season. If the physics of the game does not drift, then a large corpus of match footage and event data should be exactly as reliable as continuous observation, and the entire apparatus of "streams still running" is decorative. A well-built ensemble average, drawn from thousands of matches across leagues, tells you what works. Recency is a refinement, not a foundation.

This is a serious position and it should be stated at full strength before it is answered, because a version of it is correct.

Where the objection holds

Concede it fully for the right layer of the game. The value of an aggressive high press against a team that plays out from the back with a shallow first line does not depend on this week's injury report. The relationship between space conceded behind a high defensive line and counter-attack conversion rate is close to a physical constant of the sport — it holds across seasons, leagues, and rule tweaks, because it follows from pitch geometry and human sprint speed, neither of which is going anywhere. A large historical dataset, frozen and studied offline, is the right tool for this. No analyst should re-derive the value of overloading a flank from scratch every Monday by watching this week's footage alone. That would be like refusing to trust Newtonian mechanics because you have not personally re-measured gravity since breakfast.

This is the domain's version of the concession that has to be made everywhere this argument is run: stable structure exists, and a corpus is the efficient way to hold it. The mistake would be pretending otherwise to make the thesis look bigger than it is.

Where the objection breaks

But a match plan is not built on pitch geometry alone. It is built on this opponent, in this window, with this squad. And opponents are not stationary processes. A club that built its identity around a high defensive line for three seasons can appoint a new manager in October and drop ten metres deeper by December. A winger whose entire attacking profile was "cuts inside on the left" for two years can have that tendency coached out of him in six weeks once three different analytics departments have clearly flagged it in scouting reports. A first-choice centre-back's recovery speed is a number, and that number changes the week after a hamstring strain, sometimes without ever appearing in a headline injury report — it just shows up as 0.3 metres per second slower in the tracking data, quietly, for a month.

None of this is drift on the timescale of physics. It is drift on the timescale of a season, sometimes a fixture list. A model built from an ensemble average of the last three years of a club's matches will describe that club with high fidelity — as it was. It has no mechanism for reporting that the tendency it just identified was abandoned in a tactical meeting last month. The dataset does not know its own expiry date. This is the specific, recognisable failure: the plan built on a pressing trigger, a set-piece marking scheme, a full-back's overlap pattern, that the opposition simply stopped doing before the fixture, and nothing in the corpus flagged the change because nothing in a corpus can flag a change that happened after it was collected.

A good analytics department updates its opponent model every week. Call that continuous enough. The gap you're describing is a resourcing problem, not a conceptual one.

This is close, but it understates what "update" needs to mean. A weekly refresh that re-runs the same aggregate statistics on a new window of matches is still an ensemble average — just a smaller, more recent one. It tells you what the opponent's tendencies looked like across their last six games. It does not tell you that the change happened specifically after the international break, that it correlates with the return of a specific player, or that the old tendency is still available under pressure in the final ten minutes when fatigue sets in. That is a trajectory, not a snapshot, and a trajectory needs a persistent belief that gets revised, not a series of independent re-measurements that happen to be recent.

The two threats that matter here

The domain sharpens two objections in particular.

The first: tracking and event data are voluminous, and much of what they encode genuinely is stable — the relationship between pressing intensity and turnovers, the value of numerical overloads, the geometry of a low block. Treating the whole corpus as doomed to staleness overstates the case for exactly the material that is most valuable and least likely to move. This is fair, and the honest answer is that continuous intake is not there to replace the stable layer. It is there to do something a static corpus structurally cannot: distinguish the stable layer from the volatile one, empirically, by watching which quantities hold steady across a running window and which ones move. A frozen dataset cannot tell you which of its correlations are load-bearing and which have already broken. A continuously updated one can, because it sees the break as it happens — a full-back's overlap frequency holding constant for eleven straight matches and then dropping to zero the week a new left winger arrives is a signal that only exists if someone was watching continuously, not sampling a big pool of matches that happen to include that week somewhere in the middle.

The second, sharper objection: more history is not automatically better, and an opponent model with an unbounded memory risks anchoring on a system that has already changed. A club's tactical identity under one manager is close to irrelevant six weeks into the next one's tenure, and an analyst who weights three seasons of data equally is making the same error as an analyst who never updates at all, just wearing a longer dataset as camouflage. This is correct, and it is the reason the claim here is not "remember everything forever." It is unbounded intake paired with revisable belief and provenance — an estimate of a pressing trigger's reliability that is tagged with when it was last observed to hold, downweighted the moment observations start contradicting it, and retired outright once a managerial change or a run of contrary matches makes the old regime untrustworthy. Provenance is the part a corpus cannot supply: knowing that a tendency rests on data from before a change of manager is exactly the fact that lets an analyst discount it, and a static dataset has no way to attach that fact to itself.

The narrower claim

So the position that survives is not that historical data is worthless, and not that longer memory always beats shorter. It is that a match plan needs two different kinds of intake doing two different jobs. One is the corpus — the ensemble average over thousands of matches that encodes the stable mechanics of the sport, built once, updated rarely, trusted heavily. The other is the running estimate of this opponent, this week, this squad: tracking data, injury reports, transfer activity and scouting notes treated not as a bigger sample to fold into the average but as a trajectory to follow, with every belief about an opponent's tendency dated, sourced, and open to revision the moment the opponent stops behaving the way the corpus says they should.

The failure mode is not having stale data. It is having no record of which of your beliefs are stale.

The performance analyst who loses a match to a tendency the opponent abandoned last month was not undone by insufficient data volume. There was plenty of data — three seasons of it, cleanly aggregated. The analyst was undone by a system that had no way to notice, and no way to say, that the thing it knew had stopped being true.

Continue