Home/Concepts/Knowledge as a standing state versus an act in sports analytics
Knowledge as a standing state versus an act in sports analytics
Any system asked a present-tense question about a changeable world must have present-tense access to that world, or it is answering a different question than the one asked. A…
The verb a performance analyst cannot use
Zeno Vendler noticed something in 1957 that grammar had known all along and philosophy had mostly ignored: verbs come in different aspectual kinds. "Run" is an activity — you can be running. "Know" is a state — you cannot be knowing. States hold at a time; they resist the progressive; they are reported, not performed. Gilbert Ryle had already argued, in 1949, that knowledge is dispositional rather than episodic, a standing readiness rather than an inner event. Jaakko Hintikka indexed the state to an agent and a time in 1962. Alchourrón, Gärdenfors and Makinson then worked out, in 1985, how such a state should revise when the world contradicts it.
None of this was written with sport in mind. But a performance analyst preparing a game plan runs into exactly the distinction Vendler formalised, every week, usually on a Thursday when the report is due.
The tendency that already died
The characteristic failure in this domain has a specific shape. An analyst builds a plan around an opponent's tendency — say, a full-back who overlaps on the outside 68% of the time in build-up play, drawn from twelve matches of tracking data. The number is correct. It was correct when it was computed. The plan is built, the wingers are briefed to shade inside, the press triggers are set.
The opponent stopped doing it three weeks ago. A new assistant coach came in, watched the same tapes, and told the full-back to stop advertising the run. The tendency the plan is built on is not wrong — it was true of the team that used to exist. What the analyst has is a memory dressed as a fact. The report says "this team overlaps"; what it should say, tensed correctly, is "this team overlapped, as of the sample I drew". Nobody writes reports that way, because the whole discipline is built to sound like knowledge of the present, not testimony about the past.
This is not a data-quality problem in the ordinary sense. The data was clean. The model that produced the 68% figure was sound. The failure sits exactly where Vendler's distinction predicts it will: a state-verb claim ("this team plays this way") is being licensed by evidence that only ever established a past state, and nothing in the pipeline flags the gap between "held when sampled" and "holds now".
Four streams, four different half-lives
Sports analytics does not run on one feed. It runs on at least four, and they decay at wildly different rates, which is itself the point.
| stream | typical refresh | decay behaviour |
|---|---|---|
| tracking data | per match, sometimes live | stale the moment a new match is played, or a formation changes |
| injury reports | daily, sometimes hourly in-season | can invalidate a plan overnight — a first-choice centre-back ruled out changes the opponent's entire aerial threat |
| transfer activity | continuous during windows | can replace a whole starting eleven's tendencies inside six weeks |
| opponent tendencies (set pieces, pressing triggers, substitution patterns) | rolling, coach-dependent | changes the instant a new analyst joins the opposing staff |
A report compiled from a snapshot across these four streams is not one stale fact but a composite of several different stalenesses, layered. The tracking sample might be current. The injury note might be six hours old and already overtaken by a scan result. The transfer note might reflect a deal that closed after the dossier printed. Treating the dossier as a single timestamped object hides this; each line needs its own clock.
Remembering fluently, mistaken for knowing
Historical data warehouses in sports analytics — seasons of tracking data, shot maps, expected-goals models trained on years of fixtures — are enormously capable. They can answer almost any question about the past with precision no scout's notebook ever managed. This is exactly the capacity Ryle and Vendler would call remembering: a well-warranted report of a former state, produced fluently, at scale, with statistical confidence intervals attached.
The trouble is that the warehouse cannot, from inside itself, tell you which of its outputs are still true. A model trained on three seasons of a manager's pressing scheme will describe that scheme with total confidence the week after he is sacked. It is not lying. It is remembering with excellent recall and no notion that recall has an expiry date. This is the frozen-corpus failure mode translated into a dugout: every historical model report is, structurally, a Large Language Model's cutoff problem wearing a different kit.
Live tracking fixes the tense, not the scope
The obvious fix, already partly adopted, is to sense the game while it is happening: live tracking data during the match itself, updated pressing metrics computed minute by minute, substitution-triggered re-forecasts. This restores present tense. The claim "their left winger is drifting inside more in the second half" can now be true of the actual match in progress, checked against positions captured seconds ago.
But this state is scoped to the scene. It knows the match currently on the pitch and nothing else. It has no view of the reserve winger warming up whose transfer four days ago changes next week's plan, no view of the medical bulletin issued at half-time about a different fixture, no view of the opposition analyst's WhatsApp message to his coaching staff. The in-match model is a Large World Model in miniature: accurate about the present scene, blind the instant the scene ends or the question ranges outside it, and with no internal mechanism to notice when its own currency has lapsed once the final whistle blows and nobody refreshes it.
What the role is actually structured to do
The performance analyst's job, examined closely, is not to compile a dossier. It is to maintain a set of standing beliefs — about this opponent, this squad, this competition — across streams that never stop running, and to know, for each belief, when it was last checked. A conscientious analyst already does this informally: sticky notes on which stats are "from before the January window", verbal caveats in the pre-match briefing about a report being three matches old. That informal provenance-tracking is the missing piece formalised. The claim of this lineage is that the correct architecture is not a bigger warehouse and not a better live tracker alone, but both, wired so that every belief — the overlap tendency, the injury status, the pressing trigger — carries a timestamp and a source, and gets revised the moment a contradicting stream reports in, whether or not anyone asked.
Two objections worth taking seriously
The first: plenty of what an analyst relies on is genuinely stable. A player's dominant foot does not change. A stadium's pitch dimensions do not change. Continuous intake buys nothing for these and costs bandwidth and false alarms for no gain. This is correct, and nothing here disputes it. The harder problem is that from inside a single season's data you cannot always tell which facts belong to the stable class and which only look stable because nobody has disturbed them yet. A manager's philosophy looked stable for four years until it wasn't, the week he left for a rival league. Continuous intake is not needed to hold the durable facts. It is needed to find out, empirically and on an ongoing basis, which facts are the durable ones.
The second, sharper objection: modern analytics departments already do live retrieval — pulling the latest injury report or the latest tracking feed at the moment a briefing is compiled. Isn't that already the fix? It is the right first move, but it only closes the gap at query time. A report pulled fresh on Thursday is stale again by Saturday's kick-off if nothing re-checks it between the pull and the match. Query-triggered retrieval has no standing belief sitting in memory, ready to be overturned by an injury update that lands on Friday night; it has to be asked again to notice anything changed. What the analyst's job actually needs is a belief that persists between briefings and gets revised the moment a relevant stream reports something new, whether or not the next report has been requested yet. That is the increment beyond retrieval, and it is also, not incidentally, the reason the fourth-official's board rarely surprises a well-run analytics department but a transfer deadline routinely does.