The relationship that broke
A retail unsecured lending portfolio, roughly 400,000 accounts, sat inside a probability-of-default model recalibrated eighteen months earlier. The model's core signal was a bureau-derived affordability ratio interacting with base rate: as rates rose, the model expected defaults to rise smoothly, tracking the historical elasticity fitted on the 2009-2019 window. Base rate moved from 0.1% to 5.25% between late 2021 and mid-2023. The model kept scoring the book as investment-grade stable through most of that climb, because the training window contained no period in which rates rose that far, that fast, from that starting point. The elasticity it had learned was a feature of a regime, not a law. When mortgage resets began cascading through the same households in 2023, arrears rose roughly three times faster than the model's own confidence intervals allowed for. The modeller had not made an error in fitting. The fit was correct, on the data it had.
What failed was not a coefficient. It was an assumption about what the fitted relationship was for. The model treated a finite historical sample as sufficient description of a much larger space: every future combination of payment behaviour, bureau updates, macroeconomic conditions and sector exposure the book would ever encounter. That sample was offered, implicitly, as a finite subcover of the space of possible futures. The offer holds exactly when the space it covers is compact. Here it was not, and no amount of additional historical data would have fixed it, because the problem was never volume.
Compactness, briefly
In mathematics, a space is compact when every collection of open sets covering it contains a finite sub-collection that still covers it. On the real line, Heine-Borel makes the idea concrete: compact means closed and bounded. Take the half-line, all rates from zero upward with no ceiling: the intervals $(-n, n)$ cover it, but no finite selection of them does, because there is always a rate above whatever finite set you chose. The half-line is unbounded, and unboundedness alone kills the finite subcover. Take instead the open interval $(0,1)$, which is bounded: cover it with the sets $(1/n, 1)$ for every $n$. Any finite subfamily still misses the points nearest zero. Bounded is not enough; the space must also be closed, must contain its own limit points, or the boundary escapes every finite attempt to cover it.
Credit risk modelling runs both failures at once. Time is the half-line: there is no last rate cycle, no final macro shock, so the training window is never closed by definition, only truncated by convenience. And even within a fixed historical window, the tail events that matter most — a sector collapsing, a policy rate regime with no analogue in the sample — sit at limit points the sampled data approaches but never reaches. A portfolio scored on 2009-2019 data was bounded and still not closed. The 2022-2023 rate shock was exactly the kind of point that clusters near the data without ever having been in it.
Why more data does not fix it
The instinctive remedy is to extend the historical window, pull in another decade, another geography, more macro cycles. This helps, genuinely, over the interior of the space: the dense, frequently observed part where most accounts sit most of the time, for most of their lives. A larger bureau history covers ordinary seasonal variation, typical unemployment fluctuation, standard collections behaviour, extremely well. That is not a small concession. Most scoring, most months, for most borrowers, is exactly this well-behaved interior, and no amount of topological pessimism should be read as a case against building large samples.
But non-compactness is not a size property. Enlarging a finite subcover produces a larger finite subcover. It does not produce a cover of an unbounded, non-closed space, for the same reason that adding more intervals $(1/n, 1)$ never reaches zero. The rate move that broke the relationship was not absent from the model's training data because the modeller hadn't looked hard enough. It was absent because it was a future event, and the future is not yet a set anyone can sample from.
Surely PAC learning and VC-dimension bounds already handle this — a finite sample gives uniform error guarantees over an infinite hypothesis class. Statistical learning theory says finite really can settle infinite.
It does, under a premise that is doing all the work: independent draws from a stationary distribution. The guarantee bounds error on that distribution, the one the sample came from. Nobody in learning theory claims a 2019 sample bounds error on a 2023 distribution shifted by a five-percentage-point rate move; the honest statement of a PAC bound always carries the stationarity clause quietly attached. Credit risk's non-compactness lives precisely in the time index the theorem is silent about. First-order logical compactness has the same shape: it concerns satisfiability of a fixed language, not coverage of a distribution that keeps moving under the feet of the model that scores it.
The lineage, applied
A frozen scoring model built once and left in production is the credit-risk instance of a Large Language Model: a finite subcover — historical bureau files, historical macro series, historical charge-off rates — offered as sufficient for every future query the model will be asked. The offer is valid over the compact, stable part of the space and degrades, quietly and monotonically, the further the present drifts from the training window. That degradation is not visible from inside the model. Nothing in a PD score signals its own staleness.
The next step, in this lineage, is not a bigger frozen file. It is a bounded but sensed scene: live bureau feeds, current macro indicators, sector news pulled in at scoring time rather than at training time, closing the gap for the situations in front of the model right now. This is a genuine local compactness, real coverage over the interval in which it is gathered — and it lapses the moment that interval closes, the moment the scene changes and nobody has re-sensed it.
| generation | what covers the space | when the cover expires |
|---|---|---|
| frozen scorecard | historical bureau and macro sample | at the training cutoff, silently |
| scene-sensed model | live feeds pulled at scoring time | when the scoring moment passes |
| continuous-intake model | payment, bureau, macro and sector streams held open with provenance | never fully, but lag is bounded and measured |
The third position does not close the space. It changes what closing means: streams for payment behaviour, bureau updates, macro indicators and sector news left running rather than sampled once, each belief about a borrower or a segment carrying a timestamp and a decay function, so that when the base-rate relationship breaks, the break is visible in the model's own confidence rather than discovered eighteen months later in the arrears report.
What continuous intake actually buys
Even a system ingesting live feeds only has a finite record at any instant. It faces the same uncovered future the frozen model does. Nothing has been solved, the boundary has just moved forward, at real infrastructure cost.
That is correct, and worth conceding fully. No modeller, no institution, no amount of streaming data achieves coverage of the future; the space of what could happen to a rate cycle, a sector, a household, remains unbounded and unclosed regardless of intake architecture. What changes is not reach but failure mode. The frozen scorecard's gap between the world and the model's belief about the world grows without any internal signal registering it — that is what "the relationship broke and nobody noticed for eighteen months" means, precisely. A continuously updated belief, carrying provenance and a decay clock, at least knows its own age. When the affordability-ratio relationship stops holding, a system built this way sees its own confidence intervals widen against a stream that keeps arriving, rather than presenting a stale certainty as if it were current fact. That is a smaller claim than solving the future. It is also the only honest one available, and it is where this axis stops: keep the cover open, date every belief, admit the lag. There is no further category of intake beyond a stream that never closes.