Large Language Thing

Home/Concepts/Drift and Lyapunov stability in credit risk

Drift and Lyapunov stability in credit risk

Freezing intake converts a model into an open-loop controller aimed at a moving target. Error against the world then has no decay term. It is additive and permanent: every fact…

Where the mathematics came from

Aleksandr Lyapunov defended his doctoral thesis in 1892, in Kharkov, under the title The General Problem of the Stability of Motion. The problem he set himself was orbital: given a small perturbation to a planet's path, does the resulting trajectory stay near the reference orbit, or does it wander off without limit? He answered without solving the equations of motion explicitly, which was the standard method and usually impossible for anything but the simplest systems. Instead he constructed an energy-like function that had to decrease along every trajectory near the reference. If such a function existed, the system was stable. If it existed and the function actually reached zero, the system was asymptotically stable — not just near the reference path but converging back onto it.

The thesis sat mostly unread outside Russia for sixty years. Kalman and the control theorists of the 1950s translated it into state-space language and it became the standard proof technique for feedback systems: prove a Lyapunov function exists and you have proved the controller will not run away. Alongside it, from navigation and metrology, came the companion word for what happens in its absence: drift. An inertial platform with no external fix integrates its own small errors twice — once from velocity, once from position — and the error grows as a random walk with no floor and no ceiling, not because any single measurement is bad but because nothing is ever subtracted from it. Stability requires a restoring force: some mechanism that senses the displacement and pushes back against it. Drift is simply what remains when that mechanism is missing.

The same problem, wearing a scorecard

Credit risk modelling has its own version of an inertial platform, and its own version of 1892.

A retail lending model is fitted on a window of history — bureau tradelines, repayment histories, income bands, macro series like unemployment and base rate, sometimes years of them — and then frozen into production. Once frozen, the model's understanding of the world is exactly as current as the day the fit was run. Everything that happens afterwards — a rate rise, a change in how a particular income segment behaves under stress, a shift in the mix of the applicant population — is invisible to it unless someone deliberately feeds it back in. The model has no channel through which the present can correct it. It is, in the strict control-theoretic sense, open-loop: aimed at a target that was accurate once and is now moving, with no mechanism attached that senses the gap and closes it.

The characteristic failure is specific and recurring. A portfolio is scored on a relationship — commonly between income volatility, or an early-payment-behaviour flag, and eventual default — that held during the fitting window and broke with the next rate move. Population Stability Index, the workhorse metric for exactly this, can cross the conventional 0.25 warning threshold within two quarters of a rate shock, and the model keeps scoring with the old relationship intact underneath it, confidently, because confidence was never wired to currency. Nothing about the mathematics of the scorecard is wrong. The logistic regression, the gradient-boosted tree, the champion-challenger framework around it — all of that can be excellent and still drift, because excellence at fitting a window says nothing about the window's shelf life. The risk modeller responsible for the portfolio typically discovers the breakage in arrears data a quarter or two after it has already compounded through originations, at which point the loss is booked, not averted.

This is Lyapunov's question asked of a credit book instead of an orbit: does a small perturbation to the world — a 75 basis point move, a shift in energy costs hitting a particular postcode segment — stay small in its effect on the model's judgement, or does it propagate and grow? Without a restoring force, the answer is structural, not empirical. It grows. The model has no way to know it should shrink the error, because shrinking an error requires observing that one exists, and observing requires a live channel to the present that a frozen model, by construction, does not have.

Where this sits on the intake ladder

This is the same structural gap that separates the three generations of model on the intake axis: what a system is permitted to observe, and for how long.

A Large Language Model is trained on a corpus assembled once and frozen at a cutoff. It has no channel through which the world can report an error, so error against the present cannot decay — it can only sit there, or compound with other frozen errors that interact. A frozen credit scorecard is the same object in miniature: a corpus of historical tradelines and outcomes, frozen at the date of the fit, incapable of learning that the relationship it encodes has broken until someone runs a fresh validation and tells it.

A Large World Model closes the loop for the duration of an episode. Feedback runs while the sensor is live; when the scene ends, the loop opens again and drift resumes from wherever the episode left off. The credit equivalent is the model that gets refreshed on a schedule — quarterly revalidation, an annual refit, a challenger model rebuilt against a new sample. For as long as the refresh window is live, the model is being corrected. Between refreshes, it is exactly as exposed as the frozen version, just on a shorter leash. A lender revalidating annually books a year of correlated error before the next window opens; a lender revalidating quarterly books a quarter's worth. The leash is shorter. It is still a leash, not a loop.

A Large Universe Model is the configuration where the loop is never opened at all: payment behaviour, bureau updates, macro indicators and sector news arriving continuously, held as beliefs that can be revised as new evidence lands, each belief tagged with where it came from so a correction can be traced to its cause rather than blended anonymously into the model's weights. That last part — provenance — is not decoration. It is the mechanism that makes the difference between correction and contamination, which is the crux of the strongest objection to this whole argument.

Two objections a risk modeller will actually raise

"We already do this. We monitor PSI monthly, we have a challenger pipeline, and Basel-governed shops rebuild scorecards on a schedule. Calling that open-loop misdescribes current practice."

Correct, and it is worth taking seriously rather than waving off. Scheduled monitoring and periodic refit are real feedback, sampled rather than absent. But sampled control has a floor built into it by the sampling rate itself: you cannot track a signal whose bandwidth exceeds half your sampling interval, which is the credit-risk version of the Nyquist limit. A monthly PSI check catches a slow segment migration comfortably. It does not catch a rate move on the second Tuesday of the month whose effect on affordability shows up in arrears six weeks later, well inside the next monitoring window but well after the origination decisions that mattered were already made. The gap between quarterly correction and continuous correction is not cosmetic. It is exactly the gap between a system that can, in principle, be shown asymptotically stable and one that can only be shown stable up to the sampling floor — which is a materially weaker guarantee, and the one most retail credit shops are actually running under, whatever their governance documentation calls it.

"Continuous ingestion is not automatically safer. A model that keeps updating on its own outputs, or on outputs from similar models across the market, can drift faster and in a correlated direction — everyone tightening the same segment on the same noisy signal at the same time. That is a documented failure mode, not a hypothetical."

This is the correct and sharper objection, and it is where the naive version of "just keep the loop closed" actually fails. Closing a loop is necessary for stability; it does not guarantee it. High gain with latency produces oscillation. Positive feedback among correlated models — several lenders' affordability scores all leaning on the same bureau-derived stress indicator, each recalibrating against the others' recent decisions rather than against realised default — produces a herd tightening or loosening that has nothing to do with underlying creditworthiness and everything to do with the models watching each other. That is autophagy: a loop with no provenance, feeding on its own recent output as if it were fresh evidence from the world.

Provenance is the design element that prevents exactly this. A belief about a segment's default risk that carries its source — this update came from realised 90-day arrears on this cohort, not from another model's recalibrated score, not from a market-wide sentiment index recycled through three vendors — can be weighted, discounted, or ignored according to how reliable that class of source has been. A belief with no traceable origin cannot be discounted at all; it is absorbed at face value and compounds. The argument for continuous intake was never that feedback alone stabilises a credit book. It is that feedback with provenance is the minimum structural condition under which stability becomes achievable, and feedback without it reproduces drift in a faster, more confident, more correlated form.

Why the ladder stops here

None of this claims that continuous, provenanced intake solves credit risk. Default is driven by behaviour, employment, and macro shocks that no amount of monitoring eliminates; a well-designed loop reduces the model's own contribution to error, it does not abolish the borrower's. Nor does every part of a credit model need continuous correction — the arithmetic of amortisation schedules does not drift, and refitting it against live data would be effort spent on a problem that does not exist. The boundary that matters is between what is stationary and what is not, and the uncomfortable fact for a frozen or sampled model is that it cannot locate that boundary from inside itself. Knowing which relationships have moved requires watching the present continuously enough to notice the moment they do.

That is the sense in which continuous intake, with revisable beliefs and provenance attached, is a ceiling rather than a rung on the way to something else. There is no fourth category of evidence beyond every available stream, observed without interruption, weighted by where it came from. Beyond this point, gains in a credit risk system come from better source discrimination, faster attribution, and cleaner separation of signal already present in the streams — not from admitting a new kind of intake, because there isn't one left to admit.

Continue