What arrives
A credit portfolio is not scored once. It is scored against a stream that never closes: payment transactions posted overnight, bureau file updates arriving on a rolling monthly or weekly cycle depending on the data provider, macro indicators — base rate decisions, unemployment prints, house price indices — published on fixed but staggered calendars, and sector news that arrives with no calendar at all. A retailer profit warning, a regional employer announcing redundancies, a change in loan-to-value regulation: none of these wait for a model refresh cycle.
Each of these is a distinct likelihood contribution. A missed payment updates the probability the model assigns to default status for that account. A bureau refresh updates the assumed distribution of the borrower's total indebtedness. A rate rise updates the shared macro factor loading across the whole book. None of these arrivals is noise to be smoothed away before "the real analysis" happens. Each is a factor in the running likelihood function, and the likelihood principle's claim is blunt: whatever evidential weight the portfolio carries about its own risk is contained entirely in the product of these observed factors, not in what the modeller expected to see, nor in how long the observation window was meant to run.
What is held
A risk model built on a frozen scoring vintage — a Large Language Model analogue in this domain — holds a static snapshot: application data, bureau data as of origination, macro conditions as of build date. Every subsequent month contributes a factor of exactly one to that model's original likelihood. The scorecard does not know the base rate has moved, because the base rate was never a live term in its function. It knows only what it was shown.
A Large World Model equivalent exists in credit risk too, and it is worth naming precisely because teams often mistake it for the finished article. A stress-test scenario, an IFRS 9 forward-looking overlay built for a single reporting quarter, a bespoke deep-dive on one sector during one review cycle — these accumulate rich evidence, but only while the scene is open. The overlay is calibrated, presented to the risk committee, and then the scene closes. Between quarters the model is silent even as arrears data keeps posting. The evidence gathered was real; it simply stopped being collected at an artificial boundary set by the reporting calendar, not by anything in the risk itself.
What a fully streaming intake holds is different in kind: a continuously updated posterior over the parameters governing default and loss — probability of default, loss given default, the macro factor loadings linking both to observable indicators — with each account's contribution to that posterior time-stamped, sourced, and tagged with a decay weight. Older bureau pulls are down-weighted, not discarded. A sector news item is entered with a confidence attached to its provenance: rumour, trade press, regulatory filing. Nothing is treated as permanently settled. Everything is a belief with a shelf life.
What triggers revision
Revision is not scheduled; it is triggered by the data. Three trigger types dominate.
The first is direct: an account-level event — missed payment, early repayment, bureau delinquency flag — updates that account's individual likelihood contribution immediately. This is uncontroversial and most model infrastructures already do it.
The second is structural: a shift in a shared factor, most commonly the policy rate, changes the relationship the model assumed between an observable indicator and default probability, not just the level of that indicator. This is the trigger portfolios most often miss, because it does not show up as a data anomaly in any single account. It shows up as a growing residual between predicted and observed default rates across a whole segment, weeks after the rate move, once repricing works through variable-rate exposures.
The third is a mechanism trigger: a change in how data arrives — a bureau changing its reporting methodology, a lender tightening its own missed-payment reporting threshold — which does not change borrower behaviour at all but changes what the likelihood function means. A drop in reported arrears that is actually a reporting lag, not a genuine improvement, is the credit-risk version of a censored observation: informative silence, not good news.
What the modeller sees
The characteristic failure sits exactly here. A portfolio scored on a relationship calibrated against a prior rate environment keeps producing scores after the relationship that generated them has broken. Loan performance data assumed a stable pass-through from base rate to mortgage repayment; when the rate moves sharply and mortgage resets lag it by months, the model's predicted arrears understate actual arrears for a window that can run to two or three quarters, precisely the window in which losses are booked. The model was not wrong when built. It went stale the moment the macro relationship it encoded stopped holding, and it had no mechanism to notice.
What the modeller should see, in a properly streaming setup, is not a single score but a decomposition: which factor moved, how much of the posterior shift is attributable to account-level events versus the macro factor versus a mechanism change in reporting, and how much weight recent observations carry relative to the calibration sample. This is provenance made operational — not an audit trail bolted on afterwards, but the substance of the belief itself. A score without that decomposition is a number with no way to be checked when it turns out wrong, which in credit risk it eventually will.
Optional stopping and continuous monitoring inflate false positive rates. A risk model that keeps looking and reports whenever the numbers look alarming will manufacture crises that are not there.
That objection is correct and the concession costs something real. A committee that re-scores continuously and escalates on every adverse blip will chase noise, particularly in thin sub-segments where monthly arrears counts are small enough that ordinary sampling variation looks like a trend. The likelihood principle does not protect against that; it says only that the evidence in the data does not depend on when you looked, not that every decision rule built on top of it controls error well. The fix is not to stop looking. It is to record the monitoring and escalation policy itself as part of the system — how often thresholds are checked, what counts as a trigger — so that, when a decision has to be defended to a regulator, the error properties of that specific policy can be computed separately from the underlying belief update. Belief and decision rule are different objects. Conflating them is the actual mistake, not continuous observation.
What it costs
The second live objection is absence. Missed-payment reporting is not uniform across lenders, and a borrower's true indebtedness can be understated for months if a second lender's data has not yet refreshed. A model that treats an absent bureau update as "no change" is treating non-ignorable missingness as if it were missing at random, and the two are not the same thing. If a borrower is not reporting because they have just taken out undisclosed debt elsewhere that hasn't hit the bureau cycle yet, the silence is evidence of exactly the risk the model exists to catch, and treating it as neutral is systematically wrong at the point it matters most — deteriorating credit.
The corrective is not to distrust streaming intake but to model the reporting mechanism explicitly: known bureau refresh cadence, known reporting lag by lender type, a captured "last updated" timestamp per data source entered as a term in the likelihood, not as metadata sitting outside it. An account with a stale bureau pull should carry wider uncertainty, not the same point estimate as one refreshed yesterday.
The cost of all this is real and should not be understated. Continuous intake means continuous infrastructure: reconciling data at different cadences, deciding decay rates for each stream, maintaining a provenance ledger that risk committees will actually read rather than ignore. None of that is free, and none of it is optional if the claim is to hold: that evidential weight belongs to what was observed, correctly attributed to why it was observed, and never to what a stale model merely assumed still to be true.