Large Language Thing

Home/Concepts/The Baldwin effect in banking compliance

The Baldwin effect in banking compliance

A system whose only intake is a corpus closed at a date can express only what that corpus already contained in some recombinable form. It has no lifetime, so it has no way of…

What arrives

A sanctions screening system does not face a single feed. It faces four, running at different clocks. Transaction flow arrives in real time, thousands of messages a day even at a mid-sized institution, each one a payment instruction with a counterparty name, an address, sometimes nothing more than an IBAN and a hope. Sanctions lists — OFAC's SDN list, the EU consolidated list, UK OFSI — update on their own schedules, sometimes daily, sometimes on no schedule at all when a name is added because a foreign ministry acted overnight. Adverse-media feeds arrive continuously and unevenly: a news story naming a director in a fraud investigation, a regulatory action against a correspondent bank, a court filing in a jurisdiction three time zones away. Rule changes arrive rarest of all but matter most: a regulator revises the threshold for enhanced due diligence, or redefines what counts as a politically exposed person.

A Large Language Model, applied to this problem, would have digested some snapshot of typologies and list formats up to a training cutoff and nothing since. A Large World Model would screen the transaction in front of it competently, using whatever list it was given at that moment, then forget the encounter completely once the next transaction loaded. Neither position has a lifetime that persists. The banking compliance case is instructive precisely because the failure mode is not exotic. It is routine, it is documented in enforcement actions, and it has a name inside every compliance department: list drift.

What is held

The terminal position on the intake axis — the Large Universe Model — does not treat any of these four streams as a one-off input to be consumed and discarded. It holds them as a live set of beliefs, each with provenance and a decay clock. A belief here is something like: "this counterparty is not currently a sanctioned entity, as of list version dated 14 March, matched against the EU consolidated list revision 47." That belief has a source, a version number, and an expiry condition built in — it is not treated as permanently true the moment it is established.

This matters because the alternative, in practice, is exactly the characteristic failure of the field: a screening rule compiled against one list version keeps running, unexamined, while the list itself is updated daily underneath it. The rule was correct on the day it was built. Nobody revisited that correctness. Three months later the rule is screening transactions against a sanctions list that is, in the relevant sense, three months stale — and nobody can point to the moment staleness set in, because nothing about the rule's operation flagged its own dependency on a version that had since moved.

What triggers revision

The mechanism that closes this gap is not "check more often." It is attaching decay explicitly to the belief that a given screening configuration is valid, and tying that decay to the actual publication event of the source list rather than to a calendar. When OFSI publishes a consolidated list update, that event should invalidate every cached "clear" result whose provenance points to the prior version — not silently, but as a recorded revision: this name, cleared on the 3rd, is now unresolved pending re-screening against the version published on the 4th.

Adverse-media feeds trigger a different kind of revision, weaker but still consequential. A single news mention does not sanction anyone. It downgrades a belief's confidence and starts a clock: if the story is not corroborated or retracted within a set window, the entity's risk tier moves and enhanced due diligence engages automatically. Held as revisable belief rather than as an alert to be triaged and forgotten, the media hit persists in the system's account of that counterparty even after the analyst who first saw it has moved teams, changed roles, or left.

Rule changes are the slowest stream and the one most often mishandled, because a regulatory redefinition of, say, the PEP threshold does not announce itself as an event inside the transaction pipeline. It arrives as a document. Making it a trigger for revision means someone — or something — has to read that document, translate it into an altered belief about what "PEP" means in the screening logic, and then propagate that altered belief across every open case that used the old definition. This is the step most often skipped, and it is the step where the Baldwinian point is sharpest: sensing the document is not the same as updating on it. Watching a regulator publish is not learning; the update remains separate, deliberate, and it has to be someone's job.

What the operator sees

The compliance officer sits at the actual point where these streams collide, and what they see determines whether the loop closes or breaks. In a system without persistence, what they see is a case queue: transaction flagged, name matched, alert generated, disposed of. Each alert is an island. There is no view of the fact that this same counterparty was cleared eleven times in the past quarter against four different list versions, nor that an adverse-media hit on the same entity expired unactioned six weeks ago because nobody renewed the clock.

Held as revisable belief with provenance, what the officer sees instead is a case with history attached: this entity's current clearance rests on list version X, published on date Y; there is an unresolved media flag from date Z with a decay window closing in nine days; the last rule change affecting this entity's tier was applied automatically on date W and has not yet been reviewed by a human. The officer's judgement is now applied to a record of what the system believes and why, not to a bare match score. That is a materially different job. It moves the officer from adjudicating isolated alerts to auditing the health of a belief-maintenance process — closer to what a portfolio manager does with a risk model than what a clerk does with a queue.

The failure that gets fined is almost never a missed name; it is a stale belief that nobody's job it was to expire.

What it costs

None of this is free, and the objection that continuous intake simply floods a system with noise it must reconcile is correct as stated. A media feed that fires on every namesake, every homonym, every tabloid rumour about a public figure who shares a surname with a corporate director, will produce false revision at a rate that swamps genuine signal. This is the plasticity-cost problem in its compliance form: constant re-evaluation is not automatically an improvement over periodic review, and a screening team drowning in re-triggered alerts will start ignoring them, which is worse than never triggering at all.

Retraining the screening model every quarter against fresh list snapshots achieves the same effect as continuous revision, without the operational overhead of tracking provenance on every belief.

This is the retraining-cycle objection, and in banking compliance it has real force, because periodic refresh against a new list is exactly the current industry practice at many institutions. The trouble is what that quarterly cycle cannot see. A sanctioned entity added and then delisted within a six-week window, inside the quarter, leaves no trace in a snapshot taken at quarter's end — the transactions that cleared against it during the sanction period are invisible to any later audit, because the record of which list version applied at the moment of each transaction was never kept. Provenance is not bureaucratic overhead here; it is the only thing that lets a regulator, or an internal audit, reconstruct what was known when a specific payment cleared. A quarterly refresh gives you a better model in three months. It does not give you an answer, today, to "was this payment screened correctly against the list as it stood on the day it was made."

The discipline that survives contact with both objections is the same discipline named at the outset: not more data, and not more frequent retraining, but revisable belief with attached provenance and decay, applied selectively enough that noise does not overwhelm signal. Continuous intake without that discipline is worse than a quarterly cycle. Continuous intake with it is the only version of the loop that can answer, after the fact, what was believed, on what evidence, and when that evidence stopped being enough.

Continue