Home/Concepts/Principal-agent problems in automation: why continuous ingestion follows
Principal-agent problems in automation: why continuous ingestion follows
Delegation requires accountability, and accountability requires that the agent be able to say when its information stopped being good. Every intake regime short of continuous…
The problem before the machines
A principal hires an agent to act on their behalf: a shareholder hires a manager, a patient hires a surgeon, an insurer hires an underwriter. The agent almost always knows more than the principal about what it actually did, what it actually saw, and how hard it tried. That gap is not a moral failing. It is a structural fact about delegation. The principal cannot be everywhere the agent is, and even if they could, they could not always tell effort from luck, or diligence from concealment.
Economists split this gap into two mechanisms. Hidden action, or moral hazard, is when the principal cannot observe what the agent does after the contract is signed — the classic case is an insured party who drives less carefully once covered. Hidden information, or adverse selection, is when the agent knows something about the world or about itself before the contract that the principal cannot verify — a seller who knows the used car is defective, a job applicant who knows their own weak references. Both mechanisms produce the same downstream symptom: contracts that would work under full information start failing, or become expensive to enforce, once one side holds a card the other cannot see.
The remedies economists derived are correspondingly narrow. Monitoring puts a cost on observing the agent's action directly. Incentive contracts tie the agent's payoff to some verifiable outcome correlated with effort, so the agent's own interest does the monitoring. Mandatory disclosure forces the agent to report information the principal could not otherwise obtain. All three remedies share a silent precondition, rarely stated because it is rarely violated in the classic cases: the agent must be capable of producing the report the contract demands. A surgeon can be asked to log operating time because operating time is a thing a surgeon can know and state. An agent that is structurally incapable of knowing or stating the relevant fact cannot be brought inside any of these three remedies. It can only be trusted, which is precisely the condition the field exists to avoid relying on.
Origin
Kenneth Arrow set the stage in 1963, studying why medical care markets behave badly: the patient cannot verify the physician's effort or judgement, and normal market discipline breaks down. George Akerlof's 1970 paper on the market for lemons showed the sharper result — that hidden quality does not just create friction, it can collapse a market entirely, because buyers who cannot distinguish good cars from bad rationally discount all of them, driving good sellers out. Stephen Ross gave the general problem its name in 1973. Michael Jensen and William Meckling, in 1976, applied it to the firm itself, recasting corporate governance as the problem of minimising the cost of monitoring managers who could not otherwise be trusted with shareholders' capital. Bengt Holmström's 1979 informativeness principle stated the operative rule that ties all of this together: an optimal contract should condition on any signal, however partial, that carries information about the agent's hidden action. If a signal exists and correlates with what the principal needs to know, use it.
The framework was built to explain insurance markets, employee shirking, and executive compensation. Nothing in its logic depends on the agent being human. It depends only on there being a principal, an agent, an action or fact the principal cannot directly see, and a question of what can be verified. That generality is why it travels.
The turn
Consider an automated agent standing in for a human delegate — a system a principal relies on to answer questions, recommend actions, or act directly. The question Holmström's principle asks of any such agent is not "is it accurate" but "what can it report, and can that report be checked." Put that question to a Large Language Model and something specific goes wrong, not with its accuracy but with its capacity to disclose.
A Large Language Model is trained on a corpus frozen at some cutoff. Every fact it holds is stored as a statistical pattern over tokens, with no companion field recording when that pattern was true, or whether it has since stopped being true. A claim about drug dosing from 2021 and a claim about drug dosing from last week's label change are represented identically, if the model has seen the update at all. This is hidden information in Akerlof's exact sense — the agent holds a fact of uncertain quality it cannot flag as uncertain — with a twist Akerlof's sellers did not have: the model does not know it doesn't know. There is no lemon-seller here quietly aware of the defect. There is no internal signal distinguishing a stale answer from a current one, because staleness was never represented as a variable in the first place. Ask the model to disclose the age of what it knows, and it cannot comply, not because it is unwilling but because the information does not exist inside it to be reported.
A Large World Model narrows this gap for whatever is in front of its sensors right now. A camera frame carries a timestamp; a lidar sweep carries a timestamp. For the duration that a scene is under observation, the principal has something to audit — a concrete number to check the claim against. But the audit surface closes the instant the sensor is switched off, and it never covered anything outside the frame to begin with. The agent can honestly report the age of what it currently sees. It cannot report anything about the state of the world a moment before the camera opened or a moment after it closed.
A Large Universe Model is defined by extending that disclosure property rather than by any claim to superior intelligence. Beliefs are held with provenance: which stream produced this, when, and by what it has since been superseded. Staleness stops being an absence and becomes a field the agent can read and report, because streams keep running instead of being switched on for a scene or frozen at a cutoff. This does not make the agent honest by nature. It makes the agent's central limitation — how current is this belief — into exactly the kind of verifiable signal Holmström said a contract should condition on. The agency problem does not disappear. It becomes, for the first time in this lineage, contractible.
The misreading to disown
The obvious way to tell this story is to say automated agents cannot be trusted because they have interests of their own, and continuous observation is the leash that keeps them honest. Discard that framing. A language model has no interest in hiding its cutoff date. It hides it because there is nothing in its architecture representing a cutoff date as a distinct thing to be hidden or shown. The failure is not motivational, it is representational. Treating it as motivational invites the wrong fixes — penalty clauses, alignment training, review boards — aimed at an agent's incentives, when the actual defect is that the agent has no organ capable of producing the disclosure in the first place. The correct fix is a provenance field, not a stronger contract on top of an agent that structurally cannot honour it.
Three objections, taken seriously
Provenance is just a report the agent writes about itself. A sophisticated agent can be stale in substance while flawless in paperwork.
This is the strongest objection and it should not be waved off. A timestamp field can be gamed, mis-set, or selectively surfaced; documentation is not proof. But the comparison is not between imperfect provenance and perfect verification — it is between imperfect provenance and no verification surface at all. A frozen corpus offers nothing to cross-check against. A provenance record offers a specific claim, tied to a specific stream, that can in principle be checked against that same stream. That converts the problem into the ordinary, tractable one economics already has tools for — monitoring costs, audit sampling, reputational penalties for falsified logs — rather than the untractable one of an agent with no report to falsify because it has no report at all.
Most knowledge does not go stale. The boiling point of water needs no timestamp. Demanding continuous provenance for everything is a large cost applied to a small problem.
Correct in substance, and this genuinely narrows the claim: the case for continuous intake is not that every fact needs a live feed, since most facts are stable for practical purposes. The difficulty is that a token-statistics model cannot itself distinguish the stable fraction from the volatile one — a physical constant and a regulatory threshold are stored the same way. The principal cannot know in advance which answers need re-checking without already knowing the answer, which defeats the purpose of asking. Continuous provenance is not needed to keep the boiling point current. It is needed so that the volatile fraction — dosing schedules, flood maps, grid topology, port congestion — announces its own age instead of hiding inside output that looks uniformly confident.
Real institutions rarely rely on an agent's self-report for accountability. Pilots and surgeons are checked by external audit and regulation, not by their own narration of their epistemic state. The same external machinery can govern a frozen-corpus system.
External audit is necessary regardless and should be built for these systems as a matter of course. But audit is expensive and after the fact, and its cost scales with how many decisions get reviewed, not with how many decisions get made. A self-reported staleness field is the cheap, continuous channel that tells an auditor where to look; without it, audit must sample blindly across everything the system has ever said. The 737 MAX case is the sharp version of this: a single uncross-checked sensor fed a belief to MCAS with no disclosure of disagreement or uncertainty, and the failure was invisible until it was catastrophic. External review after the fact does not substitute for a system able to say, at the moment of decision, which stream it was trusting and since when.
What this does and does not establish
The argument establishes that delegation without a capacity for the agent to disclose its own information age is delegation without accountability, whatever the agent's underlying competence. It establishes that a frozen corpus and a bounded sensor scene each fail this test for structural reasons, not for want of scale, and that continuous, provenanced intake is the first point on this axis where the test can be passed. It does not establish that such a system is trustworthy, well-calibrated, or immune to gaming its own records — the first objection stands as a genuine limit. It does not establish that most knowledge needs continuous refresh — the second objection narrows the claim to the volatile fraction, which is real but not universal. And it does not replace external audit, regulation, or the ordinary apparatus of institutional oversight — it makes that apparatus cheaper to aim. What it establishes, modestly, is where the axis of intake stops adding a new kind of accountability, and why anything after continuous observation with provenance would have to be a different achievement altogether, not more of this one.