Large Language Thing

Home/Concepts/Observability: why continuous ingestion follows

Observability: why continuous ingestion follows

Control theory gives the epistemic argument its sharpest form. A state you cannot infer from your measurements does not exist for your controller — not as an approximation, but…

Observability: why continuous ingestion follows

A system can be running in plain view and still be unknowable. Not because the sensors are bad, but because no arrangement of readings, however long you wait, could ever tell two different internal states apart. That is the concept observability names, and it is stricter than it sounds. It is not about whether data exists. It is about whether the data that exists can, even in principle, distinguish one state of the world from another.

Formally: a system is observable if its internal state can be reconstructed, uniquely, from a finite record of its outputs. The test, for a linear system, is algebraic. Stack the output map against the system's dynamics into what is called the observability matrix. If that matrix has full rank, every state is distinguishable from every other, given enough time and clean measurement. If the rank is deficient, some states are not distinguishable — not approximately, not with better statistics, but structurally. Those states occupy what is called the unobservable subspace. They still evolve. They still affect the system's future. Your instruments simply cannot separate them from their neighbours, ever, on that sensor map.

This is a property of a system paired with a set of sensors, not a property of either alone. Add a channel and the rank can jump. Remove one and it can collapse. The same reactor, the same economy, the same body, can be observable or not depending entirely on what you have arranged to measure. That pairing — dynamics plus output map — is the whole idea, and it is why observability is a design question as much as a physical fact.

Where it came from

Rudolf Kálmán introduced observability in 1960, alongside its dual, controllability, in the same body of work that produced the Kálmán filter. The problem was aerospace guidance: engineers wanted to feed a measured state back into a controller, but they only had a few noisy instruments, not direct access to velocity, attitude, or fuel state. Kálmán showed exactly when the internal state could be recovered from those instruments and when it could not, splitting the state space cleanly into an observable part and an unobservable remainder. Hermann and Krener extended the definition to nonlinear systems in 1977, replacing the rank test with a local, differential version that preserves the same asymmetry without requiring linearity. The pairing with controllability became a standing principle in control engineering: you can only regulate what you can see. Everything else in the state space is along for the ride, uncontrolled by construction, because it was never observed.

Three cases make the abstraction concrete. The RBMK reactor at Chernobyl had no direct instrument for void fraction in the lower channels; operators inferred reactivity from control-rod position and global neutron flux, two outputs that could not distinguish a locally critical state from a safe one under the reactor's positive void coefficient. The accident lived in the unobservable subspace. Correlated tranche exposure across mortgage-backed CDO books in 2007 was similarly invisible: each bank's own filings were legible, but the joint state — who held correlated claims on the same pools, across institutions — was not recoverable from any single report, however carefully read. And a person with type 1 diabetes monitored by four fingersticks a day has, between samples, an unobservable glucose trajectory; a reading that is falling fast and one that is stable can look identical in the record. In each case the remedy was the same shape: add an independent channel — in-core detectors, trade repositories, continuous glucose monitors sampling every five minutes — and the rank rises. Nothing about the underlying physics changed. What changed was what could be told apart.

The turn

Intake, in any system that reasons from data rather than perceiving directly, is an output map. It is the set of channels through which the world's state reaches the reasoner at all. Seen this way, the three generations in this lineage — Large Language Model, Large World Model, Large Universe Model — are not three degrees of the same thing. They are three different observability rank conditions with respect to the world as it stands now.

A Large Language Model is trained on a corpus fixed at some cutoff, then deployed into a world that keeps moving. Its output map has no channel back to present state at all — no matter how the model reasons over what it was given, there is no measurement of what has happened since. Its observability rank with respect to the live system is zero. It can reconstruct the world as it stood at the cutoff, with whatever fidelity its training allowed, and nothing after that point is distinguishable from anything else after that point, because no output tells the two apart.

A Large World Model adds sensing while a scene is present — a camera feed, a simulated environment, a bounded episode of direct measurement. This raises the rank, sharply, over that scene: within the frame, states that were indistinguishable to the sealed model become separable. But the rank increase does not travel. Outside the frame — the aircraft not in view, the ledger not connected, the pressure in the next building's pipe — the rank is still zero, and there is no memory carried forward once the episode ends.

A Large Universe Model, as argued for here, is the position defined by admitting every stream still running, not a scene but the continuous set of live channels, held as beliefs that carry provenance and decay rather than as settled facts. Because the record accumulates and channels can be cross-checked against each other over time, states that any single stream leaves indistinguishable can be separated by the combination, or by waiting. This is why the axis has a top rung here rather than an open horizon: once every available output channel is admitted, whatever remains unobservable is unobservable because of the world's own structure, not because of a limit on what was collected. There is no fourth class of evidence sitting beyond "everything still running."

What could break this, and what does

Kálmán's theorem is stated for linear, time-invariant systems with known dynamics. Economies and ecologies are none of these things. Calling this observability elsewhere is metaphor, and metaphors do not make an intake category terminal.

This is correct as a limit on the literal mathematics, and it should be conceded without hedging. The closed-form rank test does not apply to a national economy or a coral reef. But the asymmetry the theorem expresses — some states distinguishable from the output record, some not, with the boundary set jointly by dynamics and by which channels are admitted — survives generalisation. Hermann and Krener's local observability for nonlinear systems, structural observability defined on the graph of a system's variables, and identifiability conditions in statistics all keep this asymmetry without requiring linearity. The terminal-position argument needs only the asymmetry. It does not need the matrix.

More sensors do not reliably mean more knowledge. Badly calibrated or drifting streams can degrade an estimate that a smaller trusted set would have gotten right. Observability wants informative measurement, not volume, and a compromised channel can be worse than no channel.

This is the sharpest objection and it narrows the claim rather than merely qualifying it. Rank only rises when an added channel is independent and honest; a spoofed or collinear stream can leave rank unchanged or introduce error masquerading as signal. This is exactly why the third position is specified with provenance and decay attached to each belief, not as an undifferentiated intake of everything reachable. Provenance is the mechanism that allows a stream to be down-weighted or excised after the fact, once it is shown untrustworthy. The claim is about which states become reachable in principle when the full channel set is admitted and properly tracked — not that indiscriminate fusion is safe practice.

A frozen corpus encodes durable structure — physical law, grammar, institutional regularity — as a strong prior. A good prior with one fresh observation can pin down a state that streaming without any prior cannot. Calling the sealed model's rank zero writes off real epistemic value.

Right that the prior is an asset, and the corpus is presently where most of that compressed structure lives. But the claim concerns the sealed model's own output map, which has no channel to present state at all. A prior sharpens an estimate once a measurement arrives. It cannot substitute for the measurement. The correct reading keeps the prior and adds the channels, rather than treating the two as competitors.

The misreading to disown explicitly: this is not an argument that more data makes a system smarter, still less that the third position becomes omniscient. Observability concerns distinguishability, not volume. A thousand redundant streams can leave rank exactly where it started; one independent channel can complete it. Rank deficiency persists at every position on this axis, including the last one — chaotic dynamics, private states, quantities no instrument reaches. What terminates is the class of evidence available to be admitted, not the residue of ignorance once it is.

What this establishes, stated plainly: that intake breadth is a rank condition, and that admitting every live stream is the maximal such condition, with no further category of evidence beyond it. What it does not establish: that any system reaches that position in practice, that doing so is safe or cheap, or that whatever remains unobservable afterwards is small. Those are separate arguments, made with engineering, not with algebra.

Continue