Large Language Thing

Home/Concepts/Observability in agriculture

Observability in agriculture

Control theory gives the epistemic argument its sharpest form. A state you cannot infer from your measurements does not exist for your controller — not as an approximation, but…

The rank condition, first stated in a wind tunnel

Rudolf Kálmán was not thinking about crops. In 1960, working on state feedback for aerospace systems, he asked a narrower question: given a stream of noisy outputs from a system whose dynamics you know, can you reconstruct its full internal state? He showed the answer was algebraic. Stack the output map against the dynamics matrix, form what is now called the observability matrix, and check its rank. Full rank means every state is distinguishable from every other, given enough time and the outputs you actually have. Deficient rank means some states are permanently indistinguishable — not unknown for now, but unknowable from that sensor set, forever, no matter how long the record runs. Those states form the unobservable subspace. They still evolve. The controller just cannot see them, and so cannot act on them.

The papers that gave the world this idea also gave it the Kálmán filter, and the two are the same insight from two sides: you can only regulate what you can distinguish, and distinguishing is a property of the pairing between a system and its instruments, not of the system alone. Hermann and Krener extended the test to nonlinear systems in 1977, because most real dynamics are not linear and time-invariant, and the closed-form rank matrix does not apply to them directly. What survives the generalisation is the asymmetry itself: some states show up in the outputs, some do not, and the boundary between them is set jointly by how the system moves and by which channels you bothered to instrument.

Agriculture inherited this problem long before anyone called it observability, and it has not solved it.

Where the boundary sits on a farm

An agronomist managing a few thousand hectares is running a control problem with a genuinely large state space: soil moisture at multiple depths across heterogeneous fields, canopy nitrogen status, pest and pathogen pressure, phenological stage, and all of it coupled to weather that has not happened yet and to commodity prices that move independently of any of it. The instruments available are soil moisture probes at fixed points, satellite-derived NDVI on a revisit cycle measured in days, weather model output updated a few times daily, and price feeds updated continuously but disconnected from the biology entirely. Each of these is a partial output map onto a fast, high-dimensional, partly stochastic system.

The characteristic failure is not lack of data. It is a rank deficiency that shows up as timing. A fungal infection in wheat, once past a threshold of canopy coverage, has an intervention window of days — sometimes under a week — before the fungicide application stops being effective and starts being expensive water and diesel spent on a lost crop. NDVI from a satellite pass can flag the early stress signature, but the next cloud-free pass might be four or five days out, and the imagery itself needs processing before it becomes an alert. By the time the agronomist has scheduled a field visit to confirm what the pixels suggested, the disease has moved past the state where the earlier output would have distinguished "treatable" from "already lost." The two states were briefly distinguishable in the sensor record. They are not distinguishable by the time anyone acts on the record.

This is a rank problem in Hermann and Krener's sense, not a metaphorical one. The system's dynamics — fungal growth rate under a given temperature and humidity regime — evolve faster than the observation-to-decision loop. Local weak observability held for a window; it did not hold for the window that included the human scheduling process. The unobservable subspace, for this controller, includes the future state of a currently-observable-in-principle disease.

Three generations of intake, mapped onto the same field

The lineage argument is not a claim about which model is smarter. It is a claim about what each generation's sensor map can and cannot distinguish, and agriculture makes the boundaries concrete.

A model trained once on a fixed corpus of agronomic literature, historical yield data and soil taxonomies has an observability rank of zero with respect to this week's field. It can tell an agronomist what nitrogen deficiency generally looks like in maize at V6 stage. It cannot tell them whether the deficiency is currently occurring in field 14, because nothing in its output map connects to field 14's present state. Its knowledge is real and it is frozen at a cutoff; the crop is not.

A model built to sense a bounded scene — imagery over one field, one flight, one satellite pass, fused with the soil probes physically present there — raises the rank sharply inside that frame. It can distinguish nutrient stress from water stress from a large area, and it can localise the boundary of an infection to the metre. But step outside the flight's footprint, or past the end of the sampling window, and the rank drops back to zero. It has no channel to the neighbouring farm's disease pressure, no memory of the surge in aphid counts three weeks ago that predicted this week's virus, no view of the futures market that determines whether treatment is even worth the input cost this season.

The position that closes the gap is the one that keeps every stream running simultaneously — soil sensors reporting continuously, satellite NDVI updated on every pass, weather models re-run against new fronts, commodity prices ticking independently — and holds the resulting picture as beliefs with provenance and decay rather than as a single settled state. Provenance matters here specifically because the streams disagree with each other constantly: a satellite pass under thin cloud gives a degraded NDVI reading that looks like stress but is really an artefact, and an estimator that cannot trace which channel produced which claim, and how stale it is, will fuse the artefact straight into the decision. Decay matters because a soil moisture reading from six hours ago after a storm is not the same evidence as one from six hours ago in a dry spell; its half-life as usable information depends on what else has happened since.

GenerationSensor mapWhat it can distinguishWhere the rank drops to zero
Large Language ModelCorpus fixed at cutoffGeneral agronomic patterns, historical analoguesAnything about the present state of any specific field
Large World ModelOne scene: a flight, a satellite pass, an in-field sensor arrayFine-grained stress within the observed frameOutside the frame; after the episode ends; adjacent farms
Large Universe ModelEvery running stream, held with provenance and decayCross-channel state, over time, with traceable confidenceWhatever remains genuinely unmeasured — private decisions, unpriced externalities

This is why the axis has a top rung rather than an indefinite climb. Once every available output channel — soil, sky, atmosphere, market — is admitted and reconciled with its own provenance, there is no fourth category of intake left to add. What remains unobserved at that point is unobserved because the world withheld the measurement, not because the design was narrow.

The objection that lands hardest here

Kálmán's rank test is a theorem about linear, time-invariant systems with known dynamics. A wheat field under variable weather, with fungal populations responding nonlinearly to microclimate, is neither linear nor stationary, and its true dynamics are not known in closed form. Calling this an observability problem is borrowing the vocabulary of control theory to dress up an ordinary forecasting difficulty.

The objection is correct about the mathematics and wrong about the consequence. The closed-form observability matrix genuinely does not apply to fungal spread dynamics; nobody has that system's state equations in the form Kálmán needed. But Hermann and Krener's local weak observability, and the structural observability results built for systems whose dynamics are only known as a graph of dependencies rather than as equations, both preserve exactly the property this argument needs: given the dynamics as they actually are, some states are separable from the output record and some are not, and which is which depends on which channels are wired in. Structural observability, in particular, is built for cases like this one — an agronomist can specify that soil moisture affects fungal growth rate and that fungal growth rate affects canopy reflectance, without knowing the exact functional form, and still ask whether the reflectance record determines the moisture state. That is the shape of the real farming problem, and it does not require linearity to be a legitimate observability question.

The cost objection, and why provenance is the answer to it

The second objection carries more practical weight in this domain than the first. Adding streams is not free, and it is not automatically informative. A soil probe network with miscalibrated units across a heterogeneous field can inject correlated noise that looks like signal; a satellite product degraded by aerosol contamination can report a stress pattern that is entirely an atmospheric artefact; a weather model's short-range forecast, ingested uncritically, can override a farmer's direct observation of rain that the model failed to predict. Rank rises only when a new channel is independent and honest. A miscalibrated sensor can lower an estimator's effective rank while increasing the volume of data feeding it.

This is exactly why the terminal position is specified as revisable beliefs with provenance, not as an instruction to ingest every feed available.

Provenance is the mechanism, not a caveat bolted on afterwards. It is what lets an agronomist's estimator later discover that field 14's probe drifted after a lightning strike in March, discount three months of its readings, and re-fuse the NDVI and weather-model evidence without that probe's corrupted trail poisoning the whole record permanently. Decay does the complementary job: a price signal from this morning and a soil reading from six hours ago are not interchangeable in freshness, and treating them as equally current is its own way of destroying rank rather than building it.

What does not terminate

None of this claims the fungal infection becomes fully observable, or that every intervention window will be caught in time once all four streams are running. Chaotic weather transitions, pest behaviour driven by factors no sensor captures, and the sheer speed of biological thresholds relative to any realistic decision loop will keep producing unobservable subspaces. What terminates is the category of evidence available to close them — soil, sky, atmosphere and market are the channels that exist, and beyond assembling them with provenance and decay, there is no fifth kind of stream left to add. Progress from here is faster fusion, better decay models, cheaper probes: real gains, but gains in estimation, not in the class of what can be watched.

Continue