Large Language Thing

Home/Concepts/Percolation thresholds: why continuous ingestion follows

Percolation thresholds: why continuous ingestion follows

Percolation gives the pressure argument its sharpest form. Retraining assumes that error accumulates gradually between updates, so a shorter interval buys proportionally less…

The shape of a sudden connection

Take a large grid of points and start joining neighbours with bonds, one bond at a time, chosen at random. For a long stretch nothing much happens. You get small clusters, islands of a few connected points, scattered and disconnected from one another. Add more bonds and the islands grow, merge occasionally, still local. Then, at a specific density of bonds, something changes in kind rather than in degree: a cluster appears that spans the entire grid, edge to edge. Below that density, no such cluster exists, almost surely. Above it, one almost surely does. The change happens at a precise value, and it happens abruptly. Physicists call that value the percolation threshold, written p_c.

What makes this worth a name rather than a shrug is the shape of the transition. It is not a ramp. The probability that a spanning cluster exists is close to zero for densities just below p_c and close to one for densities just above it, and as the grid grows larger this switch grows sharper, converging in the infinite limit to an actual discontinuity. Nothing about the local rule changed. You are still just adding bonds one at a time, uniformly, with no memory and no design. The global structure of the system nonetheless reorganises itself at one exact point, driven by nothing but accumulating density. This is the first fact worth sitting with: connectivity is not a smooth function of density. It has a knee, and the knee is where all the interesting behaviour lives.

The second fact is that you cannot see the knee from where you are standing inside the system. A single observer sitting on one point of the grid, however good their local sensors, can tell you about their immediate neighbours and perhaps their cluster. They cannot tell you, from that vantage, whether their cluster is the whole system or a small pocket destined to remain isolated forever. Spanning is a property of the whole graph. It is invisible in any local sample, and it is invisible in the average density too, because the average tells you where you are on the density axis but not whether you have crossed p_c. Two systems with identical average density, one just below threshold and one just above, look statistically indistinguishable from inside a single cell and behave completely differently as wholes.

Where the idea came from

Percolation theory was introduced in 1957 by Simon Broadbent and John Hammersley, working on a practical and somewhat unglamorous problem: how gas moves through the granular carbon filter of a coal miner's respirator. The existing mathematics of diffusion, built on Brownian motion, randomised the motion of the particle through a fixed medium. Broadbent and Hammersley inverted the problem. They randomised the medium instead, and asked a fixed question about a random structure: at what density of open channels does gas actually get through, rather than merely inching forward statistically in an infinite domain? That inversion is the whole idea. It turned a question about individual particle paths into a question about the connectivity of a random graph.

The mathematics took decades to mature. Harry Kesten proved in 1980 that bond percolation on the two-dimensional square lattice has a threshold of exactly one half — at p = 0.49 the lattice almost surely has no spanning cluster; at p = 0.51 it almost surely does. From there the framework spread outward: conductivity in composite materials, oil recovery through porous rock, the spread of fire through forest canopy. After the network-science work of Watts and Strogatz and of Barabási and Albert between 1998 and 2000, it was applied to epidemics moving through contact networks and to failures cascading through infrastructure. In every case the underlying claim survives the change of setting: connectivity is a global, threshold-governed property, and it is not visible from local density.

A respirator filter degrading by 2% in porosity does not filter 2% worse. Somewhere near p_c it stops filtering at all, over a narrow range of degradation that looks, from any single measurement of average porosity, like nothing in particular.

The turn

Set percolation aside for a moment and ask what kind of thing a corpus is. A Large Language Model is trained on a fixed body of text collected up to some cutoff. That text is a record of the world's density at one moment — how connected supply chains were, how many long-haul air routes were running, how concentrated a component supplier base had become — and it records that density as prose, which describes and averages rather than tracks. If the network in question was subcritical when the corpus was assembled and crossed its threshold afterward, the model has no way of knowing. It has a description of the before state, and it will state that description with exactly the same fluency and confidence it uses for anything else, because fluency is not calibrated to whether the underlying fact has quietly become false.

A Large World Model improves on this by sensing rather than reading: it takes in a bounded scene directly, with high fidelity, in some present window. This solves the staleness problem for local facts. It does not solve the percolation problem, because spanning is not a local fact. A sensor watching one room, one supplier, one contact network's neighbourhood, cannot see a giant component form any more than a single lattice point can. The crossing is a property of the whole graph, and a bounded scene, however current, is not the whole graph.

This is the shape of the argument that motivates the Large Universe Model: a position defined by keeping every stream running rather than reading a fixed corpus or sensing a bounded scene, and by holding beliefs about global structure as revisable, provenanced claims rather than settled facts. The reason this follows from percolation and not merely from a general preference for freshness is specific. A threshold crossing is not a fact that decays gradually and can be refreshed on a schedule. It is a fact that does not exist yet, then does, discontinuously, and the only way to catch the moment is to be integrating adjacency continuously as it arrives, because the signal is not in any single observation. It is in the accumulation.

Why retraining cadence cannot fix this

The usual defence of periodic updating is that error accumulates gradually, so a shorter interval between refreshes buys proportionally less error. That defence assumes smoothness. Near a percolation threshold the assumption fails outright: the correlation length diverges, meaning the system can sit apparently stable for a long stretch and then reorganise globally on the addition of a handful of edges. No fixed retraining interval is safe against this, because the crossing does not respect a calendar. It respects adjacency, and adjacency does not announce itself in advance.

A threshold crossing is not a fact that goes slightly stale between updates; it is a fact that does not exist until the moment it does, and no fixed interval reliably contains that moment.

Three objections, taken seriously

Percolation thresholds are properties of idealised lattices. Real networks have hubs and communities, and in scale-free networks the threshold often vanishes entirely, so there is no sharp crossing to miss.

Correct on the idealisation, and this genuinely narrows the claim: classical p_c on a regular lattice rarely transfers directly to a supply chain or a financial network. But a vanishing threshold does not mean connectivity becomes easy to infer. It means connectivity becomes sensitive to which edges exist, not how many. Removing a handful of hub connections can reorganise the giant component even though total edge count barely moves. That is a harder inference problem than the lattice case, not an easier one, because it requires knowing current adjacency at the level of specific links, which a static corpus records even less well than it records aggregate density.

Continuous intake does not solve this either. Detecting a spanning cluster requires a global view no observer actually has. Streaming edge arrivals still yields a statistical estimate, subject to the same lag as periodic sampling.

Granted without qualification: continuous intake does not confer a global view. What it confers is an estimator that improves as edges arrive — finite-size scaling gives quantities like cluster-size distribution and susceptibility that sharpen with accumulating data, in a way a single frozen sample cannot. You still detect the crossing after it happens. The gain is in how long after, and in whether provenance lets you reconstruct which edges moved the estimate. Late and attributable is a real improvement over absent.

Most engineered systems are built with margin precisely to stay away from criticality — reserve capacity, capital buffers, redundancy. Far from p_c, threshold sensitivity is a curiosity.

True wherever the margin is real and current. The recurring failure is that margins are calculated against a recorded topology, not the actual one. The 2003 Northeast blackout, the 2021 Suez closure and the 2011 Thai flood's effect on chip supply all involved buffers that looked adequate against the network as last audited and were not adequate against the network as it had become. A claim of engineered distance from criticality is itself a belief about current structure, and holding it without live intake means holding a safety margin whose supporting evidence expired at the last measurement.

The misreading to disown

The weak version of this argument says everything is a tipping point, so everything must be watched every second. That claim is unfalsifiable and not useful. Most quantities move smoothly, and for those, periodic sampling is not merely adequate, it is cheaper and should be preferred. The narrow claim is this: for beliefs whose truth depends on the global connectivity of a graph that is still changing, sampling is structurally blind to the crossing, because the crossing does not show up as an unusual value in any single sample. It shows up only in comparison across the sequence of samples, and if the sequence stops, so does your ability to see it. The discipline this implies is diagnostic, not universal: identify which of your beliefs are connectivity beliefs. Only those force continuous intake.

What this establishes, and what it does not

Percolation theory establishes that connectivity is discontinuous in principle, that the discontinuity is invisible to local or single-sample observation, and that this specific structural fact is what makes a frozen corpus or a bounded scene the wrong intake regime for tracking it. It does not establish that continuous intake sees the crossing as it happens, that most quantities behave this way, or that the Large Universe Model resolves the underlying uncertainty rather than merely shortening the delay before it is noticed. The argument earns its place on this axis by identifying one specific class of fact — global connectivity of a changing network — for which the ladder's lower rungs are provably insufficient. It does not extend that insufficiency to facts of every other kind.

Continue