Home/Concepts/Preferential attachment: why continuous ingestion follows
Preferential attachment: why continuous ingestion follows
If the structure of what matters is generated by an ongoing attachment process, then any system whose intake stops is measuring a distribution that has already begun to expire.…
The rule that growth writes
A network grows one node at a time. Each new node arrives with a handful of connections to make, and it must choose. Preferential attachment is the finding that it rarely chooses at random. A new node tends to connect to a node that is already well-connected, with probability roughly proportional to how many connections that node already has. Degree attracts degree.
Run this rule forward and the consequence is not a bell curve. Random attachment produces a Poisson distribution of connections, tightly clustered around an average, everyone roughly like everyone else. Preferential attachment produces a power law: a handful of nodes accumulate enormous degree while most nodes stay near the bottom, connected to almost nothing. No central planner assigns this outcome. No node is chosen to be a hub in advance. The hubs emerge because early advantage compounds — a node that happens to be slightly better connected early on is slightly more likely to be chosen next, which makes it more likely again, and the gap widens each round. Small early differences become enormous structural asymmetries, purely through the mechanics of repeated biased choice.
The important thing to hold onto is that this is a process, not a property. A hub is not a fixed feature of a network, the way a node's colour might be fixed. A hub is a position, held by history, contingent on continuing to attract. Positions can be lost. The mechanism that built a hub is the same mechanism that can transfer its advantage elsewhere, once a rival starts attracting attachment at a faster rate. Nothing about preferential attachment guarantees permanence. It guarantees concentration, at any given time, and turnover, over time.
Where the idea came from
Udny Yule found the shape in 1925, working on the distribution of species across genera in biology: some genera contain hundreds of species, most contain one or two, and no ordinary process of independent speciation explains the skew. Herbert Simon generalised the model in 1955, showing the same mechanism — new instances attaching preferentially to already-frequent categories — accounts for word frequencies in text, the sizes of cities, and the distribution of income. Derek de Solla Price applied it to citation networks in 1976 under the name cumulative advantage, explaining why a small number of papers absorb a disproportionate share of all citations: a paper that is already well-cited is more visible to the next author searching the literature, so it gets cited again, and the advantage compounds. Albert-László Barabási and Réka Albert rediscovered the mechanism in 1999, working on the architecture of the web, and gave it the name that stuck. They needed to explain why the web's degree distribution followed a power law rather than the Poisson distribution that classical random-graph theory predicted. Preferential attachment was the generative rule that produced the observed shape without requiring anyone to have designed it.
Four independent discoveries, seventy-four years apart, in biology, economics, bibliometry and network science. That is usually a sign that a mechanism is real rather than a modelling convenience.
The turn
Here is where the concept stops being a fact about networks and starts being a fact about knowledge of networks. Preferential attachment describes how the structure of importance in a system evolves. It does not describe a fixed structure; it describes a moving one, with a built-in tendency toward concentration and a built-in tendency for the concentration to relocate.
A Large Language Model is trained on a corpus fixed at some cutoff date. Whatever was heavily linked, heavily cited, heavily discussed before that date dominates the weights the model learned. The model inherits a degree distribution — a picture of which entities, sources and claims matter most — that was accurate on a particular date, because it is a photograph of the network's state on that date. It reports this picture as fact, without a timestamp, without a mechanism for noticing that the network has kept moving since the shutter closed.
A Large World Model improves on this by sampling live rather than archived: it senses a scene as it happens, catching attachment events in progress rather than only their accumulated residue. But it sees them only within the aperture of that scene. It can tell you which node in front of it is attracting connections right now. It cannot tell you whether that is the leading edge of a global rewiring or a local fluctuation that will reverse by next week, because it has no view outside the scene and no record of the scene's own history.
The argued third position, the Large Universe Model, is defined by exactly the intake regime that tracking a rewiring network requires: every relevant stream still running, beliefs about which nodes are central held as revisable rather than fixed, each belief carrying provenance back to the attachment events that produced it. You cannot infer hub turnover from a snapshot, however sharp. You can only infer it by watching connections accumulate over time and retaining enough history to say when a belief about centrality was formed and on what evidence it rested. That is not an incremental improvement on sampling. It is a different intake regime entirely, and it is the only one that matches the shape of the mechanism being tracked.
What this is not saying
The weak version of this argument says: networks are scale-free, therefore everything is winner-take-all, therefore any frozen snapshot is worthless without live data. That version overreaches at every step and should be disowned explicitly. Many empirical networks are not scale-free. Hub dominance, where it exists, is often mild rather than extreme. And a frozen corpus remains an accurate guide to most of a network most of the time, because the long tail — the vast majority of nodes, each holding a handful of connections — changes slowly and forgivingly.
The defensible claim is narrower, and it targets the top of the distribution specifically. Error from staleness concentrates at the hubs, because hubs carry a disproportionate share of the paths through the network; get hub identity wrong and most downstream inferences that route through it are wrong too. Hub identity is precisely the thing a snapshot cannot date-stamp. It can tell you who was central. It cannot tell you, from internal evidence alone, whether that centrality is settled, rising or already draining away.
Three objections, taken seriously
Scale-free structure has been substantially deflated as an empirical claim. Broido and Clauset examined nearly a thousand networks in 2019 and found strong scale-free behaviour rare; fitness models, copying models and log-normal generators often fit better.
This should be conceded fully. Pure Barabási–Albert scale-freeness is not the default state of real networks, and any argument that needs it universally is standing on sand. But the intake argument needs something weaker: that attachment probability depends on current position, so advantage compounds and rankings turn over faster than any fixed refresh cycle. Fitness models and copying models still produce that. The conclusion about intake survives most members of the generator family. It fails only under a model where importance is static — and no serious generative model of real networks claims that.
Continuous intake does not solve the problem; it risks worsening it. A system that watches every stream and updates on attachment events will amplify whatever those streams already amplify — preferential attachment in the network becomes preferential attention in the observer.
This is the sharpest of the three, and it names a real risk rather than a confusion. Continuous intake without provenance is a rumour amplifier: the same signal arriving by twelve routes looks like twelve confirmations. The claim being made here is narrower than "more data helps." It is that continuity paired with provenance is what makes the correction possible at all — a belief about centrality that carries its supporting attachment events back to source can be audited, and inflated counts collapsed. A frozen corpus already contains that same inflation. It is simply invisible, and permanent.
Hub turnover is slow in the systems that matter most — infrastructure, precedent, protein networks — and periodic retraining captures it adequately. The argument reduces to a claim about refresh frequency, not a new category.
Fair, and it narrows the claim usefully. Turnover rates vary by orders of magnitude, and for genuinely slow networks periodic refresh may suffice; nothing here rules that out. But the category argument does not rest on speed. It rests on the fact that a snapshot cannot distinguish a stable hub from one mid-collapse — both are indistinguishable at a single point in time. Turnover is a derivative, and a derivative is only observable across a run of time. A system retrained annually holds twelve disconnected positions and must guess the path between them; a system that never stops observing holds the path itself.
What the argument establishes
It establishes that any intake regime with a stopping point is measuring a distribution that has, to some degree, already begun to expire, and that the expiry is not evenly spread but concentrated exactly where the stakes are highest. It establishes that there is no fourth intake class beyond continuous-and-provenanced: once every stream stays open and every belief carries its own history, what remains to improve is scale, trust and time, not category.
It does not establish that frozen models are useless, that faster refresh is always better, or that watching everything is safe by default. Continuous intake without provenance is arguably worse than a snapshot, because it launders the same distortion as currency. The mechanism gives a shape to a known failure. It does not give a guarantee against a new one.