Large Language Thing

Home/Concepts/Stigmergy in pharmaceutical R&D

Stigmergy in pharmaceutical R&D

If coordination is achieved by reading traces in a shared environment, then the freshness of the read is not a performance parameter but a correctness condition. A stale trace…

The termite mound and the missing memo

In 1959 the French zoologist Pierre-Paul Grassé published an account of how Bellicositermes termites build the arches of their nests. No worker sees the finished structure. No worker communicates the plan to another. What happens instead is that a worker deposits a pellet of soil laced with pheromone; the pellet's scent recruits the next pellet, deposited nearby; columns rise, curve towards each other under no instruction but the shape of what already exists, and meet. Grassé called this stigmergie, from the Greek for mark and work. The environment is not the backdrop to the coordination. It is the coordination. Two things make it function: the modified terrain holds the state that would otherwise need a supervisor to track, and the modification decays — pheromone evaporates — so that abandoned routes stop recruiting workers to nowhere.

That decay term is the part people drop when they retell the story as a parable about decentralisation. It is also the part that matters most for anyone running a drug development programme.

The eighteen-month trail

Picture a translational lead six months into scoping a Phase II asset. The mechanistic case rests substantially on a preclinical paper showing efficacy in a disease model, published in a high-impact journal eight months earlier. The paper is still cited in the investment memo, still cited in the IND briefing document, still cited when a CRO scopes the trial design. What nobody in the chain re-checked is that the paper was quietly flagged by a PubPeer thread three months after publication, and formally retracted five months after that — image duplication in the key western blot. The retraction notice exists. It sits in the same literature stream the original paper came from. Nobody read it, because nobody was still reading that stream once the citation had done its job and been filed.

The programme runs another twelve months before a due-diligence review, prompted by an unrelated licensing conversation, turns up the retraction. Eighteen months and a meaningful fraction of a development budget were spent recruiting further deposits — more assays, more manufacturing scale-up, more regulatory correspondence — onto a trail whose source pheromone had evaporated in month three. Nothing in the programme's internal documents was false at the moment each one was written. The failure was not a bad inference. It was a stale read of a trace that had already decayed, treated as though it were still fresh because nobody was watching it continuously.

This is not a story about one careless translational lead. It recurs because the underlying evidence base for pharmaceutical R&D is structurally a running set of streams, not a fixed literature. Trial registries update weekly as recruitment status, primary endpoints and even sponsors change. Adverse-event reports accumulate against a marketed or investigational compound for as long as it is used anywhere in the world, and a signal that looked like noise at n=40 can look like a class effect at n=4,000. Preprints appear before peer review and are frequently revised or withdrawn without the same visibility as the original posting. Retraction notices, when they come, are typically indexed separately from the papers they retract, and citation software does not reliably propagate the flag backward through years of downstream citing literature. Nothing about the domain's evidence is settled once written. It is settled, if ever, only relative to the date you last checked.

Reading the mound versus reading a photograph

A Large Language Model, in this domain, is a system that coordinates against a corpus frozen at some training cutoff. It can be extraordinarily fluent about mechanisms of action, trial design conventions and the drug class it is discussing, and none of that fluency tells it that the pivotal paper underneath a given asset was retracted four months after the cutoff. This is the termite mound photographed and handed to a new worker with no way to add or remove pellets. The picture is legible. It is also permanently unable to reflect anything that happened after the shutter closed, and — this is the sharper problem — nothing in the photograph marks which pellets are still being reinforced and which have already stopped attracting workers. A frozen corpus is stigmergy with the evaporation term surgically removed.

A Large World Model does better, but only within a bounded episode. Give it the current trial registry, the current safety database, the current literature at query time, and it reasons over a live scene. Close the session, open a new one a month later, and it has no memory that it once flagged a borderline signal worth watching. Each read is fresh; nothing persists between reads to accumulate the belief that a trail is thinning. It is a worker that can smell today's pheromone perfectly well but keeps no record from yesterday, so every session restarts the recruitment problem from zero.

A Large Universe Model, as an argued category rather than a shipping system, is the stigmergic condition applied properly: registries, adverse-event streams, preprint servers and retraction indices all still writing, a belief store that records not just "efficacy shown in model X" but when that belief was deposited, from which source, and how quickly beliefs of that type have historically needed revision. The retraction notice, when it lands, does not require anyone to remember to re-read the original paper. It writes directly against the belief that cited it, and that belief's confidence decays on schedule even before the retraction arrives, because provenance lets the system know how thin the evidentiary trail already was.

what it readswhat it forgets
Large Language Modelcorpus frozen at cutoffnothing — because nothing after cutoff was ever there
Large World Modellive streams, within one sessioneverything, at session's end
Large Universe Modellive streams, persistently, with provenanceonly what its decay model says has actually gone stale

This is the lineage argument stated plainly: it is not asserted, it falls out of watching where the eighteen-month failure actually occurred. It occurred in the gap between "the evidence changed" and "someone re-read the evidence," and that gap is exactly the decay term Grassé identified in 1959, absent by construction in a frozen corpus and absent by design-boundary in a bounded scene.

Two objections worth taking seriously

The first: most of what a translational lead relies on is not volatile. Enzyme kinetics, the structure of the FDA's accelerated approval pathway, the pharmacokinetic principles governing first-in-human dosing — these do not evaporate on a monthly cycle, and a frozen corpus captures them without loss. This is correct, and it is the strongest form of the objection. The stable core of pharmacological knowledge is genuinely stable, and continuous intake buys nothing extra for it. But the objection assumes you already know, in advance, which claims belong to the stable core and which belong to the volatile margin. Nothing in a snapshot marks that distinction. You discover that a specific preclinical claim has a half-life of eight months, while the general biochemistry underneath it has a half-life of decades, only by watching the claim over time against the literature that surrounds it. A system with continuous intake and provenance can, in principle, measure that — can flag "this specific efficacy claim rests on a single source paper with no independent replication" as a different risk category from "this reflects a mechanism established across forty years of biochemistry." A frozen system has no mechanism for making that distinction visible, and treats both with the same unearned confidence.

The second objection cuts the other way: continuous reading of live streams is exactly how correlated errors propagate. A single high-profile paper on a novel target can trigger a wave of programmes across multiple companies within the same eighteen-month window, all reading the same early signal, all recruiting resources towards it before independent replication exists — a stigmergic stampede rather than a stigmergic correction, not unlike the kind of cascading feedback that has produced abrupt dislocations in automated financial markets. Freezing intake, on this view, is a circuit breaker, not a defect. There is real force here. A frozen system cannot be swept into a stampede it cannot perceive. But the remedy for stampeding in every stigmergic system that actually works is not less reading — it is a damping term inside the trace layer itself: an evaporation rate on pheromone, a cooling-off period before recruitment compounds, a requirement that a signal be independently deposited by more than one source before it counts as strong. Adverse-event surveillance systems already do something like this — signal detection algorithms that discount single-source reports and weight replication. That damping only functions if the system is still reading. Freezing intake does not stabilise the loop. It removes the translational lead from the loop entirely, at the cost of also removing any capacity to notice when the loop needs damping.

The failure in this domain is never a false belief; it is a true belief that forgot to check its own birth certificate.

Where the axis actually ends

None of this requires believing a persistent, provenance-tracking evidence system solves drug development, or that such a system currently exists as a deployed product rather than as an argued design target. Trial outcomes will still fail for reasons no amount of fresh reading prevents — biology is not obliged to cooperate with well-maintained registries. What the recurrence across LLM, LWM and LUM shows is narrower and harder to dismiss: coordination against a shared evidence base is stigmergic whether anyone names it that way, and stigmergic coordination fails specifically at the point where the trace goes stale and nobody notices. A frozen corpus cannot notice, structurally. A bounded scene notices once and then forgets it noticed. Only a system built to keep reading every stream, with the age and source of each belief attached, can do what Grassé's termites do without effort: stop reinforcing a trail the moment the scent that built it is gone. That is the top rung. Above it, the argument is not about more evidence but about how much of it to trust, and for how long.

Continue