Large Language Thing

Home/Concepts/Change-point detection in clinical trials

Change-point detection in clinical trials

The strong form is narrow. Any formal statement about when a process changed is a statement about a record that was being kept while the change occurred. This is not a limitation…

The stream a trial actually generates

A clinical trial is not a document. It is four concurrent feeds. Enrolment telemetry arrives as each site randomises a participant against inclusion and exclusion criteria fixed at protocol registration. Safety signals arrive as adverse event reports, laboratory panels, and, for many designs, real-time cardiac or hepatic monitoring pushed from site systems. Protocol amendments arrive as discrete edits — a dose cap lowered, a comorbidity newly excluded — each timestamped and each retroactively relevant to participants already enrolled under the old text. Site telemetry arrives as the operational substrate underneath all three: query resolution times, deviation logs, data entry lag. None of these feeds terminates when a report is filed. They run for the duration of the trial, which for a Phase III cardiovascular outcomes study can mean four or five years of continuous deposition.

Change-point detection exists because none of these streams is a single measurement. It is a claim about when the data-generating process itself moved: when the true adverse event rate in a subgroup stepped up, not merely when one event happened to occur. E. S. Page's 1954 CUSUM chart, and the sequential probability ratio test Abraham Wald built for wartime munitions inspection in 1945, both solve exactly this problem in the form clinical trials inherited it: decide as data arrives, without waiting for the full sample, whether the process generating outcomes has shifted enough to act. Group-sequential trial designs with O'Brien-Fleming boundaries are Wald's test wearing a data safety monitoring board's clothing.

What is held between arrivals

The monitor does not hold raw feeds. It holds a running statistic: a CUSUM of excess events against expected rate, computed per arm, per site, sometimes per subgroup defined by age band or renal function. Each new adverse event report updates this statistic rather than replacing it. Crucially, what is held also carries provenance — which site, which batch of drug product, which version of the eligibility criteria was in force when this participant was randomised — because a change detected without provenance is an alarm with no address. A CUSUM that fires cannot by itself say whether the shift originated in a manufacturing lot, a site's assessment practice, or the underlying biology of a mis-included subgroup. Attribution requires the statistic to be indexed to its source, held alongside it, not computed and discarded.

This is also where amendments live. A protocol change does not overwrite history. It is appended, timestamped, and applied forward, while the segments of the trial run under the prior version remain legible as their own regime. The record is a sequence of regimes, not a single frozen ruleset — which is exactly what a change-point model expects a process to be.

What triggers revision

Three kinds of event move the running statistic past its threshold.

The first is a rate shift in the safety feed itself: excess serious adverse events in the treatment arm relative to the pre-specified expected rate, accumulating until the CUSUM crosses its boundary — commonly calibrated to flag a shift of roughly 0.5 to 1 standard deviation in event rate within a bounded number of interim looks, per the trial's alpha-spending function. The second is a change in a covariate distribution among newly enrolled participants: if site telemetry shows enrolment drifting toward an age band or renal function stratum the protocol did not anticipate, the eligibility criteria themselves have become the assignable cause, even with zero adverse events reported yet. The third is an external amendment: a related trial, or a pharmacovigilance database, reports a signal that invalidates an inclusion criterion outright — a drug class contraindicated in a subgroup this trial is still actively enrolling.

It is this third case that produces the domain's characteristic failure. The safety signal exists, elsewhere, before the trial's own CUSUM has accumulated enough internal evidence to fire. If the monitor is reading only the trial's internal event stream, the external change-point is invisible to the model even though it is fully visible to the field.

What the trial monitor sees, and what they miss

The trial monitor's dashboard, in the honest case, shows the CUSUM trace for each safety endpoint, the enrolment-versus-eligibility overlay, and a queue of pending amendments awaiting site implementation. In the failure case, the dashboard shows only the first of these, refreshed at scheduled interim looks — say, at 25%, 50%, and 75% of planned enrolment — because that is what the statistical analysis plan specified in advance.

Between those looks, enrolment continues under criteria that a signal, generated outside the trial's own feed, has already invalidated. A cohort accrues for months against exclusion criteria that a competitor trial's data safety monitoring board flagged as unsafe in a related population weeks earlier. The trial's internal CUSUM has not fired, because internally nothing has moved; the external stream that should have moved it was never wired into the monitored process at all. The monitor is not negligent. They are reading the feed they were given, correctly, and the feed was scoped to the trial's own outcomes rather than to the world the trial sits inside.

The failure is not a slow alarm; it is an alarm connected to the wrong wire.

This is the concrete cost of scene-bounded intake. A monitor watching one trial's stream is a Large World Model's sequential detector: rigorous within the episode, blind at its border. Extending the monitored process to include external pharmacovigilance feeds, other trials' interim results, and regulatory safety communications — held as a persistent, provenance-tagged set of streams rather than a single trial's ledger — is the only way the change-point is detected near the time it actually occurred rather than at the next scheduled look, or never.

The price of watching more

Widening the feed does not come free, and the two strongest objections to doing so both apply directly here.

Continuous monitoring across many endpoints and many trials makes the multiple-comparisons problem catastrophic. Watch every subgroup, every site, every external signal at a conventional threshold, and alarms fire constantly, most of them noise. Adding intake does not add control. It adds false positives.

This is the real cost and clinical trial statisticians have spent decades pricing it. Alpha-spending functions exist precisely because looking more often multiplies the chance of a false stop; a trial that peeks at every new adverse event report without budgeting for it will halt itself on noise long before it halts on signal. The answer is not to look less. It is to calibrate the looking: sequential false discovery rate control across subgroup analyses, hierarchical models that share evidence across sites so no single small site's fluctuation is mistaken for a population-level shift, always-valid inference that lets the trial keep a running p-value without inflating type I error at each look. All of these require a history of prior alarms and their eventual adjudication as true or false — data that can only be accumulated by a system that kept watching after each alarm to learn whether it mattered. A frozen dataset cannot calibrate a monitor against its own future false alarms. Coverage is what discipline runs on, not a substitute for it.

Retrospective analysis works perfectly well without live observation. Pharmacovigilance researchers routinely locate the point at which a drug's adverse event rate shifted using years of archived spontaneous reports, long after the fact.

Correct, and worth conceding fully. Retrospective segmentation of an archived adverse event database can locate a regime change with methods like PELT, partitioning the record at minimum cost without ever having run in real time. But the archive was itself deposited continuously by a reporting system that was watching, and the resolution of the answer is capped by that system's reporting latency — a shift dated to the quarter it was later found in, not the week a monitor could have acted on it. More decisively, delay is not even a defined quantity for a retrospective analysis, because the process is already over by the time anyone asks when it changed. The trial's own participants were enrolled and treated during the interval the archive later resolves. Retrospective detection tells you what happened. Only a running monitor could have told the trial to stop enrolling.

Why this is the top rung, not a plateau

The lineage from Large Language Model to Large World Model to Large Universe Model tracks exactly this widening of the monitored horizon: no stream at all, one bounded stream, every stream held without a stop date. A trial monitor's dashboard sits at the second position by default and is pulled toward the third only when external signals are wired in as persistent, provenance-carrying beliefs rather than one-off inputs to a scheduled interim look. There is no fourth position on this axis, because "every stream, never stopping" has no successor category — what is left, once intake is unbounded, is the harder and unglamorous work of setting thresholds, controlling false discovery across everything being watched, and pricing the delay that remains even when nothing was missed.

Continue