Industry: Healthcare and Epidemiology Product: ModelRisk Application: Clinical trial recruitment
Roughly 80% of clinical trials miss their original recruitment deadline, and every month of slip on a Phase III oncology trial burns somewhere between $600,000 and $1.5 million in fixed site, monitoring and supply costs. The deterministic recruitment plan — multiply 50 sites by 2 patients per site per month and declare a 10-month finish — works on the spreadsheet and fails in the field. The reason is not that the average rate is wrong; it is that the variability across sites and the long tail of site activation are first-order effects that the mean cannot represent.
A global pharmaceutical sponsor planning a 1,000-patient Phase III oncology trial across 50 sites built a probabilistic recruitment model in ModelRisk. The trial team needed not the mean finish date but the distribution — what is the 90th percentile finish month, what is the probability of hitting the 12-month plan, and what mitigation buys the most weeks back. Twenty thousand simulated trials produce the answer in one frame.
The deterministic plan promised completion at month 10. The simulated distribution puts the median finish at month 12 and the P90 at month 16 — and the probability of hitting the 12-month contractual milestone is only 62%. The long right tail is entirely a product of slow sites and slow activations that no single mean rate can express.
The standard mistake is to model trial-level enrolment as a single Poisson process with rate λ equal to (sites × per-site rate). That assumes every site is identical and that the rate is known. Real site enrolment is two-stage uncertainty: each site's true recruitment rate is itself drawn from a distribution, and conditional on that rate, the realized monthly count is Poisson. The compound is a Negative Binomial — the over-dispersed cousin of the Poisson, and the right starting point for pharmaceutical recruitment forecasting.
Site-level rate was modeled as Gamma(2.0, 1.5), giving a mean of 3.0 patients per site per month, a standard deviation of 2.1, and a 90% CI of 0.5–7.2. Two-thirds of sites recruit close to the mean. A handful enrol at 6+ per month and end up contributing 30%+ of the total. A handful never recruit a patient. The Gamma prior is what generates that long right tail of fast sites and the spike of zero-recruiters that every program manager has lived through.
Site activation was modeled as LogNormal with median 60 days and a tail extending past 150 days for sites needing extra ethics-committee submissions. The first-quartile site is active by day 35; the worst-decile site is still not active by day 130.
Screen failure was modeled as Beta(3, 7) — mean 30%, 90% CI 11%–52% — applied as a multiplicative thinning of the consented-to-randomized conversion. This is the correct way to combine a per-event probability with a count process: thin the Poisson rate by the kept fraction, not subtract a count after the fact.
Per simulated trial: draw activation day for each of the 50 sites, draw a true rate per site, then for every month after the site is active, sample a Poisson number of consented patients at rate × (1 − screen-fail).
The cumulative enrolment fan makes the same point on a different axis:
By month 12 the P50 trial has randomized about 1,075 patients — just past target, which is why P(meet by month 12) sits a little above half. The P10 trial has only randomized about 750. The width of that envelope is the contingency the deterministic plan invisibly demands. Sponsor and CRO need to be looking at this fan before the contract is signed, not after the milestone has been missed.
The two biggest movers are how many sites you open and how fast each site actually recruits — a ±10-site swing and a ±0.6 patient/site/month swing each move P(target by month 12) by roughly 30 percentage points either side of the baseline (a full spread of about 60 pp). Screen-failure rate comes next at about ±23 pp half-spread. Site-activation speed is the smallest of the four levers tested, at about ±12 pp half-spread: cutting median activation from 60 to 40 days helps, but far less than adding capacity or lifting per-site rate. The implication: the schedule is dominated by enrolment capacity — sites and their throughput — not by shaving weeks off ethics-committee turnaround.
The team modeled two interventions jointly — staggered activation prioritizing pre-validated sites (median activation 40 days) and adding 10 sites for redundancy (60 total):
The combined mitigation shifts the time-to-N distribution from a P50 of 12 months and a P90 of 16 months to a P50 of about 9 months and a P90 of 13 months — mean finish drops from 12.3 to 9.9 months. P(meet target by month 12) rises from 62% to roughly 89%. The economic case: the additional 10 sites cost about $1.8M upfront, and the saved milestone-slip cost (estimated at $1.2M per month at P90) more than covers it.
In clinical trial recruitment the headline rate is a story about the average site, but the schedule is a story about the worst-quartile site and how much enrolment capacity you actually field. Monte Carlo simulation in ModelRisk is what turns "we plan to finish in 12 months" into "we have a 62% chance of finishing in 12 months — here is the cheapest way to get to 90%."