Operations

Site selection is a data problem

Track record, competition density and the enrolment arithmetic are all measurable from public records — and the most common site-selection mistake is a counting error, not a judgement error.

Clinical Trial OS · · 5 min read

The argument in short

  • A raw trial count at a location measures experience, not contention. They are different numbers and they point different ways.
  • Competition is only the trials that could enrol a patient today — status scoping changes the answer dramatically.
  • The enrolment arithmetic is unforgiving, and the variable people over-tune is the one that matters least.

Site selection is where a feasibility assessment becomes a budget. Get it wrong and you do not find out for nine months, at which point the fix — adding sites — costs more than the original error and buys back less time than anyone hopes. It is also, unusually for clinical operations, a question that public data can answer well, because trial registrations carry locations, statuses, sponsors and dates.

What goes wrong is rarely the judgement. It is the counting.

Three different numbers wearing the same label

Ask how many trials in your condition are running in a given city, and you will typically be handed one number. There are at least three, and they mean opposite things:

What a “trials at this location” count can mean.
CountDefinitionWhat it tells you
All-timeEvery same-condition trial ever registered at this locationSite experience — this centre has done this before
RecentPrimary completion inside the last few yearsCurrency — the team was active lately, not in 2009
ActiveRecruiting, enrolling by invitation, or not yet recruitingContention — who else is chasing your patients

Experience is a reason to choose a site. Contention is a reason to avoid one, or at least to discount its projected rate. Reporting the all-time count and calling it competition inverts the signal: it steers you toward the crowded centres precisely because they are crowded, and it does so with a number that looks rigorous.

One worked example, from the public record

While building this analysis we probed a reference endometriosis study. Of 114 registered phase-2 trials in the condition, seven were in a status that could enrol a patient — and four of the six cities that the naive all-time count flagged as competitive hotspots had none. Same registry, same query, two different pictures. This is a single condition at a single point in time, not a general statistic; it is here because it is checkable, and because the size of the gap is the point.

The other counting error: population

A map of raw trial counts by country is, to a first approximation, a map of population. Large countries have more trials in every condition, so an un-normalised country ranking will put the same handful of countries at the top of every study you ever run, which is a strong hint that it is not measuring what you asked.

Normalising by population turns the count into something comparable — trials per head of population — and it changes the ranking substantially. It also introduces a dependency worth being explicit about: you now need a population denominator, from a named source, with a date. When the denominator cannot be resolved for a country, the honest output is no ratio at all rather than a quietly bundled estimate.

Track record, and the investigator underneath it

A site is an address; an investigator is a person with a history. The public record supports more of this than most teams use. Registry records name principal investigators and their institutions. Publication records show what they have published, in what indication, and how recently. Grant databases show whether they currently hold funded work — which is a proxy for whether they have any bandwidth left.

Two cautions, both learned by getting them wrong. First, names collide — a common surname and an initial is not an identity, and merging two investigators into one produces a track record that belongs to nobody. Disambiguation has to be explicit, and where it fails the honest output is a lower-confidence match rather than a confident wrong one. Second, some registry records list only a placeholder contact — a “Medical Director” with no name — and a pipeline that silently drops those sites will systematically under-count industry-run centres. The fallback is to recover the investigator from the trial’s publications instead.

The enrolment arithmetic

Everything above feeds one calculation, and it is embarrassingly simple:

patients enrolled = sites × active months × patients per site per month

Three observations about this identity, all of which get violated in practice.

  1. Active months are not calendar months. A site contributes nothing between selection and its green light. Contracting, ethics, and site initiation eat the front of the curve, and adding a site late buys you fewer active months than the plan assumes — often far fewer.
  2. The per-site rate is where the plan usually breaks. Teams negotiate hard over the number of sites, which is visible and contractual, and accept a per-site rate that came from optimism, a sponsor’s recollection, or a site’s own estimate. The rate is the term the total is most sensitive to and the one with the least evidence behind it.
  3. The rate is not independent of the other two. Add sites in a contested city and the per-site rate falls, because the same eligible patients are being recruited by more studies. Modelling sites and rate as independent inputs is what produces enrolment curves that look fine on the slide and flat in month eight.

Note what is not in that equation: eligibility. The per-site rate is a function of how many people walk through the door and what fraction of them clear the protocol, which is why site selection and criteria design are the same conversation held twice. A protocol that is one exclusion tighter needs a different site list, not just a longer timeline.

What a defensible site rationale contains

  • Each candidate site with its all-time, recent and active counts stated separately — never merged.
  • The status filter and the date the registry was read, on the number itself.
  • Population-normalised country comparison where a denominator exists, and an explicit gap where it does not.
  • Named investigators with the evidence for the match, and a lower confidence flag where the identity is uncertain.
  • An enrolment projection whose per-site rate has a stated basis — and a sensitivity showing what happens if that rate is half of it.

In the product

The Where & How pillar covers the study sites and contacts map, site and investigator selection, site density versus competition, geographic prioritisation, cost feasibility and operational benchmarks — each cited to the record it came from. See how it works.

All writing

See a verdict you can actually check.

Send us a protocol — or just a molecule and an indication. We'll return a fully cited feasibility assessment you can trace, line by line, back to public data — yours to defend in a bid, take to your board or investment committee, or hand to a regulator.