Cohort
Eligibility criteria are a cohort tax
Every criterion is added for a good reason and priced as if it were free. The arithmetic is multiplicative, the cost is invisible at the point of decision, and nobody in the room owns it.
Clinical Trial OS · · 5 min read
The argument in short
- Criteria compose multiplicatively, so the tenth costs far more than the first.
- Every criterion is added by someone protecting a real interest; none of them are charged for the enrolment they consume.
- Strictness is only meaningful against a peer set — the same criteria scored the same way on comparable trials.
A protocol’s eligibility section is written by committee, and the committee is right every time. The medical monitor wants the comorbidity exclusion because of a signal in the phase 1 data. The statistician wants the biomarker cut because it sharpens the effect size. Regulatory wants the washout because a precedent asked for one. Safety wants the organ-function floor. Each argument is individually sound.
Nobody in that room is charged for the patients their criterion removes. That is the structural problem, and it is why eligibility is the most reliable place to find a trial that will not enrol on time.
The arithmetic is multiplicative, not additive
Intuition says criteria add up. They multiply. If a criterion retains some fraction of the screened population, the eligible pool after a series of criteria is the product of those fractions, not their sum. The consequence is that the marginal cost of a criterion rises with every criterion already in the protocol — the tenth one is applied to an already-thinned population, so it removes fewer people in absolute terms but a larger share of what you had left.
Illustrative arithmetic — not a measurement
Take five criteria that each retain four fifths of the population they see. That is a mild-sounding constraint applied five times. The pool that survives all five is 0.8⁵ — about a third of where you started. Ten such criteria leave roughly a tenth. These are made-up retention fractions used to show the shape of the curve; real retention is per-criterion, per-indication and per-population, and has to be measured rather than assumed. The point is the exponent, not the numbers.
There is a second effect the multiplication hides. Criteria are not independent. Age floors, organ-function thresholds and performance-status requirements all correlate — they are largely the same underlying patients. So the true joint retention is usually better than the naive product, which is a genuine argument against panicking at the arithmetic above. But correlation cuts the other way too: a biomarker requirement and a prior-treatment exclusion can be anti-correlated, and the population that satisfies both may be far rarer than either alone suggests. Neither direction can be reasoned out from a spreadsheet. It has to be measured against real cohort data.
What kind of criterion is it?
Not all criteria tax the same population in the same way, and it helps to score a protocol along axes rather than as a single count. The eight we score are:
- Laboratory and biomarker thresholds — the cuts that require a test result to fall in a range.
- Comorbidity and safety exclusions — conditions that rule a patient out for their own protection.
- Prior and concomitant treatment, and washout — what they have had, what they are on, and how long they must come off it.
- Performance status and general fitness — ECOG, Karnofsky, and their equivalents.
- Consent, compliance, contraception and lifestyle — the requirements about behaviour rather than biology.
- Age, sex, BMI and other demographics.
- Diagnosis confirmation and disease severity — how firmly the disease must be established, and how bad it must be.
- Reproductive and cycle status — where the indication makes it relevant.
The axes matter because they have different remedies. A washout period is a scheduling tax — those patients exist and can be enrolled later, at a cost in time. A biomarker cut is a prevalence tax — those patients do not exist in your catchment at all, and no amount of extra sites in the same geography fixes it. A consent-and-compliance criterion is often a site tax, falling much harder on some centres than others. Counting criteria without distinguishing these tells you a protocol is long, which you knew.
Strict compared to what?
An absolute criterion count is close to meaningless. Twenty-two criteria is unremarkable in one oncology setting and extraordinary in another. The only useful statement is relative: this protocol sits at the Nth percentile of strictness against comparable trials in this indication and phase.
That comparison is only honest if both sides are scored the same way. It is easy — and tempting — to score your own protocol with a rich, structured extraction of its criteria, and the peer trials with a cruder pass over registry text, because that is what is available. The result is an asymmetry that makes your protocol look systematically stricter or looser than its peers depending on which classifier is more sensitive. One classifier, both sides, and where an extractor fills a gap on your side only, say so in the output.
The question to ask a benchmark
Whenever you are shown a strictness percentile: was the peer set scored by the same rule as my protocol, and which trials are in it? A percentile against an unnamed peer set is a decoration.
Pricing a criterion before it goes in
The practical fix is procedural, not analytical. Make the cost visible at the moment of the decision, in the protocol review where the criterion is proposed. For each candidate criterion, three things:
- What does it retain? A measured or cited retention estimate, with its basis — a real-world cohort, a published cohort characterisation, or an explicit “unknown”. “Unknown” is a legitimate answer and a useful one; it tells you the criterion is being added blind.
- What does it buy? The specific risk it removes or the specific effect-size gain it delivers. If nobody can state this in a sentence, that is informative.
- Which tax is it? Scheduling, prevalence, or site. This determines whether more sites, more time, or a different geography is even a possible remedy.
Run that on the ten most restrictive criteria and you will usually find two or three that cost a great deal of cohort and buy very little — typically inherited from a previous protocol in the programme, where they were load-bearing and now are not. Those are the cheapest weeks you will ever recover from a timeline.
And when the answer comes back that the criterion is genuinely necessary, you have still gained something: a protocol whose enrolment assumption is now defensible, with a named reason for the tax you chose to pay.
In the product
The Cohort pillar covers cohort sizing, eligibility complexity and real-world cohort reachability — including a per-criterion view of which criteria are costing you the most, each with its basis stated. See how it works.
Next
Related reading
Operations
Site selection is a data problem
Track record, competition density and the enrolment arithmetic are all measurable from public records — and the most common site-selection mistake is a counting error, not a judgement error.
Method
Why most feasibility assessments are unfalsifiable
A feasibility pack that cannot be wrong is not a piece of analysis. It is a document. Here is the difference, and what a checkable one looks like line by line.
See a verdict you can actually check.
Send us a protocol — or just a molecule and an indication. We'll return a fully cited feasibility assessment you can trace, line by line, back to public data — yours to defend in a bid, take to your board or investment committee, or hand to a regulator.