Method
What “probability of success” can and cannot tell you
Ours is directional, and we say so on the number. Here is why an honest directional read is more useful than a precise-looking one — and what a PoS structurally cannot see.
Clinical Trial OS · · 4 min read
The argument in short
- A PoS is a base rate moved by evidence — it inherits every weakness of the base rate.
- Ours is directional: good for ranking and for pressure-testing, not for pricing a decision to two significant figures.
- The failure modes that kill trials most often are the ones least visible in public data.
Every probability of success — ours, a bank’s, a consultancy’s, the one in your own model — is built the same way. You start from a base rate: how often trials of this phase, in roughly this indication, have historically gone on to succeed. Then you move that base rate up or down for the things that make this particular trial different. That is the whole architecture. Everything else is detail about where the base rate came from and how the adjustments are chosen.
Understanding that structure tells you almost everything about what the number can and cannot do.
What it inherits from the base rate
A base rate is a historical frequency, so it carries the history’s biases. Three matter most:
- Reporting bias. Trials that succeed are published, presented and press-released. Trials that quietly stop are often a registry status change and nothing else. Any base rate assembled from what got written up is skewed toward success, and correcting for that means taking the silent records seriously — which is a data-engineering problem before it is a statistics problem.
- Label noise. “Failure” in a training set is frequently not a failure at all. A trial terminated because a grant ended, a site closed, or a sponsor reprioritised its portfolio tells you nothing about whether the drug works. Treating those as negative outcomes teaches a model the wrong lesson, and there are a lot of them.
- Indication drift. The base rate for “oncology” in 2004 and “oncology” in 2026 describe different scientific worlds. Slice too finely and you have no sample; slice too coarsely and you have a number about a category your trial is not really in.
None of these are fatal. All of them are reasons to distrust a PoS quoted to a decimal place.
Why we call ours directional
A calibrated probability makes a strong promise: among all the trials we scored at 30%, close to three in ten went on to succeed. Making that claim honestly requires a held-out set of outcomes, a scoring rule, and a published record of how the model performed against it over time. Until that exists, quoting a calibrated-sounding number is a claim you cannot support.
A directional signal makes a weaker and more defensible promise: this trial scores lower than that one, and here is the itemised reason why. Directional numbers are good at ordering, at flagging, and at forcing an argument into the open. They are not good at pricing. If your decision changes between 34% and 41%, the PoS is not the input you should be leaning on.
Where the caveat lives
The caveat travels with the number rather than sitting in a footnote. When the engine has no fitted model for a question it says so on the tile, and when it composes several analyses into one verdict the reported confidence is the weakest leg — the minimum across the inputs, not an average that would let a strong analysis mask a thin one.
What a missing input should do
This is the design decision that separates an honest PoS from a decorative one, and it is worth asking any vendor about directly: what happens when a feature is unavailable?
The tempting answer is to fill the gap with a typical value, because a filled model runs and an empty one does not. The problem is that a default is indistinguishable, downstream, from a measurement. Six defaults later you have a confident number assembled mostly from assumptions, and nothing in the output says so.
The honest handling is for a missing feature to contribute nothing — to push the estimate back toward the base rate rather than toward an invented value — and for the fact of its absence to be recorded in the output where a reader can see it. A trial with six missing features should look uncertain, because it is.
The failure modes a PoS cannot see
Even a well-built PoS is blind to a large share of what actually stops trials, because those things are not in the public record at the time you need the estimate:
- A competitor reading out first and changing the standard of care your comparator arm assumed.
- A manufacturing or supply interruption in the investigational product.
- A protocol amendment cascade that resets your enrolment clock mid-study.
- An internal portfolio decision — a merger, a strategic shift, a budget — that has nothing to do with the science.
- The specific investigator who was going to enrol a third of your patients moving institution.
Some of these leave early traces you can watch for. Sponsor delay signals appear in filings; adverse-event signals appear in safety databases; competitive density appears in registry activity. But a number that claimed to have priced all of them would be lying, and the right response is to model the operational risks separately and explicitly rather than to fold them invisibly into one score.
How to use it
Use a PoS to rank — across a portfolio, across design options, across indications for the same asset. Use the itemised contributions to find which assumption the estimate is leaning on hardest, then go and pressure-test that assumption specifically. Treat a large move in the score as an instruction to look at what moved it, not as a result in itself.
Do not use it as a number to put in a valuation without saying what it is. A directional PoS multiplied by a peak-sales forecast produces a figure with a false air of precision, and the person who inherits that slide will not remember the caveat you said out loud.
Our position, stated plainly
Our probability of success is directional. It is not calibrated, it is not a prediction, and it is not a guarantee. It is built to be argued with — which is why every contribution to it is itemised and cited. See the verdict pillar and our trust posture.
Next
Related reading
Evidence
Reading a trial’s failure: what “why stopped” actually means
A free-text field on a registry record is doing more work in the industry’s models than almost anything else. Most of what it says is not a failure — and “completed” is not the same as “worked”.
Method
Why most feasibility assessments are unfalsifiable
A feasibility pack that cannot be wrong is not a piece of analysis. It is a document. Here is the difference, and what a checkable one looks like line by line.
See a verdict you can actually check.
Send us a protocol — or just a molecule and an indication. We'll return a fully cited feasibility assessment you can trace, line by line, back to public data — yours to defend in a bid, take to your board or investment committee, or hand to a regulator.