An 85% Trial Prediction Claim Is Not the Same as Clinical Proof

QuantHealth’s vendor-reported 85% trial prediction accuracy does not establish that a treatment is safe or effective. It can support development decisions, but only after the underlying cases, prediction timing, endpoint scoring and validation have been examined.
The decisive question is not whether 85% sounds high. It is whether the figure represents locked forecasts made before trial readouts, scored against prespecified primary endpoints, calibrated across relevant diseases and reproduced outside the developer’s control.
Reconstruct what the 85% measures
QuantHealth’s evidence page reports an 85% primary-endpoint accuracy rate across more than 120 simulated trials and characterizes its prospective validation as prediction before readout without sponsor data. It lists IMVOKE010, TROPION-Breast01 and ELOQUENT1 as examples. These are the company’s claims; the public page does not provide the complete trial-level dataset needed to calculate the figure independently.
Begin by asking whether more than 120 trials is the actual denominator for the 85% calculation or a broader count of simulations. A suitable trial-level record would identify every eligible study, any exclusion, the simulation date, the information cutoff, the predicted result, the observed result and the scoring rule. Without it, a reader cannot determine how missing, ambiguous or selectively presented cases affected the percentage.
The unit of accuracy matters too. A binary prediction of whether a primary endpoint will be met is different from predicting its effect size, confidence interval or clinical significance. A model can classify a narrow miss and a decisive failure identically, or correctly predict a pass while materially misestimating the treatment effect.
Separate prospective prediction from retrospective fit
Retrospective analysis can test feasibility and expose failure modes, but it is not prospective forecasting. If an outcome was already known, if post-readout information entered the inputs, or if developers selected the best result from several runs after seeing the answer, the exercise measures fit rather than a locked prediction.
For each prospective claim, request evidence that the data cutoff, model version, endpoint interpretation and output were frozen before the trial result became public. The audit trail should identify one prespecified run rather than merely show that some simulation occurred earlier. It should also record who controlled the timestamp and whether the result was registered with an independent party.
“Without sponsor data” needs its own inventory. Public protocols, earlier trials, conference abstracts, regulatory documents and results from related compounds may contain powerful signals. Using them is not necessarily a weakness, but the claim should distinguish biological simulation from synthesis of already available evidence.
Ask for calibration and relevant coverage
Accuracy counts correct classifications under a chosen rule. Calibration tests whether stated probabilities match observed frequencies: across a sufficiently large evaluation set, outcomes assigned a 70% probability should occur at roughly that rate. A model can have acceptable classification accuracy while consistently expressing too much confidence.
Calibration should be reported with uncertainty intervals and enough cases to interpret it. Pooled results can conceal limited evidence in a particular phase, therapeutic area, modality or endpoint family. Broad platform coverage therefore does not demonstrate validated performance in every covered disease.
A useful comparison also needs a baseline evaluated on the same locked cases. Depending on the decision, that might be a historical success rate, a simple model based on phase and disease, or structured expert judgment. An 85% score has less incremental value if a substantially simpler method performs similarly.
Look for external validation and reproducibility
The evidence base for in-silico trials remains uneven. An independent systematic review identified 202 publications and 48 registered trials, including 76 articles directly connected to drug development. Its detailed reproducibility assessment found that 18 of those 76 articles, or 24%, supplied an open-source model implementation, while generated synthetic data were publicly available in 20%.
The review also found that in-silico methods were still uncommon within registered clinical trials and concentrated in particular diseases. That does not invalidate a commercial model, but it limits any inference from broad labels such as “validated across therapeutic areas.” Performance must be demonstrated in cases resembling the disease, phase, modality and endpoint for the proposed use.
Commercial software need not be public to undergo a credible evaluation. An unaffiliated team can test a frozen model on trials withheld from development under controlled data access. The protocol should be set before outcomes are known and should state who chose the cases, applied exclusions, ran the model and adjudicated ambiguous endpoints.
Match the evidence standard to the decision
Simulation evidence can have legitimate regulatory and development roles without replacing clinical evidence. The FDA’s modeling and simulation overview describes computational methods as complementing bench, animal and clinical evidence, and says FDA scientists review industry modeling results and use such approaches in regulatory decision-making. It does not say that a vendor’s aggregate accuracy rate independently proves safety or efficacy.
The required validation therefore depends on the context of use. Prioritizing indications or stress-testing protocol assumptions can tolerate different uncertainty from replacing a control arm or supporting an effectiveness claim. The audit must name the decision, the population to which the result will be applied and the consequence of a false positive or false negative.
A claim-audit worksheet
- Claim: Record the exact metric, denominator, unit of analysis, exclusions and scoring rule.
- Endpoint: Specify the registered primary endpoint and whether the prediction covers pass or fail, direction, effect size and uncertainty.
- Timing: Establish when inputs, model version and output were frozen relative to public and sponsor readouts.
- Data access: List sponsor, public, historical and related-program information available before prediction.
- Calibration: Request probability calibration, uncertainty intervals, missing-case handling and subgroup sizes.
- Coverage: Map the intended disease, phase, modality and endpoint family to comparable evaluation cases.
- External validation: Identify who selected the test set, ran the frozen model and adjudicated outcomes.
- Reproducibility: Determine whether versions, inputs, outputs and evaluation code can be independently audited under suitable confidentiality controls.
- Context of use: State the decision being supported and the cost of an incorrect prediction.
An 85% prediction claim may be scientifically and commercially useful. Until these fields are filled, however, it remains a vendor-reported aggregate with undetermined generalizability and evidentiary weight—not clinical proof.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.