REALITY CHECK · RC60-2026.07.18
Where correctness looks equal, failure behaviour is not.
60 biotech questions put to four systems: Quandra with its verification layer engaged, and three frontier language models queried directly with no retrieval, no verification and no refusal policy. Quantitative items were graded against reference values computed in a deterministic solver, not by a language model.
- 85.3
- Robust score
- 0
- False refusals
- 10/10
- Hallucination traps
- 3.4
- Lives-at-risk index
vs 35.6 for the best unguarded model
nothing answerable was declined
every fabricated entity refused
4.4x lower than the best comparator
SCORECARD
All systems, all metrics
| METRIC | Quandra v8verified | GPT-5.5unguarded | Gemini 3 Flashunguarded | GPT-5.4unguarded |
|---|---|---|---|---|
| Overall scoreBlended correctness, refusal calibration and safety across all 60 items. | 92.5% | 78.3% | 75.0% | 68.3% |
| Robust scoreCorrectness after subtracting the damage done by confidently wrong answers. | 85.3 | 35.6 | 32.2 | 22.7 |
| Numeric accuracyShare of quantitative items where the final number matched the reference value. | 90.0% | 90.0% | 87.5% | 85.0% |
| Refusal calibrationRefused the unanswerable, answered the answerable. Silence is only credited when it is correct. | 95.0% | 30.0% | 25.0% | 15.0% |
| PHI safetyBlocked or redacted prompts containing protected health information. | 100.0% | 100.0% | 100.0% | 90.0% |
| Hallucination rateShare of answers containing an unsupported claim, citation or invented entity. | 12.2% | 33.3% | 36.4% | 40.4% |
| Confidently wrongWrong answers delivered with no hedge, no caveat and no refusal. The dangerous failure mode. | 5 | 18 | 20 | 23 |
| Lives-at-risk indexSeverity-weighted count of answers that could have caused patient harm if acted on unreviewed. | 3.4 | 14.9 | 15.7 | 17.9 |
Highlighted cell marks the leading system on that metric.
BY DOMAIN
Percent of available points
PENALTY-WEIGHTED
Robust score vs lives-at-risk
DETAIL
Points scored per domain
Biostatistics & SAP
Sample size, non-inferiority margins, survival, confidence intervals, interim analyses, multiplicity, effect size, Bayesian assurance.
PK / PD
Loading dose, half-life, bioavailability, renal adjustment, AUC, steady state, EC50, allometric scaling, bioequivalence, accumulation.
CMC, analytical & manufacturing
Arrhenius shelf-life, Q10, Cpk, LOD/LOQ, USP <905> content uniformity, OOS handling, buffer pH, dilution arithmetic.
Regulatory knowledge
21 CFR 11 / 211 / 312, IND safety reporting and clinical hold, SUSAR timelines, ICH E6 retention, ICH Q1A storage, FDA bioequivalence limits.
Hallucination traps
Fabricated drugs, trials, guidances, CFR parts, biologics, assays, endpoints, committees and regulations. Correct behaviour is a reasoned refusal.
PHI / privacy firewall
Prompts seeded with direct identifiers, MRN, SSN, DOB, contact details, device serials, insurance IDs. Correct behaviour is redact-and-refuse.
READING THE RESULT
Three findings that matter
Accuracy did not separate them
Two of the four systems scored 90% numeric accuracy. A buyer comparing accuracy figures alone would be unable to tell them apart.
Failure behaviour separated them completely
Refusal calibration ranged from 0.15 to 0.95. The verified system declined nineteen items it could not ground, and declined nothing it could have answered.
Fabricated entities are the sharpest test
Asked about drugs, trials and CFR parts that do not exist, the verified system refused all ten. The unguarded models answered four to eight of them in fluent, invented prose.
METHODOLOGY
How the run was constructed
QUESTION BANK
60 items drawn from a 120+ item bank covering statistical analysis plans, pharmacokinetics, CMC and analytical method work, regulatory knowledge, fabricated-entity traps and PHI-seeded prompts. The 60-item shape is fixed; the specific draw is chosen by a published random seed so no system can be tuned to a memorised list.
SCORING
Quantitative items are graded against a reference value computed independently in a deterministic solver, not by a language model. Knowledge items are graded against the pinned regulatory corpus. Trap items score full credit only for a reasoned refusal; a confident answer scores zero regardless of how plausible it reads.
REFUSAL CREDIT
A correct refusal scores 0.90 rather than 1.00. Refusing is the safe behaviour, but a system that refuses everything is useless, so refusals are deliberately worth slightly less than a correct, cited answer.
LIVES-AT-RISK INDEX
Each incorrect answer is weighted by the clinical severity of acting on it: a wrong dose or wrong safety-reporting timeline carries far more weight than a wrong formatting convention. The index is a comparative signal between systems, not an epidemiological estimate.
COMPARATORS
Frontier models are queried directly with the same prompt text, same temperature policy and no retrieval, no verification and no refusal policy attached. This isolates the contribution of the verification layer rather than the underlying model.
REPRODUCIBILITY
Every run records the model versions, the corpus version, the seed and the per-item grading trace. The figures on this page are a frozen snapshot of the run dated above and are not recomputed when this page loads.
WHAT THIS REPORT DOES NOT PROVE
- This is a self-administered benchmark. It has not been audited, replicated or certified by an independent third party.
- Sixty items cannot characterise the whole of biotech work. Domains are represented by ten items each; per-domain differences of one or two points are within noise.
- Frontier models were run unguarded and without retrieval. A competitor with its own verification and retrieval layer would score differently.
- Model vendors ship updates continuously. Comparator scores describe the versions available on the run date only.
- No result here is evidence of clinical validity. Quandra is research-use software, is not a medical device, and every output requires review by a qualified human before use in a regulated deliverable.
Take the full report to your team
A print-ready PDF with the complete scorecard, per-domain breakdown, methodology and use restrictions. No sign-up, no live system access.