FirsthandHealth
medRxiv PreprintsInternational11 October 2026

Do Calibrated-Probability Claims Hold Up in Clinical Diagnosis? A Pre-registered Evaluation of a Decision Model Against Frontier Language Models

This is an official announcement record

Firsthand records what medRxiv Preprints announced and links to the original. The wording below is theirs, not ours.

Background: A new class of AI "decision models" returns a probability for every option in a single pass and is marketed as calibrated, which invites use in clinical triage. Whether those probabilities hold up on clinical diagnosis has not been independently tested. Objective: To test the diagnostic accuracy, calibration, and probabilistic coherence of a decision model (Jev 1.13, TypeSafe AI) against frontier language models, with and without reasoning. Methods: In a pre-registered study (OSF, registered 26 September 2026), we compared Jev with Claude Opus 5.5 (adaptive reasoning, always on) an
— medRxiv Preprints
Read the official announcement

Opens www.medrxiv.org

More from medRxiv Preprints

This content is for informational purposes only and is not medical advice. It is not intended to diagnose, treat, cure, or prevent any disease. Consult a healthcare professional before starting any supplement, treatment, or program — especially if you are pregnant, nursing, taking medication, or managing a health condition.