parallelquant
July 19, 2026 · The Decoder

AI models reading X-rays are often confidently wrong, benchmark finds

The RadLE 2.0 benchmark tests whether AI radiology models know when to defer a diagnosis to a human rather than guess. Many models delivered incorrect findings with high confidence, while human radiologists still substantially outperformed them overall.

Why it matters: Calibrated uncertainty, knowing when not to answer, is arguably the harder unsolved problem in medical AI, more important than raw accuracy, since a confident wrong diagnosis is more dangerous than an admitted 'I don't know.' This is a concrete data point against near-term autonomous AI diagnosis in high-stakes clinical settings.

Related updates