AI models trained to read X-rays frequently make confident diagnoses that turn out to be wrong, posing serious risks in clinical settings. A new benchmark called RadLE 2.0 reveals a critical gap in medical AI: these systems struggle to recognize the limits of their own knowledge.

The benchmark tests whether radiology AI models can accurately assess their own confidence levels and defer uncertain cases to human radiologists. Results show many models fail this test dramatically. They deliver incorrect diagnoses with high confidence, potentially misleading clinicians who rely on AI predictions as decision support tools.

This phenomenon, called overconfidence or poor calibration, differs from simple accuracy problems. A model can be right 85 percent of the time but wrong in ways that matter most. In radiology, a model that confidently misses a tumor is far more dangerous than one that says it's uncertain and routes the case to a human expert.

Human radiologists significantly outperform current AI systems on this calibration task. Experienced doctors naturally know when a case is ambiguous or outside their wheelhouse. They flag uncertain diagnoses for specialist review. AI models haven't learned this discipline yet.

The implications are substantial. Hospitals deploying AI for radiology screening hope to reduce human workload and catch diagnoses faster. But if the AI confidently flags cases as normal when they're abnormal, or vice versa, it creates new dangers. Clinicians may trust the AI output too much, especially when workflow pressure mounts.

Solving this requires fundamentally different training approaches. Models need to learn not just pattern recognition but genuine uncertainty quantification. Some research explores adding explicit abstention options during training, teaching AI when to defer. Others focus on better calibration through confidence estimation techniques.

RadLE 2.0 establishes a needed standard for evaluating medical AI beyond raw accuracy metrics. Before any AI system diagnoses independently, it must demonstrate reliable