AI chatbots reading X-rays can be dangerously confident even when they're wrong
The Decoder:AI News(RSS)Aioga 编辑团队2026-07-19T07:35:20.000Z热度
The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should l...
AI资讯The Decoder:AI News(RSS)
今日 AI 情报摘要
The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a
diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.
中文正文 · AI 翻译
RadLE 基准测试的第二个版本用于检测放射学领域的 AI 系统是否能够判断何时应将诊断交给人类。许多模型会以完全自信的态度给出错误的结果,这正是它们在患者护理中存在危险的原因。
没有总体获胜者。Anthropic 的 Claude Fable 5:https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ 在可靠和安全的答案方面表现最好,领先主要指标。谷歌的 Gemini 3 Pro 则拥有最高的原始准确率。
测试的第一个版本描绘了一个更为严峻的场景:https://arxiv.org/abs/2509.25559。放射科医生的准确率为 83%,而最佳模型仅约为 30%。在三个月内,Gemini 3 Pro 已经超越了住院放射科医生的水平。准确率增长迅速,但这些模型仍然缺乏对自身能力的认知。
越来越多的人将 X 光或 MRI 扫描上传到聊天机器人:https://the-decoder.com/ai-models-confidently-describe-images-they-never-saw-and-benchmarks-fail-to-catch-it/ 并信任其回应。近期在 npj Digital Medicine 上的一项研究:https://www.nature.com/articles/s41746-026-02428-5 显示,广泛使用的聊天机器人在医疗问题上经常给出不可靠的答案。
研究团队指责高管和投资者公开夸大 AI 模型的能力。声称 AI 系统已经比 99% 的医生诊断更准确,主要基于轶事或模拟。就在今年四月,对当时被认为是最先进的 21 个模型的一项研究:https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847679 显示,它们尚未准备好进行无人监督的临床使用。
OpenAI 首席执行官 Sam Altman 多年来一直预测 AI 将以惊人的速度取代人类工作:https://the-decoder.com/openai-ceo-sam-altman-says-rapid-impact-of-ai-on-the-job-market-is-potentially-a-little-scary/,但他最近又收回了这一说法,暗示 AI 实际上可能创造了更多工作机会:https://the-decoder.com/openai-ceo-altman-is-now-pretty-sure-ai-is-net-job-creating-which-is-quite-the-pivot-from-predicting-mass-layoffs/。到目前为止,研究并不支持任何一种说法。
AI 专家可能理解他们的模型,但他们经常高估整个职业被取代的速度。这类预测现在又回到了流行趋势:https://the-decoder.com/nobel-laureates-and-ai-leaders-warn-the-window-to-prepare-for-ais-economic-impact-is-closing-fast/。就像他们构建的 AI 一样,即使是人,有时也不知道何时保持沉默会更好,因为他们处在自己专业之外:https://the-decoder.com/nvidia-ceo-jensen-huang-calls-out-tech-leaders-god-complex-over-reckless-ai-job-loss-predictions/。
保持对 AI 的关注。清晰、有用,无废话。
关注 The Decoder 了解 AI 新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
The second version of the RadLE benchmark tests whether AI systems in radiology can tell when they should leave a diagnosis to a human. Many models produce wrong findings with full confidence, and that's what makes them dangerous for patient care.
RadLE 2.0, short for "Radiology's Last Exam," was developed by the CRASH Lab at Ashoka University in India. It's the revised follow-up to a test the team first released in September 2025:https://arxiv.org/abs/2509.25559. The new version measures whether a model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it's out of its depth. The AI has to rate its answers on a confidence scale from 0 to 4 and is explicitly allowed to say "I don't know."
The test:https://crashlab.in/radle-technicalreport ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758.
The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer "I don't know," and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent.
The study tackles a point recently raised by this highly cited paper:https://www.nature.com/articles/s41586-026-10549-w: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty.
There's no overall winner. Anthropic's Claude Fable 5:https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy.
Meta's Muse Spark 1.1:https://the-decoder.com/metas-muse-spark-1-1-outperforms-glm-5-2-in-coding-and-costs-slightly-less/ was the best at knowing when to hand a case off to a human. Meta had recently cut Muse Spark 1.1's hallucination rate nearly in half because the model more often refuses to answer rather than giving a wrong one. Other frontier models trend the opposite way. Grok 4.5:https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/, for example, hallucinates significantly more than its predecessor because while it knows more, it's also more convinced of its wrong answers.
According to the research team, several models would have scored much better if they had stayed quiet more often instead of guessing:https://the-decoder.com/when-ai-models-cant-see-they-just-make-something-up/. This was especially obvious among open-weight models and those trained specifically for medical use. They tried to answer nearly every case and were often wrong, usually with high confidence.
The first version of the test painted an even starker picture:https://arxiv.org/abs/2509.25559. Radiologists hit 83 percent accuracy, while the best model managed only about 30 percent. Within three months, Gemini 3 Pro had already surpassed the level of resident radiologists. Accuracy is growing fast, but the models still lack any sense of their own limits.
More and more people are uploading X-rays or MRI scans to chatbots:https://the-decoder.com/ai-models-confidently-describe-images-they-never-saw-and-benchmarks-fail-to-catch-it/ and trusting the responses. A recent study in npj Digital Medicine:https://www.nature.com/articles/s41746-026-02428-5 showed that widely used chatbots frequently give unreliable answers to medical questions.
The research team accuses executives and investors of publicly overstating what AI models can do. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations. As recently as April, a study of 21 models that were then considered state-of-the-art:https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847679 showed they aren't ready for unsupervised clinical use.
RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced.
Two other recent studies on autonomous medical AI agents pointed in a different direction. MIRA, a system for electronic health records, and AMIE:https://the-decoder.com/ai-systems-rival-doctors-in-new-nature-studies-but-one-result-suggests-the-tech-wont-age-well/ were able to keep pace with general practitioners in simulated consultations. Both fueled expectations that AI could soon make diagnoses on its own. The RadLE 2.0 authors push back: before an AI makes decisions independently, it has to know when it's better off not doing so.
Then there's the problem of skill loss. A Polish observational study:https://the-decoder.com/doctors-detected-fewer-lesions-after-routinely-using-ai-during-colonoscopies/ from 2025 found that doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions without the tool. Detection rates dropped from 28.4 to 22.4 percent. The authors call it the "Google Maps effect": without the navigation aid, users are lost.
Radiology already went through one AI hype cycle. In 2016, AI researcher Geoffrey Hinton declared that we should stop training radiologists:https://the-decoder.com/geoffrey-hintons-wildly-overconfident-ai-prediction-failed-now-its-a-lesson-in-humility/ because deep learning would soon take over the job. Colleagues like Richard Sutton:https://the-decoder.com/turing-award-winner-rich-sutton-founds-oak-lab-to-build-ai-agents-that-learn-on-their-own/ agreed.
Nearly ten years later, radiologists are still overburdened, and Hinton had to walk back his prediction. He had reduced the profession to image analysis and overlooked the complexity of the entire field. The fact that these systems can confidently produce wrong diagnoses means humans remain indispensable.
OpenAI CEO Sam Altman spent years predicting that AI would replace human jobs at a scary pace:https://the-decoder.com/openai-ceo-sam-altman-says-rapid-impact-of-ai-on-the-job-market-is-potentially-a-little-scary/, then recently walked it back, suggesting AI may have actually created more jobs:https://the-decoder.com/openai-ceo-altman-is-now-pretty-sure-ai-is-net-job-creating-which-is-quite-the-pivot-from-predicting-mass-layoffs/. So far, research doesn't support either claim.
AI specialists may understand their models, but they routinely overestimate how fast entire professions can be replaced. Those kinds of predictions are back in fashion right now:https://the-decoder.com/nobel-laureates-and-ai-leaders-warn-the-window-to-prepare-for-ais-economic-impact-is-closing-fast/. Much like the AI they build, even people don't always know when they'd be better off staying quiet because they're outside their own expertise:https://the-decoder.com/nvidia-ceo-jensen-huang-calls-out-tech-leaders-god-complex-over-reckless-ai-job-loss-predictions/.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a Aioga 将其归入「AI资讯」方向,重点关注它对真实使用和行业竞争的影响。