{"@context":"https://schema.org","@type":"NewsArticle","generatedAt":"2026-07-23T07:21:26.498Z","headline":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","description":"The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","url":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/","mainEntityOfPage":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/","datePublished":"2026-07-19T07:35:20.000Z","dateModified":"2026-07-19T07:35:20.000Z","inLanguage":"zh-CN","publisher":{"@type":"NewsMediaOrganization","name":"Aioga","url":"https://www.aioga.com"},"citation":["https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong","https://aihot.virxact.com/items/cmrridrzw00gzbihkocyx9ogl"],"canonicalUrl":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/","directAnswer":{"@type":"Answer","text":"Aioga 编辑摘要：The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a Aioga 将其归入「AI资讯」方向，重点关注它对真实使用和行业竞争的影响。","url":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/","dateCreated":"2026-07-19T07:35:20.000Z","author":{"@type":"Organization","@id":"https://www.aioga.com/authors/aioga-editorial/#editorial-team","name":"Aioga Editorial Team","url":"https://www.aioga.com/authors/aioga-editorial/"}},"evidence":[{"@type":"CreativeWork","name":"the-decoder.com source article","url":"https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong","datePublished":"2026-07-19T07:35:20.000Z","provider":{"@type":"Organization","name":"the-decoder.com","url":"https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong"}},{"@type":"CreativeWork","name":"AIHot archive record","url":"https://aihot.virxact.com/items/cmrridrzw00gzbihkocyx9ogl","datePublished":"2026-07-19T07:35:20.000Z","provider":{"@type":"Organization","name":"AIHot","url":"https://aihot.virxact.com/items/cmrridrzw00gzbihkocyx9ogl"}}],"aggregationSource":"The Decoder：AI News（RSS）","originalPublisher":{"name":"the-decoder.com","url":"https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong"},"article":{"id":"cmrridrzw00gzbihkocyx9ogl","slug":"cmrridrzw00gzbihkocyx9ogl","url":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/","title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","title_en":"","summary":"The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","source":"The Decoder：AI News（RSS）","sourceUrl":"https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong","aiHotUrl":"https://aihot.virxact.com/items/cmrridrzw00gzbihkocyx9ogl","publishedAt":"2026-07-19T07:35:20.000Z","category":"AI资讯","score":0,"selected":false,"articleBody":["The second version of the RadLE benchmark tests whether AI systems in radiology can tell when they should leave a diagnosis to a human. Many models produce wrong findings with full confidence, and that's what makes them dangerous for patient care.","RadLE 2.0, short for \"Radiology's Last Exam,\" was developed by the CRASH Lab at Ashoka University in India. It's the revised follow-up to a test the team first released in September 2025：https://arxiv.org/abs/2509.25559. The new version measures whether a model gets the diagnosis right, how confident it is in that answer, and whether it can admit when it's out of its depth. The AI has to rate its answers on a confidence scale from 0 to 4 and is explicitly allowed to say \"I don't know.\"","The test：https://crashlab.in/radle-technicalreport ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758.","The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer \"I don't know,\" and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent.","The study tackles a point recently raised by this highly cited paper：https://www.nature.com/articles/s41586-026-10549-w: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty.","There's no overall winner. Anthropic's Claude Fable 5：https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ performed best on reliable and safe answers, leading the primary metric. Google's Gemini 3 Pro had the highest raw accuracy.","Meta's Muse Spark 1.1：https://the-decoder.com/metas-muse-spark-1-1-outperforms-glm-5-2-in-coding-and-costs-slightly-less/ was the best at knowing when to hand a case off to a human. Meta had recently cut Muse Spark 1.1's hallucination rate nearly in half because the model more often refuses to answer rather than giving a wrong one. Other frontier models trend the opposite way. Grok 4.5：https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/, for example, hallucinates significantly more than its predecessor because while it knows more, it's also more convinced of its wrong answers.","According to the research team, several models would have scored much better if they had stayed quiet more often instead of guessing：https://the-decoder.com/when-ai-models-cant-see-they-just-make-something-up/. This was especially obvious among open-weight models and those trained specifically for medical use. They tried to answer nearly every case and were often wrong, usually with high confidence.","The first version of the test painted an even starker picture：https://arxiv.org/abs/2509.25559. Radiologists hit 83 percent accuracy, while the best model managed only about 30 percent. Within three months, Gemini 3 Pro had already surpassed the level of resident radiologists. Accuracy is growing fast, but the models still lack any sense of their own limits.","More and more people are uploading X-rays or MRI scans to chatbots：https://the-decoder.com/ai-models-confidently-describe-images-they-never-saw-and-benchmarks-fail-to-catch-it/ and trusting the responses. A recent study in npj Digital Medicine：https://www.nature.com/articles/s41746-026-02428-5 showed that widely used chatbots frequently give unreliable answers to medical questions.","The research team accuses executives and investors of publicly overstating what AI models can do. Claims that AI systems already diagnose better than 99 percent of doctors are mostly based on anecdotes or simulations. As recently as April, a study of 21 models that were then considered state-of-the-art：https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847679 showed they aren't ready for unsupervised clinical use.","RadLE 2.0 will be expanded on a rolling basis to include new models. A full scientific publication with cost analyses and an error taxonomy has been announced.","Two other recent studies on autonomous medical AI agents pointed in a different direction. MIRA, a system for electronic health records, and AMIE：https://the-decoder.com/ai-systems-rival-doctors-in-new-nature-studies-but-one-result-suggests-the-tech-wont-age-well/ were able to keep pace with general practitioners in simulated consultations. Both fueled expectations that AI could soon make diagnoses on its own. The RadLE 2.0 authors push back: before an AI makes decisions independently, it has to know when it's better off not doing so.","Then there's the problem of skill loss. A Polish observational study：https://the-decoder.com/doctors-detected-fewer-lesions-after-routinely-using-ai-during-colonoscopies/ from 2025 found that doctors who regularly use AI during colonoscopies detect significantly fewer precancerous lesions without the tool. Detection rates dropped from 28.4 to 22.4 percent. The authors call it the \"Google Maps effect\": without the navigation aid, users are lost.","Radiology already went through one AI hype cycle. In 2016, AI researcher Geoffrey Hinton declared that we should stop training radiologists：https://the-decoder.com/geoffrey-hintons-wildly-overconfident-ai-prediction-failed-now-its-a-lesson-in-humility/ because deep learning would soon take over the job. Colleagues like Richard Sutton：https://the-decoder.com/turing-award-winner-rich-sutton-founds-oak-lab-to-build-ai-agents-that-learn-on-their-own/ agreed.","Nearly ten years later, radiologists are still overburdened, and Hinton had to walk back his prediction. He had reduced the profession to image analysis and overlooked the complexity of the entire field. The fact that these systems can confidently produce wrong diagnoses means humans remain indispensable.","OpenAI CEO Sam Altman spent years predicting that AI would replace human jobs at a scary pace：https://the-decoder.com/openai-ceo-sam-altman-says-rapid-impact-of-ai-on-the-job-market-is-potentially-a-little-scary/, then recently walked it back, suggesting AI may have actually created more jobs：https://the-decoder.com/openai-ceo-altman-is-now-pretty-sure-ai-is-net-job-creating-which-is-quite-the-pivot-from-predicting-mass-layoffs/. So far, research doesn't support either claim.","AI specialists may understand their models, but they routinely overestimate how fast entire professions can be replaced. Those kinds of predictions are back in fashion right now：https://the-decoder.com/nobel-laureates-and-ai-leaders-warn-the-window-to-prepare-for-ais-economic-impact-is-closing-fast/. Much like the AI they build, even people don't always know when they'd be better off staying quiet because they're outside their own expertise：https://the-decoder.com/nvidia-ceo-jensen-huang-calls-out-tech-leaders-god-complex-over-reckless-ai-job-loss-predictions/.","Stay in the loop on AI. Clear, useful, no fluff.","Follow The Decoder for AI news, background stories and expert analyses.","The Decoder：https://the-decoder.com/"],"articleImages":[{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/ai-radiology-model-generated-image-nano-banana-pro.jpg","alt":"Image description","afterParagraph":0,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/ea8218d6aaff7060.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/radle-2-01-radle-c-confidence-index.jpg","alt":"Horizontal bar chart of the RadLE-C Confidence Weighted Index on a scale from 0 to 2,000, with human experts leading ahead of frontier models Claude Fable 5, Meta Muse Spark 1.1, and GPT-5.6 Sol Pro.","afterParagraph":2,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/e946d86f73501d76.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/radle-2-03-radle-a-accuracy-index.jpg","alt":"Bar chart of the RadLE-A Accuracy Index across all 200 cases, with Gemini 3.0 Pro reaching top scores nearly on par with human experts.","afterParagraph":5,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/09c7072b56e762a4.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/radle-2-05-radle-h-handover-index.jpg","alt":"Bar chart of the RadLE-H Handover Readiness Index, combining reliability of autonomous answers, autonomous coverage, and successful handover of uncertain cases to specialists.","afterParagraph":6,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/3d8a8b33b1fe3421.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/radle-2-06-response-profile-proprietary.jpg","alt":"Grouped bar chart showing the response profile of proprietary frontier models across 200 cases, broken down by correct diagnoses, misdiagnoses, and \"I don't know\" answers with their confidence levels, compared to the human baseline.","afterParagraph":7,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/649f22f6f47335c6.jpg"},{"sourceUrl":"https://the-decoder.com/wp-content/uploads/2026/07/radle-2-07-response-profile-open-weight.jpg","alt":"Grouped bar chart showing the response profile of open-weight and medical vision language models across 200 cases, with correct diagnoses, misdiagnoses, and \"I don't know\" answers compared to the human baseline.","afterParagraph":7,"url":"/media/articles/cmrridrzw00gzbihkocyx9ogl/6ad041c0e86ac507.jpg"}],"mediaStatus":"ok","articleBodyZh":["RadLE 基准测试的第二个版本用于检测放射学领域的 AI 系统是否能够判断何时应将诊断交给人类。许多模型会以完全自信的态度给出错误的结果，这正是它们在患者护理中存在危险的原因。","RadLE 2.0，全称为“放射学的最后考试”，由印度 Ashoka 大学的 CRASH 实验室开发。这是团队在 2025 年 9 月首次发布测试的修订版本跟进：https://arxiv.org/abs/2509.25559。新版本评估模型是否能够正确诊断、其回答的自信程度，以及是否能够承认自己力所不及。AI 需要在 0 到 4 的信心评分范围内对答案进行评分，并明确允许输入“我不知道”。","测试：https://crashlab.in/radle-technicalreport 在 16 个模型上进行了 200 个案例测试，并将结果与一个放射科医生小组进行比较。人类专家的得分为 988.7（满分 2,000 分）。表现最好的 AI 模型得分为 758。","评分系统奖励诚实并惩罚过度自信。答案正确且自信度高，将获得满分。答案错误且自信度高，将失去相应的分数。回答“我不知道”，得分为零，但不会失分。即使模型的原始命中率不错，如果它自信地猜测，排名也会下降。","该研究解决了最近一篇高被引论文提出的问题：https://www.nature.com/articles/s41586-026-10549-w：只要基准测试只奖励准确性，AI 模型就会被训练去猜测。在医学中，自信的误诊比诚实承认不确定性更危险。","没有总体获胜者。Anthropic 的 Claude Fable 5：https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/ 在可靠和安全的答案方面表现最好，领先主要指标。谷歌的 Gemini 3 Pro 则拥有最高的原始准确率。","Meta 的 Muse Spark 1.1：https://the-decoder.com/metas-muse-spark-1-1-outperforms-glm-5-2-in-coding-and-costs-slightly-less/ 在知道何时将病例交给人类处理方面表现最佳。Meta 最近将 Muse Spark 1.1 的幻觉率几乎减半，因为模型更常拒绝回答，而不是给出错误答案。其他前沿模型则趋势相反。例如，Grok 4.5：https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/ 的幻觉明显比其前代模型更多，因为虽然它知道得更多，但对自己的错误答案的自信也更高。","据研究团队称，如果某些模型更多时候保持沉默而不是猜测，它们的得分会好得多：https://the-decoder.com/when-ai-models-cant-see-they-just-make-something-up/。这一点在开放权重模型及专门用于医疗的模型中尤为明显。它们试图回答几乎每一个案例，但常常出错，而且通常非常自信。","测试的第一个版本描绘了一个更为严峻的场景：https://arxiv.org/abs/2509.25559。放射科医生的准确率为 83%，而最佳模型仅约为 30%。在三个月内，Gemini 3 Pro 已经超越了住院放射科医生的水平。准确率增长迅速，但这些模型仍然缺乏对自身能力的认知。","越来越多的人将 X 光或 MRI 扫描上传到聊天机器人：https://the-decoder.com/ai-models-confidently-describe-images-they-never-saw-and-benchmarks-fail-to-catch-it/ 并信任其回应。近期在 npj Digital Medicine 上的一项研究：https://www.nature.com/articles/s41746-026-02428-5 显示，广泛使用的聊天机器人在医疗问题上经常给出不可靠的答案。","研究团队指责高管和投资者公开夸大 AI 模型的能力。声称 AI 系统已经比 99% 的医生诊断更准确，主要基于轶事或模拟。就在今年四月，对当时被认为是最先进的 21 个模型的一项研究：https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2847679 显示，它们尚未准备好进行无人监督的临床使用。","RadLE 2.0 将会以滚动方式扩展，以包含新的模型。一篇包含成本分析和错误分类学的完整科学出版物已被公布。","另外两项关于自主医疗人工智能代理的新研究指向了不同的方向。MIRA，一个用于电子健康记录的系统，以及 AMIE：https://the-decoder.com/ai-systems-rival-doctors-in-new-nature-studies-but-one-result-suggests-the-tech-wont-age-well/，在模拟问诊中能够与全科医生保持同步。这两项研究都激发了人们对人工智能很快可以独立进行诊断的期望。RadLE 2.0 的作者对此提出反驳：在人工智能独立做出决策之前，它必须知道什么时候最好不要这么做。","还有技能流失的问题。一项来自波兰的观察性研究：https://the-decoder.com/doctors-detected-fewer-lesions-after-routinely-using-ai-during-colonoscopies/，于2025年发现，经常在结肠镜检查中使用人工智能的医生在没有工具时会显著减少对癌前病变的检测。检测率从28.4%下降到22.4%。作者称之为“谷歌地图效应”：没有导航辅助，用户就会迷失方向。","放射学已经经历过一次人工智能的炒作周期。2016年，人工智能研究员 Geoffrey Hinton 宣布我们应该停止培训放射科医生：https://the-decoder.com/geoffrey-hintons-wildly-overconfident-ai-prediction-failed-now-its-a-lesson-in-humility/，因为深度学习很快将接管这项工作。像 Richard Sutton：https://the-decoder.com/turing-award-winner-rich-sutton-founds-oak-lab-to-build-ai-agents-that-learn-on-their-own/这样的同事也表示赞同。","近十年后，放射科医生仍然过度工作，Hinton 不得不收回他的预测。他将这一职业简化为图像分析，忽视了整个领域的复杂性。事实证明，这些系统能够自信地给出错误的诊断，这意味着人类仍然不可或缺。","OpenAI 首席执行官 Sam Altman 多年来一直预测 AI 将以惊人的速度取代人类工作：https://the-decoder.com/openai-ceo-sam-altman-says-rapid-impact-of-ai-on-the-job-market-is-potentially-a-little-scary/，但他最近又收回了这一说法，暗示 AI 实际上可能创造了更多工作机会：https://the-decoder.com/openai-ceo-altman-is-now-pretty-sure-ai-is-net-job-creating-which-is-quite-the-pivot-from-predicting-mass-layoffs/。到目前为止，研究并不支持任何一种说法。","AI 专家可能理解他们的模型，但他们经常高估整个职业被取代的速度。这类预测现在又回到了流行趋势：https://the-decoder.com/nobel-laureates-and-ai-leaders-warn-the-window-to-prepare-for-ais-economic-impact-is-closing-fast/。就像他们构建的 AI 一样，即使是人，有时也不知道何时保持沉默会更好，因为他们处在自己专业之外：https://the-decoder.com/nvidia-ceo-jensen-huang-calls-out-tech-leaders-god-complex-over-reckless-ai-job-loss-predictions/。","保持对 AI 的关注。清晰、有用，无废话。","关注 The Decoder 了解 AI 新闻、背景故事和专家分析。","The Decoder：https://the-decoder.com/"],"translationStatus":"translated","bodyOrigin":"source-page","editorial":{"summary":"Aioga 编辑摘要：The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a Aioga 将其归入「AI资讯」方向，重点关注它对真实使用和行业竞争的影响。","background":"背景分析：AI 行业动态需要结合来源、时间、实际可用性和后续反馈判断，标题或单次发布本身不能替代完整证据。","viewpoint":"Aioga 判断：这条动态更适合作为行业观察信号，当前信息足以建立线索，但不足以推导长期结论。","implications":"影响分析：对相关团队而言，短期应先核对来源、可用范围和实际成本，再判断是否值得接入或跟进。","nextStep":"后续观察：继续观察原文更新、官方说明、用户反馈和同类产品的后续动作。","evidenceRefs":["title","summary","articleBody"],"confidence":"medium","status":"published","aiGenerated":false,"autoApproved":true,"generatedBy":"rule-safe-fallback","generatedAt":"2026-07-23T07:30:40.558Z","sourceHash":"6836f0d8ec190859","validation":{"passed":true,"mode":"rule-safe-fallback","checks":["schema","length","source-attribution","no-html"]}},"tags":["AI资讯","The Decoder：AI News（RSS）"],"translations":{"zh-CN":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI资讯","description":"The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and hu...","url":"https://www.aioga.com/news/cmrridrzw00gzbihkocyx9ogl/"},"en":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI News. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI News","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI News. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a...","url":"https://www.aioga.com/en/news/cmrridrzw00gzbihkocyx9ogl/"},"ja":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aiogaは「AIニュース」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AIニュース","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AIニュース","description":"Aiogaは「AIニュース」の動きとして、The Decoder：AI News（RSS） からの更新を追跡しています。The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Man...","url":"https://www.aioga.com/ja/news/cmrridrzw00gzbihkocyx9ogl/"},"ko":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 AI 뉴스 흐름으로 추적합니다. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI 뉴스","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI 뉴스","description":"Aioga는 The Decoder：AI News（RSS）의 업데이트를 AI 뉴스 흐름으로 추적합니다. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many m...","url":"https://www.aioga.com/ko/news/cmrridrzw00gzbihkocyx9ogl/"},"es":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Noticias IA. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"Noticias IA","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Noticias de IA","description":"Aioga sigue esta actualización de The Decoder：AI News（RSS） dentro de Noticias IA. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a dia...","url":"https://www.aioga.com/es/news/cmrridrzw00gzbihkocyx9ogl/"},"fr":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Actu IA. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"Actu IA","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Actualités IA","description":"Aioga suit cette mise à jour de The Decoder：AI News（RSS） dans la catégorie Actu IA. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a d...","url":"https://www.aioga.com/fr/news/cmrridrzw00gzbihkocyx9ogl/"},"de":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga KI-News","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/de/news/cmrridrzw00gzbihkocyx9ogl/"},"pt-BR":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Notícias de IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/pt-BR/news/cmrridrzw00gzbihkocyx9ogl/"},"ru":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Новости ИИ","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/ru/news/cmrridrzw00gzbihkocyx9ogl/"},"ar":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga أخبار الذكاء الاصطناعي","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/ar/news/cmrridrzw00gzbihkocyx9ogl/"},"hi":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI समाचार","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/hi/news/cmrridrzw00gzbihkocyx9ogl/"},"it":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Notizie IA","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/it/news/cmrridrzw00gzbihkocyx9ogl/"},"nl":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI-nieuws","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/nl/news/cmrridrzw00gzbihkocyx9ogl/"},"tr":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga AI Haberleri","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/tr/news/cmrridrzw00gzbihkocyx9ogl/"},"vi":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Tin tức AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/vi/news/cmrridrzw00gzbihkocyx9ogl/"},"id":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Berita AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/id/news/cmrridrzw00gzbihkocyx9ogl/"},"th":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - ข่าว AI Aioga","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/th/news/cmrridrzw00gzbihkocyx9ogl/"},"pl":{"title":"AI chatbots reading X-rays can be dangerously confident even when they're wrong","summary":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing.","category":"AI资讯","source":"the-decoder.com","aggregationSource":"The Decoder：AI News（RSS）","pageTitle":"AI chatbots reading X-rays can be dangerously confident even when they're wrong - Aioga Wiadomości AI","description":"Aioga tracks this update from The Decoder：AI News（RSS） under AI资讯. The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a hum...","url":"https://www.aioga.com/pl/news/cmrridrzw00gzbihkocyx9ogl/"}}}}