这些风险并非假设性的。在最新的放射学基准测试 RadLE 2.0 中,测试的 16 个 AI 模型没有一个的表现能与人类放射科医生相比:https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong/。主要问题是,聊天机器人在给出错误结果时会高度自信,而不是承认它们已经达到了极限,而人类放射科医生在承认不确定性方面要好得多。这种过度自信、说服力和谄媚的混合——聊天机器人验证用户而不是挑战错误假设——也导致了严重的心理健康危害:https://the-decoder.com/openais-sam-altman-warned-of-ais-superhuman-persuasion-in-2023-and-2025-proved-him-right/。Ad DEC_D_Incontent-2
与此同时,一些报告显示,AI 在健康数据中发现了医学专业人员可能遗漏的模式,且有时结果非常显著:https://the-decoder.com/chatgpt-helped-identify-a-genetic-mthfr-mutation-after-a-decade-of-missed-diagnoses/。用于电子健康记录的系统 MIRA 和 AMIE 在模拟诊疗中表现大致与初级保健医生相当:https://the-decoder.com/ai-systems-rival-doctors-in-new-nature-studies-but-one-result-suggests-the-tech-wont-age-well/。一位研究人员将这样的 AI 代理比作飞机的自动驾驶系统:“这些系统可以通过接管例行任务来支持和减轻医疗专业人员的负担,但最终责任始终由医生承担。” Ad
保持对 AI 的关注。内容清晰、实用,没有冗余。
关注 The Decoder 获取 AI 新闻、背景故事和专家分析。
The Decoder:https://the-decoder.com/
OpenAI is rolling out "Health in ChatGPT" to U.S. users aged 18 and older after unveiling and testing the feature since January. Paying for a subscription could literally save your life.
Users can connect Apple Health, medical records, and wellness apps to review lab results, prepare for doctor's appointments, and analyze sleep or activity data. OpenAI says it won't use connected health data for model training or advertising.
Users on the free version of ChatGPT receive lower-quality health advice:https://openai.com/index/health-in-chatgpt/. OpenAI powers the feature with GPT-5.5 Instant:https://the-decoder.com/chatgpts-new-health-upgrade-beats-doctor-written-answers-openai-says/, which scores lower on health benchmarks than the new flagship model, GPT-5.6 Sol:https://openai.com/index/gpt-5-6/, reserved for paying users. OpenAI will likely defend this two-tier system on ethical grounds by pointing out that both models beat doctors' answers on the HealthBench Professional:https://the-decoder.com/openai-says-its-latest-models-outperform-doctors-in-medical-benchmark/ test. Ad
Even when benchmark results appear decisive, they come from artificial test environments designed to measure knowledge, and doctors may score lower for several reasons. They may be under time pressure, dealing with fatigue, or taking the test without tools such as patient records or input from colleagues. Ad DEC_D_Incontent-1
Benchmarks also can't capture much of what happens during an actual medical exam, where doctors can examine patients in person, pick up on nonverbal cues, and draw on years of experience to assess their overall condition. OpenAI itself repeatedly states in the announcement that ChatGPT can still make mistakes and can't replace medical advice. The company says more than 260 physicians helped develop the Health features.
OpenAI's early tests also found that more than 70 percent of participants asked health questions outside the dedicated Health section because switching to it was too cumbersome. The company has since made Health available in any conversation while keeping the separate Health section for managing data and accessing past health chats. Ad
Those risks aren't hypothetical. In the latest radiology benchmark, RadLE 2.0, none of the 16 AI models tested performed as well as human radiologists:https://the-decoder.com/ai-chatbots-reading-x-rays-can-be-dangerously-confident-even-when-theyre-wrong/. The main problem was that chatbots gave incorrect findings with high confidence instead of admitting when they had reached their limits, while human radiologists were far better at acknowledging uncertainty. That mix of overconfidence, persuasion, and sycophancy, where chatbots validate users rather than challenge false assumptions, has also contributed to serious mental health harms:https://the-decoder.com/openais-sam-altman-warned-of-ais-superhuman-persuasion-in-2023-and-2025-proved-him-right/. Ad DEC_D_Incontent-2
At the same time, some reports show AI spotting patterns in health data that medical professionals miss, sometimes with striking results:https://the-decoder.com/chatgpt-helped-identify-a-genetic-mthfr-mutation-after-a-decade-of-missed-diagnoses/. MIRA, a system for electronic health records, and AMIE both performed about as well as primary care doctors in simulated consultations:https://the-decoder.com/ai-systems-rival-doctors-in-new-nature-studies-but-one-result-suggests-the-tech-wont-age-well/. One researcher compared AI agents like these to an airplane's autopilot: "These systems can support and relieve medical professionals by taking over routine tasks, but ultimate responsibility will always remain with the physicians." Ad
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:OpenAI 向美国 18 岁以上用户推出"Health in ChatGPT"功能,可连接 Apple Health 等数据源。 Aioga 将其归入「产品更新」方向,重点关注它对真实使用和行业竞争的影响。