一项新研究揭示,个性化大语言模型普遍存在过度推断(OI)现象,即编造超出证据支持的用户属性。
在 MirageBench 基准测试中,12 个模型均有 35%-49% 的推断被判定为虚构(均值 41.6%)。
更关键的是,模型自我评估的 OI 与外部评测结果呈负相关(rho = -0.60),表明自我报告的可信度是误导性信号,外部验证才是更可靠的个性化基础。
A new study reveals that personalized large language models generally exhibit an over-inference (OI) phenomenon, which means fabricating user attributes be...
A new study reveals that personalized large language models generally exhibit an over-inference (OI)
phenomenon, which means fabricating user attributes beyond what the evidence supports. In the MirageBench benchmark tests, all 12 models had 35%-49% of inferences judged to be fictitious (average 41.6%). More importantly, the models' self-assessed OI was negatively correlated with external evaluation results (rho = -0.60), indicating that self-reported reliability is a misleading signal, and external verification is a more reliable basis for personalization.
一项新研究揭示,个性化大语言模型普遍存在过度推断(OI)现象,即编造超出证据支持的用户属性。
在 MirageBench 基准测试中,12 个模型均有 35%-49% 的推断被判定为虚构(均值 41.6%)。
更关键的是,模型自我评估的 OI 与外部评测结果呈负相关(rho = -0.60),表明自我报告的可信度是误导性信号,外部验证才是更可靠的个性化基础。
研究指出,个性化大语言模型会过度推断用户属性,即生成超出已有证据支持的画像信息。MirageBench 测试中的 12 个模型均出现这一现象。
材料将这种问题称为过度推断(OI)。在 MirageBench 基准测试中,12 个模型有 35%-49% 的推断被判定为虚构,平均比例为 41.6%。
Aioga 判断,这项研究最值得关注的并非只有虚构比例,还包括模型自我评估与外部评测呈负相关,说明自我报告不能直接充当可信依据。
Aioga 判断,依赖模型自行判断个性化推断是否可靠,可能产生误导。研究给出的 rho 为 -0.60,材料据此强调外部验证是更可靠的个性化基础。 Aioga 建议,后续关注该研究对过度推断的判定方法及外部验证机制,并在个性化应用评估中分别记录模型自我报告与外部评测结果。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: Summary aggregation · Source domain: arxiv.org
Source: HuggingFace Daily Papers(社区热门论文)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: summary-fallback · Updated: 2026-08-05T00:00:00.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.