对157家企业的调查显示,50%的组织在过去一年曾部署通过内部评估但导致客户故障的AI智能体或大语言模型功能,5%的企业完全信任自动化评估,29%认为评估与现实结果对齐不佳是最大局限。
尽管信任度低,66%的企业已允许或正计划在12个月内实现低风险智能体的全自动、无人工干预部署。
A survey of 157 companies shows that 50% deployed AI agents or large language model features that passed internal evaluations but caused customer issues ov...
A survey of 157 companies shows that 50% deployed AI agents or large language model features that
passed internal evaluations but caused customer issues over the past year, 5% of companies completely trust automated evaluations, and 29% believe poor alignment between evaluations and real-world results is the biggest limitation. Despite low trust, 66% of companies have allowed or plan to allow fully automated deployment of low-risk agents without human intervention within 12 months.
对157家企业的调查显示,50%的组织在过去一年曾部署通过内部评估但导致客户故障的AI智能体或大语言模型功能,5%的企业完全信任自动化评估,29%认为评估与现实结果对齐不佳是最大局限。
尽管信任度低,66%的企业已允许或正计划在12个月内实现低风险智能体的全自动、无人工干预部署。
一项覆盖157家企业的调查显示,50%的组织曾在过去一年将通过内部评估的AI智能体或大语言模型功能投入生产,随后出现客户故障,反映内部测试结果与真实环境表现存在偏差。
调查中仅5%的企业表示完全信任自动化评估,29%将评估与现实结果对齐不佳视为最大局限。与此同时,66%的企业已允许或计划在12个月内对低风险智能体实行全自动、无人工干预部署。
Aioga判断,材料呈现的核心矛盾并非企业是否开展评估,而是评估结果能否代表生产环境。自动化评估信任度有限与自动部署推进并存,值得关注其风险识别和上线决策是否充分衔接。
这可能意味着,仅以内部评估通过作为部署依据,仍不足以排除客户侧故障。随着低风险智能体自动部署扩大,企业可能需要更谨慎地界定“低风险”,并持续核对评估结果与现实表现。 值得关注企业是否补充生产环境反馈、客户故障记录及人工复核机制,并明确自动部署的适用范围。后续报道还应核查调查对象构成、评估标准和“客户故障”的具体定义。
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Ingestion channel: RSS · Source domain: venturebeat.com
Source: VentureBeat:AI(RSS)
Original link: Open original source
Aioga archive: Open intelligence page
Content record: summary-fallback · Updated: 2026-07-16T16:40:48.000Z

统一接入主流 AI 模型 API,为开发、测试与生产环境提供稳定调用入口。
立即访问 api.w173.comAioga aggregates global AI updates and preserves source information for verification and citation.