Anthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection:https://the-decoder.com/claude-cowork-hit-with-file-stealing-prompt-injection-days-after-anthropics-launch/, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card:https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf#page=76. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved:https://the-decoder.com/openai-admits-prompt-injection-may-never-be-fully-solved-casting-doubt-on-the-agentic-ai-vision/. In a general prompt injection test by security firm Gray Swan:https://www.grayswan.ai/, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.
That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork:https://www.anthropic.com/engineering/claude-code-auto-mode. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
情报判断
Aioga 编辑摘要
Aioga 编辑摘要:Anthropic 称其 Opus 5 模型在自家软件中几乎免疫提示词注入攻击。 Aioga 将其归入「模型更新」方向,重点关注它对真实使用和行业竞争的影响。