Anthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection:https://the-decoder.com/claude-cowork-hit-with-file-stealing-prompt-injection-days-after-anthropics-launch/, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card:https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf#page=76. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved:https://the-decoder.com/openai-admits-prompt-injection-may-never-be-fully-solved-casting-doubt-on-the-agentic-ai-vision/. In a general prompt injection test by security firm Gray Swan:https://www.grayswan.ai/, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.

That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork:https://www.anthropic.com/engineering/claude-code-auto-mode. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero.
Stay in the loop on AI. Clear, useful, no fluff.
Follow The Decoder for AI news, background stories and expert analyses.
The Decoder:https://the-decoder.com/
