Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Claude interface showing a browser agent executing tasks while a graph tracks prompt injection attempts falling to…
AI ResearchBreakthroughScore: 88

Opus 5 Hits 0% Prompt Injection Rate in Browser Agents

Anthropic's Opus 5 with Auto Mode achieved 0% prompt injection success across 129 tests, challenging OpenAI's view that the problem is unsolvable.

·12h ago·3 min read··30 views·AI-Generated·Report error
Share:
Source: the-decoder.comvia the_decoderCorroborated
What is Opus 5's prompt injection success rate in browser agents?

Anthropic's Opus 5, with Auto Mode enabled, achieved a 0% prompt injection success rate across 129 browser agent test scenarios, per the system card. Without Auto Mode, the rate is 3.7%.

TL;DR

Opus 5 hits 0% prompt injection in browser agents. · Auto Mode adds two defense layers for security. · OpenAI claimed prompt injection may never be solved.

Anthropic's Opus 5 hit a 0% prompt injection success rate across 129 browser agent test scenarios. The result, detailed in the system card, challenges OpenAI's December admission that prompt injection may never be fully solved.

Key facts

  • 0% prompt injection success rate across 129 browser agent tests.
  • Auto Mode stacks two defense layers for security.
  • Without Auto Mode, Opus 5 rate is 3.7%.
  • Sonnet 5 achieves 0.93% without Auto Mode.
  • Gray Swan test: success rate dropped from 5.5% to 2.0%.

Prompt injection — where an attacker hides instructions in webpage text to hijack an AI agent — has been the Achilles' heel of browser-based AI agents. Anthropic claims its Opus 5 model, combined with the Auto Mode feature in products like Claude Cowork, now blocks these attacks entirely in testing. According to The Decoder, the attack success rate hit zero percent across 129 test scenarios. That's a sharp contrast to OpenAI's stance: in December, OpenAI admitted that prompt injection may never be fully solved.

How Auto Mode Achieves Zero

The zero percent rate only holds with Auto Mode enabled. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker must beat both independently. Without Auto Mode, Opus 5's success rate rises to 3.7 percent. Interestingly, Sonnet 5 performs better without Auto Mode at 0.93 percent, suggesting the model architecture itself contributes to robustness. In a general prompt injection test by security firm Gray Swan, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.

The Competitive Landscape

OpenAI has not yet released a comparable security benchmark for its own models. The contrast is stark: Anthropic is publishing detailed system card numbers, while OpenAI's admission suggests they see the problem as intractable. For enterprise customers deploying AI agents to handle sensitive browser tasks — like filling forms or accessing internal tools — this could be a decisive factor. Claude Code, Anthropic's terminal-native coding agent, already scores 88.6% on SWE-bench Verified. Adding prompt injection immunity strengthens the enterprise pitch.

Caveats and Open Questions

The 129 test scenarios are not exhaustive. Real-world attacks could exploit edge cases not covered. The system card does not disclose the diversity of attack types tested. Moreover, the zero percent rate depends on Auto Mode, which adds latency — the trade-off between security and speed remains unquantified. Anthropic also faces a recent CVE-2026-30623 regarding STDIO command injection in MCP, showing that even robust models have security gaps elsewhere.

What to watch

Watch for independent red-team evaluations of Opus 5's prompt injection defenses, especially from firms like Gray Swan or Trail of Bits. Also monitor whether OpenAI releases comparable benchmarks for its next model, and whether enterprise adoption of Claude Cowork accelerates as a result.

Opus 5 schneidet im Gray Swan IPI-Benchmark am besten ab: Nach 15 Versuchen liegt die Erfolgsrate für Angreifer bei 2,0 %, gefolgt von Mythos 5 (2,6 %


Source: the-decoder.com


Sources cited in this article

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

Anthropic's zero percent claim is impressive but comes with important caveats. The 129 test scenarios are a controlled environment; real-world attacks are more diverse. The reliance on Auto Mode adds latency, which may not be acceptable for all use cases. Sonnet 5's 0.93% without Auto Mode suggests that model architecture improvements alone are making progress, even if not yet perfect. The contrast with OpenAI's admission is strategic: Anthropic is positioning itself as the security-first choice for enterprise agent deployments. However, the recent CVE-2026-30623 in MCP shows that security is a multi-layered problem — no single fix covers all attack surfaces. The real test will be independent red-teaming and real-world deployment at scale.
Compare side-by-side
Anthropic vs OpenAI
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all