Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane July 25, 2026 1 min read
Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents

Anthropic reports that Opus 5 achieves a zero percent success rate against browser-based prompt injection across 129 specific test scenarios. This result marks a significant improvement over previous versions, where the rate dropped from 5.5 percent in Opus 4.8 to 2.0 percent in general testing by security firm Gray Swan. The achievement comes despite OpenAI previously stating that fully solving this flaw may never be possible. The near-immunity applies specifically to products like Claude Cowork when Auto Mode is enabled.

The success relies entirely on a dual-layer defence system rather than the underlying model alone. One layer scans incoming data for hidden instructions before the AI processes them, while a second layer blocks dangerous actions before execution. Without both measures active, the success rate remains at 3.7 percent for Opus 5 and 0.93 percent for Sonnet 5. The combination of protective software and model updates is required to reach the zero percent figure.

  • Auto Mode must be enabled to activate the dual defence layers.
  • Zero percent success was recorded across 129 browser-based test scenarios.
  • General prompt injection rates still show vulnerabilities without specific protections.
Scroll to Top