Opus 5 is a great model for coding, data analysis, design, biology, knowledge work.
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.
And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon.
www-cdn.anthropic.com/c5fbac…
Claude (@claudeai)
On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:
— https://nitter.net/claudeai/status/2080699497064083942#m