Significant jump in problem solving for Opus 5!
I think the ARC benchmarks and similar, the ones where open source models normally fall short is where you end up seeing real world comparisons
Excited to for the vibe check on this Claude upgrade
Claude (@claudeai)
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.
— https://nitter.net/claudeai/status/2080699504576045299#m