Twitter/X

@nummanali: Significant jump in problem solving for Opus 5! I think the ARC benchmarks and similar, the ones w...

Significant jump in problem solving for Opus 5!

I think the ARC benchmarks and similar, the ones where open source models normally fall short is where you end up seeing real world comparisons

Excited to for the vibe check on this Claude upgrade

Claude (@claudeai)

On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.

— https://nitter.net/claudeai/status/2080699504576045299#m