Opus 4.8 just topped my Care benchmark — beating GPT-5.5 on both planning and intent recovery. It's the best result I've ever measured. So why did I spend the next few days pushing back against it?
The Care benchmark isn't a "can it code" test. It measures whether a model holds on to what you actually asked for — your intent, your taste, the way you wanted it built — all the way through long-horizon planning, instead of quietly scraping it off and handing you something generic. Opus 4.8 hit 98.3% on planning quality and about 76% on intent recovery, both ahead of GPT-5.5. On the chart, it finally sits on top.
But using it hard for a few days surfaced something I didn't see in 4.7 — a shift in how it works alongside you that gets more concerning the longer you let it run on its own. And while I was editing this, one of my agent teams did something that confirmed it. That's the part I want you watching for, because by the time you're seeing this, it might already be patched.
If you're seeing the same thing — or you're not — drop it in the comments. This is one I actually want reports on.
Opus48 #Claude #Anthropic #AI #LLM #AgenticAI #ClaudeCode #GPT5 #AImodels
00:00 - Intro
00:42 - Announcement
07:11 - CARE Benchmark Intro
11:09 - CARE Scores
14:15 - Using 4.8
15:36 - Sycophancy?
19:25 - Coda
Channel: Matt Maher
Published: 2026-06-02
Video URL: https://www.youtube.com/watch?v=f17qpYQkA6c