The part of what you say to a model that explains why you want it built a certain way — that's the part I've been watching disappear. And it's not the model being lazy. It's something the design of these systems does to your words, every time, on purpose.
So I built a benchmark for it. I'm calling it CARE — Capture and Recovery Eval. It measures how much of your intent (not just your features) survives the planning step these models always run before they go to work. The old benchmark I'd been running on feature coverage is basically maxed — both GPT-5.5 and Opus 4.7 sit around 98% on it — so I needed something that catches what those scores don't.
In this one I'm sharing where the latest models land today, the roughly five-point gap between them, and one finding I keep going back to: turning the reasoning effort all the way up doesn't help — and the same shape shows up in both providers, not just one of them. I'm not ready to claim why that's happening. The numbers are stark enough that I want to keep looking at them out loud.
If you build with Claude Code, Cursor, Copilot, or any agent that turns a back-and-forth conversation into a task list before it starts working, this is for you. If you spend any real time on prompts, planning, or designing agent workflows — same. And if you care about AI evaluation more broadly, CARE is one attempt at measuring the layer most existing benchmarks aren't catching: not whether the features made it through, but whether the reasons you wanted them did.
Topics this touches: AI benchmarks, model evaluation, prompt engineering, agent workflows, planning and reasoning in LLMs, GPT vs Claude comparisons, reasoning effort settings, intent preservation.
📊 CARE benchmark dashboard + repo — coming soon. Subscribe so you don't miss it.
🎬 Prior rigor video (the feature-coverage benchmark this one builds on): https://youtu.be/mkH4N6IXnic?si=yOveropYG71XAJ2G
AI #ArtificialIntelligence #Benchmark #ClaudeCode #PromptEngineering
00:00 - Intro
01:18 - The Old Benchmark
03:28 - What Matters
05:26 - IRL Example
09:38 - BENCHMARK SCORES
14:17 - Plans aren't just from Planning Mode
16:24 - The Curiosity
19:05 - Conclusion
Channel: Matt Maher
Published: 2026-05-14
Video URL: https://www.youtube.com/watch?v=O3ALG38xEDU