YouTube

The Thing GPT and Claude Quietly Drop in Every Conversation

Brief

Matt Maher presents a video presentation of CARE (Capture and Recovery Eval), a new benchmark that quantifies how much user intent survives the planning step agents perform before executing tasks. CARE reveals a roughly five-point gap between GPT-5.5 and Opus 4.7, highlights that prior feature-coverage tests (~98%) miss this layer, and shows maxing reasoning effort does not fix the loss.

Why it matters

Matt Maher (video published 2026-05-14) introduces CARE (Capture and Recovery Eval), a benchmark that measures how much user intent survives the planning step agents run before executing tasks.

Key details

  • Prior feature-coverage tests were ~98% for GPT-5.5 and Opus 4.7, but CARE surfaces a roughly five-point gap between those models on intent preservation.
  • Turning reasoning-effort settings to maximum (stronger planning/reasoning) does not improve intent capture/recovery; the same failure shape appears across providers and agent-style systems (Claude Code, Cursor, Copilot).
Source evidence

The part of what you say to a model that explains why you want it built a certain way — that's the part I've been watching disappear. And it's not the model being lazy. It's something the design of these systems does to your words, every time, on purpose.

So I built a benchmark for it. I'm calling it CARE — Capture and Recovery Eval. It measures how much of your intent (not just your features) survives the planning step these models always run before they go to work. The old benchmark I'd been running on feature coverage is basically maxed — both GPT-5.5 and Opus 4.7 sit around 98% on it — so I needed something that catches what those scores don't.

In this one I'm sharing where the latest models land today, the roughly five-point gap between them, and one finding I keep going back to: turning the reasoning effort all the way up doesn't help — and the same shape shows up in both providers, not just one of them. I'm not ready to claim why that's happening. The numbers are stark enough that I want to keep looking at them out loud.

If you build with Claude Code, Cursor, Copilot, or any agent that turns a back-and-forth conversation into a task list before it starts working, this is for you. If you spend any real time on prompts, planning, or designing agent workflows — same. And if you care about AI evaluation more broadly, CARE is one attempt at measuring the layer most existing benchmarks aren't catching: not whether the features made it through, but whether the reasons you wanted them did.

Topics this touches: AI benchmarks, model evaluation, prompt engineering, agent workflows, planning and reasoning in LLMs, GPT vs Claude comparisons, reasoning effort settings, intent preservation.

📊 CARE benchmark dashboard + repo — coming soon. Subscribe so you don't miss it.
🎬 Prior rigor video (the feature-coverage benchmark this one builds on): https://youtu.be/mkH4N6IXnic?si=yOveropYG71XAJ2G

AI #ArtificialIntelligence #Benchmark #ClaudeCode #PromptEngineering

00:00 - Intro
01:18 - The Old Benchmark
03:28 - What Matters
05:26 - IRL Example
09:38 - BENCHMARK SCORES
14:17 - Plans aren't just from Planning Mode
16:24 - The Curiosity
19:05 - Conclusion

Channel: Matt Maher
Published: 2026-05-14
Video URL: https://www.youtube.com/watch?v=O3ALG38xEDU