My old AI planning benchmark died in six months, so I built a smarter one with agent teams.
This is the actual workflow: multi-agent ideation, promptware, throwaway tools, a PRD pipeline, and a new benchmark question — when an AI turns a messy request into a plan, does the why survive, or only the features?
Along the way I show the way I actually work with AI when a project gets too large for one chat: starting from an empty folder, designing agents with different jobs, making dense outputs readable, and checking the early numbers when the benchmark starts telling a more complicated story than expected.
Related videos:
- Claude Code and Codex CLI Just Got Quietly Replaced: https://youtu.be/IEHQrRX5aGc
- Why Getting Good at AI Made Everything Harder: https://youtu.be/sgQAoCrsn9Q
- My Biggest AI Unlock - It Does Everything: https://youtu.be/FcWi9j0FiYk
- GPT-5.5 vs Opus 4.7: OpenAI Finally Closed the Gap: https://youtu.be/BUhaXIVbRT4
- Opus 4.7 Hit 97% on My Hardest Benchmark: https://youtu.be/Bbh8ldDwz3g
Chapters:
00:00 - Intro
01:34 - Dead Benchmark
03:30 - Let's make ideas
03:50 - Agent Team 001
10:22 - But then what?
10:48 - Transient Software
15:22 - New Benchmark Intro
15:57 - Plan vs PRD?
18:37 - Let's Make Another Agent Team
18:48 - Agent Team 002
23:26 - Simple, right?
24:19 - The Results-ish
28:18 - Conclusion
Channel: Matt Maher
Published: 2026-05-09
Video URL: https://www.youtube.com/watch?v=mkH4N6IXnic