YouTube

How I Actually Used AI Agents to Build a Benchmark

Brief

Matt Maher’s 2026-05-09 video (tutorial/demo) shows how he rebuilt a failed planning benchmark using multi-agent 'agent teams', promptware, transient tools, and a PRD pipeline. He walks through designing Agent Team 001 and Agent Team 002 (timestamps 03:50 and 18:48), working from an empty folder, turning messy requests into readable plans, and evaluating early benchmark metrics.

Why it matters

Matt Maher’s prior AI planning benchmark died within six months; on 2026-05-09 he released a new approach that uses multi-agent 'agent teams' to build a more durable benchmark.

Key details

  • His workflow combines multi-agent ideation, 'promptware', throwaway (transient) tools, and a PRD pipeline to turn messy user requests into executable plans while evaluating whether the 'why' (rationale) survives alongside features.
  • He demonstrates two agent teams—Agent Team 001 (03:50) and Agent Team 002 (18:48)—starting from an empty folder, assigning distinct agent roles, formatting dense outputs for readability, and checking early benchmark metrics when results reveal unexpected complexity.
Source evidence

My old AI planning benchmark died in six months, so I built a smarter one with agent teams.

This is the actual workflow: multi-agent ideation, promptware, throwaway tools, a PRD pipeline, and a new benchmark question — when an AI turns a messy request into a plan, does the why survive, or only the features?

Along the way I show the way I actually work with AI when a project gets too large for one chat: starting from an empty folder, designing agents with different jobs, making dense outputs readable, and checking the early numbers when the benchmark starts telling a more complicated story than expected.

Related videos:
- Claude Code and Codex CLI Just Got Quietly Replaced: https://youtu.be/IEHQrRX5aGc
- Why Getting Good at AI Made Everything Harder: https://youtu.be/sgQAoCrsn9Q
- My Biggest AI Unlock - It Does Everything: https://youtu.be/FcWi9j0FiYk
- GPT-5.5 vs Opus 4.7: OpenAI Finally Closed the Gap: https://youtu.be/BUhaXIVbRT4
- Opus 4.7 Hit 97% on My Hardest Benchmark: https://youtu.be/Bbh8ldDwz3g

Chapters:
00:00 - Intro
01:34 - Dead Benchmark
03:30 - Let's make ideas
03:50 - Agent Team 001
10:22 - But then what?
10:48 - Transient Software
15:22 - New Benchmark Intro
15:57 - Plan vs PRD?
18:37 - Let's Make Another Agent Team
18:48 - Agent Team 002
23:26 - Simple, right?
24:19 - The Results-ish
28:18 - Conclusion

Channel: Matt Maher
Published: 2026-05-09
Video URL: https://www.youtube.com/watch?v=mkH4N6IXnic