30 engineering teams entered her company's internal hackathon. A PM won it.
The idea she used was sitting in a public Anthropic blog post. All 30 teams could have read it.
Here's what she built with it.
One agent writes. A second agent gets configured with what actually matters at her company, then attacks the first agent's output. Flaws route back, the first agent revises, the loop runs until the work clears the bar.
She pointed it at the real codebase and shipped it into production.
GANs, applied to agents.
Now look at where the leverage actually sits in that setup.
The builder agent is commodity.
Every team in that room had the same models and the same context window.
The scarce input was the config: a written definition of what "good" means at that company, precise enough that a machine could enforce it with no human in the room to interpret it.
That's a spec. PMs have claimed to own the spec for twenty years. This is the version where vagueness has a cost you can measure.
Jyothi Nookula has been an AI PM since before that was a title. She rebuilt the evaluator live in this episode, starting from an empty folder.
Most PMs are still trying to get the agent to produce the right answer.
The people winning are building the thing that decides whether the answer is right.
Write the evaluator.
Video
Aakash Gupta (@aakashgupta)
She literally showed how she uses adversarial agents in Claude Code to win hackathons and ship features:
2:47 - How she won her hackathon
4:30 - The 5-layer Claude stack
7:17 - Which model to actually use
9:55 - Chat vs Desktop vs Chrome
14:43 - Cowork automations
26:18 - Skills that beat prompts
34:28 - The AI chief of staff build
59:24 - MCPs every PM needs
1:02:16 - Claude Design is here
1:08:53 - Adversarial agents, live
1:13:01 - The new AI builder role
1:15:57 - The 2026 AI PM interview
1:29:06 - Self-improving product loop
Video
— https://nitter.net/aakashgupta/status/2077470028480549202#m