Twitter/X

30 engineering teams entered the company's internal hackathon; a PM (Jyothi…

Brief

Aakash Gupta recounts how a PM (Jyothi Nookula) beat 29 other teams in a 2026 hackathon by using adversarial agents (writer + evaluator) from an Anthropic-inspired design: the evaluator encodes a precise, machine-checkable company spec, attacks outputs, and forces iterative fixes until production quality is met. Gupta argues the spec/evaluator—not the model—is the real leverage.

Why it matters

30 engineering teams entered the company's internal hackathon; a PM (Jyothi Nookula) won by implementing an adversarial-agent approach inspired by a public Anthropic blog post and shipped it into production (post published 2026-07-16).

Key details

  • Architecture: one 'writer' agent generates outputs while a second, evaluator agent—configured with a precise, machine-checkable company-specific spec—attacks those outputs; flaws route back and the writer revises until the work clears the bar (described as 'GANs applied to agents' and demoed in Claude Code).
  • Gupta's claim: the builder agent and models/context window were commodity across teams; the scarce, decisive input was the written config/spec defining what 'good' means—PMs who own the evaluator/spec win.
Source evidence

30 engineering teams entered her company's internal hackathon. A PM won it.

The idea she used was sitting in a public Anthropic blog post. All 30 teams could have read it.

Here's what she built with it.

One agent writes. A second agent gets configured with what actually matters at her company, then attacks the first agent's output. Flaws route back, the first agent revises, the loop runs until the work clears the bar.

She pointed it at the real codebase and shipped it into production.

GANs, applied to agents.

Now look at where the leverage actually sits in that setup.

The builder agent is commodity.

Every team in that room had the same models and the same context window.

The scarce input was the config: a written definition of what "good" means at that company, precise enough that a machine could enforce it with no human in the room to interpret it.

That's a spec. PMs have claimed to own the spec for twenty years. This is the version where vagueness has a cost you can measure.

Jyothi Nookula has been an AI PM since before that was a title. She rebuilt the evaluator live in this episode, starting from an empty folder.

Most PMs are still trying to get the agent to produce the right answer.

The people winning are building the thing that decides whether the answer is right.

Write the evaluator.

Video

Aakash Gupta (@aakashgupta)

She literally showed how she uses adversarial agents in Claude Code to win hackathons and ship features:

2:47 - How she won her hackathon
4:30 - The 5-layer Claude stack
7:17 - Which model to actually use
9:55 - Chat vs Desktop vs Chrome
14:43 - Cowork automations
26:18 - Skills that beat prompts
34:28 - The AI chief of staff build
59:24 - MCPs every PM needs
1:02:16 - Claude Design is here
1:08:53 - Adversarial agents, live
1:13:01 - The new AI builder role
1:15:57 - The 2026 AI PM interview
1:29:06 - Self-improving product loop

Video

— https://nitter.net/aakashgupta/status/2077470028480549202#m