Lenny's Podcast: Product | Career | Growth

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Brief

Diane Penn, head of product for Anthropic’s Research and Labs teams, recounts joining Anthropic in 2023 when the product org was tiny and the company was still finding its identity. She credits Anthropic’s early, tightly aligned culture and bottoms-up energy—engineers and designers voluntarily assembling prototypes—for creating the conditions to move fast. A vivid anecdote about Golden Gate Cloud (a 24-hour demo built from an interpretability finding that made Claude repeatedly reference the Golden Gate Bridge) shows how research discoveries were turned into public-facing experiences quickly and how those moments helped the team understand their unique product voice.

Diane traces a sequence of technical and product inflection points: Opus 3 (launched around March 2024) as the moment that rallied cross-functional trust and demonstrated the feasibility of a frontier model, and Opus 4/4.5 as the phase when model capability and product vehicles (notably Claude Code / Cloud Code) unlocked broad, end‑to‑end use cases like long-form coding. She highlights the emergent, jagged-edge nature of capabilities: while scaling laws show smooth gains in loss, many higher-level abilities appear discontinuously, which means evals (automated tests that reproduce specific failure modes) are essential for discovering, quantifying and shipping real user impact. Diane and the host agree that product managers must now “sweat the tokens as much as the pixels” — interacting with models directly to shape hypotheses — and that experimentation often wins when it’s communal rather than solitary.

On how labs work, Diane describes small, experimental pods focused on high‑variance, high‑upside bets. Labs’ cultural traits are deliberate: hire people who relish zero-to-one work, keep pods small to stay fast, and treat failed prototypes as learning that can be revisited with future model generations. Product-for-research is an explicit role: Diane’s team converts vague user complaints ("Claude hallucinated") into actionable diagnostics (was it a tool‑use failure, search/synthesis failure, or alignment issue?) and builds evals that researchers can iterate against. She says PRDs haven’t disappeared, but their role has shifted—evals serve as executable requirements for many model-improvement problems while PRDs still align large stakeholder groups and stretch‑goals.

Diane also addressed safety, access and organizational readiness. As models become more capable (the Fable/Mythos era), scrutiny and access restrictions increase; product teams need evolving pre-release testing, fallback UXs and stronger safety packages to maintain user experiences while reducing risk. On hiring and skills, Diane hasn’t changed Anthropic’s PM hiring loop for three years: look for first‑principles thinkers, hands‑on builders who tinker and ship with the models, and low‑ego teammates who enable rapid collaboration. She closes with practical people advice—pair up to find joy and accelerate discovery, keep managers hands‑on to maintain theory-of-mind with the tech, and preserve human judgment and persistence as the most durable human strengths in a rapidly automating world.

Why it matters

Diane Penn joined Anthropic in 2023 as the company's first technical product manager when the product team was about five engineers (Diane Penn).

Key details

  • Opus 3 was an internal inflection point (launched around early March 2024) that rallied research, inference and fine-tuning teams and built trust across product and research (Diane Penn).
  • A 24-hour lab-to-product experiment called "Golden Gate Cloud" showcased an interpretability quirk (a model feature that made Claude obsess about the Golden Gate Bridge) and reached roughly 2,000 people, demonstrating early bottoms-up product culture (Diane Penn).
  • Diane: 'Evals are the new PRDs'—Anthropic product teams now define user problems via automated, reproducible evals (examples: JSON/schema-following tests) to make feedback actionable for researchers and to track regressions across model versions.
  • Diane and Lenny discussed the $100,000/year 'token maxing' hypothesis (attributed to YC's Gary Tan via the hosts): heavy token spend today approximates how people will work in 2028, and Diane says token-driven experimentation is a key input for rapid learning.
  • Labs at Anthropic operate as small, founder-like pods (examples: Claude Code, Cloud Code, skills, TAG, MCP, Claude Design) that pursue discontinuous bets with strong convictions about themes and weak convictions about exact prototypes (Diane Penn).
Reader · no content

No body text on file.

Open the original to read the full piece.