YouTube

Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club

Brief

YC Paper Club (presentation, 2026-07-29) featured five researchers presenting practical systems work: Stuart Sul on the 'Parallel Kittens' method for simplifying multi‑GPU kernels; Jon Saad‑Falcon proposing an 'Intelligence per Watt' metric for local vs cloud inference; Mark Saroufim on AI‑written GPU kernels; Misha on heterogeneous inference infrastructure; and Brennan on a GPU‑only game engine for RL.

Why it matters

Stuart Sul (Stanford / Cursor) presented 'Parallel Kittens' (arXiv:2511.13940) at 7:16, proposing a systematic, practical approach to simplifying multi‑GPU AI kernels to make multi‑device implementations easier to build and reason about.

Key details

  • Jon Saad‑Falcon (21:29) introduced 'Intelligence per Watt' (arXiv:2511.07885), a metric to quantify and compare energy efficiency of local versus cloud AI inference to inform on‑device deployment trade‑offs.
  • Other talks: Mark Saroufim (31:05) demonstrated AI‑generated GPU kernels and benchmarking methods; Misha Smelyanskiy (47:04) argued for heterogeneous hardware and infrastructure for inference; Brennan Shacklett (1:04:33) showed a high‑throughput, GPU‑only game engine for RL (Madrona Engine / Shacklett SIGGRAPH23).
Source evidence

At our latest YC Paper Club, researchers and builders presented on multi-GPU kernel optimization, intelligence per watt for local inference, AI-generated GPU kernels and benchmarking, heterogeneous inference infrastructure design, and GPU-accelerated game engines for reinforcement learning. Thanks to the following presenters:

Stuart Sul (Stanford / Cursor), John (Stanford), Mark (PyTorch / GPU Mode / CoreAuto), Misha (Marlo), and Brennan (Stanford)

Transcript: https://www.ycrootaccess.com/p/multi-gpu-kernels-intelligence-per

Chapters:
0:00 – Francois Chaubard: The case for chip and kernel specialization
7:16 – Stuart Sul: Parallel Kittens - Systematic and Practical Simplification of Multi-GPU Al Kernels (https://arxiv.org/abs/2511.13940)
21:29 – Jon Saad-Falcon: Intelligence per Watt - Measuring the Intelligence Efficiency of Local and Cloud AI (https://arxiv.org/abs/2511.07885)
31:05 – Mark Saroufim: When Al Starts Writing Systems Code
47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware
1:04:33 – Brennan Shacklett: Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU (https://madrona-engine.github.io/shacklett_siggraph23.pdf)
1:15:35 – Wrap-up & what's next

Apply to Y Combinator: https://www.ycombinator.com/apply
Work at a startup: https://www.ycombinator.com/jobs

Channel: Y Combinator
Published: 2026-07-29
Video URL: https://www.youtube.com/watch?v=n8dz2FX0_uY