From our team at Impossible Research - an agent harness that plays games, writes code, reasons like a physicist, and saturates the ARC-AGI-3 benchmark.
Haven Feng (@HavenFeng)
Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set.
[schema] makes an LLM think like a physicist. 🧵
Video
— https://nitter.net/HavenFeng/status/2077770348876247502#m