Twitter/X

Opus 5 scored 95.5% on ARC‑AGI‑3 on 2026-08-06 (Prime Intellect claim), up from…

Brief

Prime Intellect's Prime Agent harness reportedly raised Opus 5's ARC‑AGI‑3 score to 95.5% as of 2026-08-06, a dramatic rise from 30% ~1.5 weeks earlier and <1% five weeks earlier. The open-source coding harness claims lower token consumption, broad model improvements over proprietary harnesses, and value equivalent to ~3× a model upgrade.

Why it matters

Opus 5 scored 95.5% on ARC‑AGI‑3 on 2026-08-06 (Prime Intellect claim), up from 30% about 1.5 weeks earlier and from <1% five weeks earlier — a 65.5 percentage-point jump that surpasses the human-expert baseline.

Key details

  • @cryptopunk7213 reports Prime Intellect's open-source "Prime Agent" general-purpose coding harness reduces token use, delivers major cross-model gains versus proprietary harnesses, and is "basically worth 3x" the model upgrade itself.
Source evidence

prime intellects new model harness is basically worth 3x the model upgrade itself which is fckin nuts

5 weeks ago the best ai model scored <1% on ARC AGI-3.

1.5 weeks ago opus 5 scored 30% (#1 at the time)

today that same opus 5 model scored 95.5% beating human experts. 60+ point jump

all while using fewer tokens!

you get a smarter model for a much cheaper rate. why wouldn’t you use this?

(open source btw)

Prime Intellect (@PrimeIntellect)

Prime Agent is a general-purpose coding harness

On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific.

We see major improvements across models when compared to their proprietary harnesses:

— https://nitter.net/PrimeIntellect/status/2085087000764568010#m