Twitter/X

Benchmark posted 2026-08-07 compared 10 models on a reverse‑Jenga physics sim…

Brief

A benchmark thread (posted 2026-08-07) tested 10 models stacking blocks in a reverse‑Jenga physics sim where higher builds increase collapse risk. Opus 5 won with 11.07 m using 15 blocks and locking its turn; GPT, Deepseek, and Kimi each used all 30 blocks with lower heights. A key constraint: you can observe position OR velocity of a block, never both.

Why it matters

Benchmark posted 2026-08-07 compared 10 models on a reverse‑Jenga physics sim; Opus 5 scored 11.07 m using only 15 blocks and ended its turn to lock that score.

Key details

  • GPT pushed to all 30 blocks and reached 7.92 m before the top buckled; Deepseek placed all 30 blocks and reached 8.18 m; Kimi used 30 blocks and finished at 4.88 m; the top three were Claude models.
  • Simulation rule: an agent can know exactly where a block will land OR exactly how fast it’s moving, but never both; community requested a competitive Jenga variant where models can sabotage each other.
Source evidence

SOMEONE BENCHMARKED 10 MODELS ON STACKING BLOCKS IN A PHYSICS SIM

its basically jenga in reverse

every block you place could topple the whole tower, and going higher always means more risk

> opus 5 hit 11.07m using only 15 blocks, then ended its turn to lock in the score instead of gambling for one more
> gpt got 30 blocks in and only reached 7.92m, kept pushing and the top of the tower buckled
> deepseek went for maximum risk on every single placement, all 30 blocks, and still only got to 8.18m
> kimi used all 30 blocks and finished at 4.88m
> the top 3 were all claude models

so the winner used half the blocks and still built the tallest thing on the board

theres a genius constraint too, you can know exactly where a block will land, or exactly how fast its moving, but never both

everyones asking for a jenga version next where the models sabotage each other

Video