Twitter/X

Opus 5 (Max) ranked #1 with 1,669 pts; GPT-5.6 Luna ranked #21 with 1,497 pts (a…

Brief

Burke Holland reports Code Arena's Image-to-WebDev leaderboard where Opus 5 (Max) tops with 1,669 pts and GPT-5.6 Luna sits at #21 with 1,497 pts (172-point gap). Holland asserts it's cheaper to iterate with Luna than attempt a one-shot with Opus; the benchmark tests image-to-website generation with agentic, multi-step coding and tool use.

Why it matters

Opus 5 (Max) ranked #1 with 1,669 pts; GPT-5.6 Luna ranked #21 with 1,497 pts (a 172-point gap). Other reported scores: #4 GPT-5.6 Sol 1,581; #5 Grok-4.5 1,578; #6 Kimi K3 (Max) 1,571; #8 Muse Spark 1.1 1,550; #13 GPT-5.6 Terra 1,531.

Key details

  • @burkeholland claims it's cheaper to iterate with GPT-5.6 Luna than to try a one-shot approach using Opus 5 (Max).
  • Image-to-WebDev in Code Arena evaluates models' ability to generate websites from images/screenshots using agentic coding workflows that involve multi-step reasoning and tool use.
Source evidence

Basically, it's cheaper to iterate with Luna than try and one-shot with Opus

Arena.ai (@arena)

Update: Scores are in for Image-to-WebDev in Code Arena for Opus 5 (Max), GPT-5.6 Sol, Luna and Terra, Grok-4.5, Kimi K3 (Max), and Muse Spark 1.1. See how they stack up.

  • #1 Opus 5 (Max) with 1,669 pts
  • #4 GPT-5.6 Sol (xHigh) with 1,581
  • #5 Grok-4.5 with 1,578 pts
  • #6 Kimi K3 (Max) with 1,571
  • #8 Muse Spark 1.1 with 1,550 pts
  • #13 GPT-5.6 Terra with 1,531 pts
  • #21 GPT-5.6 Luna with 1,497 pts

Image-to-WebDev in Code Arena, allows you to view overall rankings across AI models on their ability to generate websites from images and screenshots, alongside agentic coding workflows that involve multi-step reasoning and tool use. Learn more below.

— https://nitter.net/arena/status/2083596490539511856#m