Twitter/X

On completed frontend designs, Kimi K3 and Claude Sonnet 5 produced comparable…

Brief

Jess (@daRubberDuckiee) ran an Aug 7, 2026 eval comparing Kimi K3 vs Claude Sonnet 5 for frontend design agents: Kimi matched Sonnet on visual similarity for finished outputs (0.739 vs 0.698) but showed slower runs and more failures—affected by Moonshot's overloaded serving—while Kimi's listed $3/$15 per‑M token price is actually higher than Sonnet, making it cheaper only by using fewer turns.

Why it matters

On completed frontend designs, Kimi K3 and Claude Sonnet 5 produced comparable visual quality: Kimi scored 0.739 vs Sonnet 0.698 on visual similarity (eval run published 2026-08-07 by @daRubberDuckiee).

Key details

  • Kimi suffered higher latency and more failed outputs during the test; Moonshot was the primary host at launch and was overloaded, and later hosts (Baseten, Fireworks) picked Kimi up, so timeouts likely reflect serving rather than model quality.
  • Kimi K3's listed pricing is $3/$15 per million tokens, which is higher than Sonnet; Kimi only appeared cheaper per run because it used fewer conversational turns, not because of a lower token price.
Cleaned source text

I ran an eval comparing Kimi K3 vs Claude Sonnet 5 as the agent model for frontend design tasks.

When I ran this, Moonshot was basically the only place serving Kimi K3 and getting slammed with traffic right after launch. Since then Baseten, Fireworks, and others picked it up, so the latency and timeout numbers here probably say more about the serving stack that week than the model itself.

But two interesting findings regardless:

1. On the designs Kimi finished, quality was basically a tie. 0.739 vs Sonnet's 0.698 on visual similarity. The open model is genuinely good.

2. Cheaper because it's open" doesn't hold. Kimi K3 lists at $3/$15 per million tokens, higher than Sonnet. It only came out cheaper per run by using fewer turns.

Jess (@daRubberDuckiee)

Article

Kimi K3 vs Claude Sonnet 5 for frontend design agents

👊Main takeaways

1. Kimi K3 matched Sonnet on completed designs, but slower runs and failed outputs changed the result across the full evaluation.

2. Kimi produced similar quality as Sonnet when runs

— https://nitter.net/daRubberDuckiee/status/2085810812699185657#m