I ran an eval comparing Kimi K3 vs Claude Sonnet 5 as the agent model for frontend design tasks.
When I ran this, Moonshot was basically the only place serving Kimi K3 and getting slammed with traffic right after launch. Since then Baseten, Fireworks, and others picked it up, so the latency and timeout numbers here probably say more about the serving stack that week than the model itself.
But two interesting findings regardless:
1. On the designs Kimi finished, quality was basically a tie. 0.739 vs Sonnet's 0.698 on visual similarity. The open model is genuinely good.
2. Cheaper because it's open" doesn't hold. Kimi K3 lists at $3/$15 per million tokens, higher than Sonnet. It only came out cheaper per run by using fewer turns.
Jess (@daRubberDuckiee)
Article
Kimi K3 vs Claude Sonnet 5 for frontend design agents
👊Main takeaways
1. Kimi K3 matched Sonnet on completed designs, but slower runs and failed outputs changed the result across the full evaluation.
2. Kimi produced similar quality as Sonnet when runs
— https://nitter.net/daRubberDuckiee/status/2085810812699185657#m