Twitter/X

Qwen-3.8 Max scores 56 on the Artificial Analysis Intelligence Index (up from 46…

Brief

Qwen-3.8 Max (Alibaba) climbs into the top five with a 56 on the Artificial Analysis Intelligence Index and big per-benchmark gains, but it trails Kimi K3 by 1 point while costing $1.14 per index task versus Kimi's $0.86. Alibaba says Qwen-3.8 Max is a 2.4T MoE (≈95B active), 1M-token multimodal model whose weights will be released next week; higher costs are driven by much larger token usage and some regressions (AA-Omniscience/hallucination).

Why it matters

Qwen-3.8 Max scores 56 on the Artificial Analysis Intelligence Index (up from 46 for Qwen-3.7 Max), placing it in the top 5 but 1 point behind Kimi K3 (57); it costs $1.14 per Intelligence Index task versus Kimi K3's $0.86 (≈25% cheaper), making ROI worse unless Qwen outperforms Kimi on coding or offers discounts.

Key details

  • Model specs and availability: Alibaba reports Qwen-3.8 Max is a 2.4T-parameter Mixture-of-Experts (≈95B active params per forward pass), 1M-token context window, multimodal (text/image/video → text), with first-party API pricing $2/$6 per 1M input/output tokens and weights slated for release next week; it would be ~6× larger than Qwen3.5 (397B) and second only to Kimi K3 (2.8T) among open weights.
  • Benchmark details: Qwen-3.8 Max posts 1739 Elo on GDPval-AA (+468 vs 3.7), ahead of Kimi K3 (1685) and roughly tied with Claude Fable 5 (1743); gains appear in agentic, scientific reasoning and coding (Terminal-Bench +6, CritPt +7, SciCode +4) but AA-Omniscience regressed −10 (hallucination rate +23% → 40%). Cost per task rose because GDPval-AA input tokens ≈15× and output tokens increased 45% (to 145M).
Source evidence

Qwen-3.8 Max enters the top 5 on intelligence index.

BUT

it falls slightly behind Kimi K3 and costs more. so the ROI is off unless it surpasses Kimi on coding or Qwen thinks of some discount prices.

Artificial Analysis (@ArtificialAnlys)

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task ($0.86)

@Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. Alibaba has announced it plans to release the weights next week, a shift in strategy as it has typically kept its Max class of models proprietary. Once released, Qwen3.8 Max would be ~6x larger than Alibaba's largest open weights release to date (Qwen3.5 397B) and the second largest open weights model behind Kimi K3 (2.8T)

Note: we earlier published results showing Qwen3.8 Max scoring 53 on the Artificial Analysis Intelligence Index. Those runs were affected by intermittent issues on the endpoint we were evaluating, and we have re-run all evaluations on Alibaba's public API endpoint

Key takeaways:
➤ Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, up 10 points from Qwen3.7 Max (46). It is in line with Claude Opus 4.8 (max, 56) and sits second among labs from China, ahead of GLM-5.2 (max, 51) but behind Kimi K3 (max, 57)
➤ Qwen3.8 Max scores 1739 Elo on GDPval-AA, a 468 Elo gain over Qwen3.7 Max. This places it ahead of Kimi K3 (1685), effectively tied with Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), and behind only Claude Opus 5 (max, 1852)
➤ Gains over Qwen3.7 Max span agentic evaluations, scientific reasoning and coding: Terminal-Bench v2.1 +6 points, CritPt +7 points, SciCode +4 points and HLE +3 points, with GPQA unchanged. AA-LCR (-2 points) and AA-Omniscience (-10 points, driven by a hallucination rate rising 23% to 40%) regress
➤ Qwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Max ($0.53), at ~1.3x Kimi K3 (max, $0.86) and ~2x GLM-5.2 (max, $0.57). Cost is driven in part by more turns on agentic evaluations, with GDPval-AA input tokens rising ~15x over Qwen3.7 Max, and output token usage up 45% to 145M
➤ The 𝜏³-Bench Banking result (42%) appears out of distribution. It is a 32 point gain over Qwen3.7 Max and places Qwen3.8 Max ahead of models that outscore it on other evaluations

Key model details:
➤ Size: 2.4T total parameters, ~95B active per forward pass (MoE)
➤ Context window: 1M tokens
➤ Multimodal: text, image and video input with text output
➤ Pricing: $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API, with a $0.25 cache hit price. This is lower than Qwen3.7 Max across the board ($2.50/$7.50, with a $0.50 cache hit price)
➤ Availability: Alibaba Cloud first-party API. Alibaba states the weights will be released next week

— https://nitter.net/ArtificialAnlys/status/2085270415614828675#m