Twitter/X

Multiple Chinese labs are training models with over 5 trillion parameters (5T+)…

Brief

Chinese LLM landscape: As of 2026-08-07 multiple labs are training 5T+ parameter models (comparable to 'GPT 5.5'/'Opus 5'), with ByteDance reportedly pursuing a >5T model while Kimi K3 is 2.8T. Vendors tout cost gaps—DeepSeek v4 Flash claimed up to 105× cheaper than Fable—and teams are shifting compute to Huawei chips as semiconductor firms (e.g., CXMT IPO, DUV tooling) scale up.

Why it matters

Multiple Chinese labs are training models with over 5 trillion parameters (5T+), described as 'GPT 5.5' / 'Opus 5' territory; ByteDance is reportedly discussing a >5T model while Kimi K3 is 2.8T and ByteDance's founder is said to oppose distilling Western models.

Key details

  • Chinese models claim major cost advantages: DeepSeek v4 Flash is reported as up to 105× cheaper than Fable on some tasks; several labs are abandoning free tiers and shifting to pay-per-use pricing.
  • Most Chinese labs are moving training and inference to Huawei chips, and semiconductor capacity is scaling to support this (examples cited: CXMT IPO and increased DUV tooling).
Source evidence

whew.. that’s western frontier llm territory. just to recap where chinese models are at right now:

  • 5T+ models being trained by multiple labs. this is gpt 5.5, opus 5 territory

  • models are 1/10th the price eg deepseek v4 flash is 105x cheaper than fable on some tasks.

  • some chinese labs are killing the free pricing, shifting to pay-per-use.

  • most labs shifting inference and training to huawei chips

  • semiconductor manufacturers are scaling i.e. cxmt ipo, duv tooling.

Lentils (@Lentils80)

🚨 Reports indicate ByteDance is discussing training a massive LLM model with over 5 trillion parameters

Looks like China is REALLY scaling up now. For reference, Kimi K3 is "only" 2.8 trillion params

ByteDance's founder is also supposedly against distilling western models

— https://nitter.net/Lentils80/status/2085472000886411332#m