whew.. that’s western frontier llm territory. just to recap where chinese models are at right now:
5T+ models being trained by multiple labs. this is gpt 5.5, opus 5 territory
models are 1/10th the price eg deepseek v4 flash is 105x cheaper than fable on some tasks.
some chinese labs are killing the free pricing, shifting to pay-per-use.
most labs shifting inference and training to huawei chips
semiconductor manufacturers are scaling i.e. cxmt ipo, duv tooling.
Lentils (@Lentils80)
🚨 Reports indicate ByteDance is discussing training a massive LLM model with over 5 trillion parameters
Looks like China is REALLY scaling up now. For reference, Kimi K3 is "only" 2.8 trillion params
ByteDance's founder is also supposedly against distilling western models
— https://nitter.net/Lentils80/status/2085472000886411332#m