🚨 @ZAI_ORG JUST DROPPED GLM-5.2, AND IT IS PUNCHING RIGHT AT THE LEVEL OF CLAUDE OPUS 4.8 🤯
The kicker?
It’s a 753B parameter model with a true 1M-token context, released fully open-source under an MIT license
What makes this release technically interesting:
→ IndexShare Attention: They reuse a single indexer across every 4 sparse layers, cutting per-token FLOPs by 2.9× at a 1M context.
→ Better Speculative Decoding: An improved MTP layer increases acceptance length by up to 20%.
→ Adjustable Compute: Flexible thinking-effort levels let you explicitly trade off between performance and latency.
It’s also Day-0 ready.
You don’t have to wait for the ecosystem to catch up, it’s already supported in transformers, vLLM, and SGLang.
Repo, weights and paper in 🧵 ↓