Nemotron 3 Ultra (550B-A55B) is here - our strongest open-weight model and full training recipe to date.
Heavy emphasis on real-world inference efficiency for long-context agentic workloads.
Everything is open 🤗: base, post-trained, reward checkpoints, NVFP4 quantized versions, training data, and recipes.
Key technical highlights ‼️:
- 550B total / 55B active parameters
- Hybrid Mamba2-Transformer (~4:1 Mamba:Attention)
- Pretrained in NVFP4 on 20T tokens
- LatentMoE architecture
- Two-stage MOPD post-training
- Native MTP
Technical details in the thread 👇