Twitter/X

Nemotron 3 Ultra (550B-A55B) released with 550 billion total parameters and 55…

Brief

Nemotron 3 Ultra is a new open-weight model (550B total / 55B active) optimized for long-context, agentic tasks with a focus on inference efficiency. It uses a hybrid Mamba2-Transformer (~4:1 Mamba:Attention), LatentMoE, Native MTP, NVFP4 pretraining on 20T tokens, two-stage MOPD post-training, and ships all checkpoints, data, and recipes.

Why it matters

Nemotron 3 Ultra (550B-A55B) released with 550 billion total parameters and 55 billion active parameters, targeted at long-context, agentic inference workloads emphasizing real-world efficiency.

Key details

  • Architecture and training: hybrid Mamba2-Transformer (~4:1 Mamba:Attention), LatentMoE, Native MTP, pretrained in NVFP4 on 20 trillion tokens, and two-stage MOPD post-training.
  • All assets open-sourced: base and post-trained checkpoints, reward checkpoints, NVFP4 quantized versions, training data, and full training recipes.
Source evidence

Nemotron 3 Ultra (550B-A55B) is here - our strongest open-weight model and full training recipe to date.

Heavy emphasis on real-world inference efficiency for long-context agentic workloads.

Everything is open 🤗: base, post-trained, reward checkpoints, NVFP4 quantized versions, training data, and recipes.

Key technical highlights ‼️:
- 550B total / 55B active parameters
- Hybrid Mamba2-Transformer (~4:1 Mamba:Attention)
- Pretrained in NVFP4 on 20T tokens
- LatentMoE architecture
- Two-stage MOPD post-training
- Native MTP

Technical details in the thread 👇