Twitter/X

Nemotron 3 Embed released with an 8B variant that NVIDIA claims is…

Brief

Nemotron 3 Embed models, announced by NVIDIA AI (and shared by @BarathAnandan7) on 2026-07-16, include an 8B SOTA variant that topped the RTEB retrieval benchmark and a 1B NVFP4 variant trading slight accuracy loss for major performance gains, positioned to improve agent context and response quality.

Why it matters

Nemotron 3 Embed released with an 8B variant that NVIDIA claims is state-of-the-art and reached #1 overall on the RTEB benchmark (announcement tied to NVIDIA AI on 2026-07-16).

Key details

  • A 1B NVFP4 variant is available that sacrifices a small amount of accuracy for a “huge” performance boost, aimed at latency/efficiency-sensitive deployments.
  • RTEB measures retrieval accuracy across real-world tasks, and NVIDIA frames improved retrieval as a way to provide agents more relevant context and better response accuracy.
Source evidence

Nemotron 3 Embed models are here and we're SOTA!

There's a 8B variant if you want the best of them all.

There's a 1B NVFP4 variant if you want a tiny hit on accuracy for a huge performance boost.

Chart for reference and links below if you want to play with it

NVIDIA AI (@NVIDIAAI)

Today we released Nemotron 3 Embed 8B and it reached #1 overall on RTEB 🏆

RTEB benchmarks retrieval accuracy across real-world tasks. Better retrieval gives agents more relevant context, helping improve response accuracy.

— https://nitter.net/NVIDIAAI/status/2077786069840318800#m