Uploaded June 2026 | Updated September 2026, 2 weeks ago
NVIDIA Nemotron 3 Ultra is a 550B total, 55B active-parameter hybrid Mamba-Transformer MoE.
It’s an open-frontier level model specifically designed to be great at operating agentic harnesses, like Hermes and OpenCode.
🛠️ Key Technical Highlights:
Hybrid Backbone: Interleaving Mamba-2 for sequence efficiency and Transformer layers for precision reasoning.
Latent MoE: Routing compressed tokens to 4x as many experts for the same inference cost.
Native NVFP4: Pretrained specifically for NVIDIA Blackwell to cut memory requirements and speed up inference.
API Control: Implementation of enable_thinking, reasoning_budget, and low_effort modes for granular control.
NVIDIA Technical Blog: nvda.ws/3PZORCq
NVIDIA Technical Report: research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf
Nemotron 3 Ultra on Hugging Face: huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
NVIDIA Nemotron 3 Ultra is a 550B total, 55B active-parameter hybrid Mamba-Transformer MoE.
It’s an open-frontier level model specifically designed to be great at operating agentic harnesses, like Hermes and OpenCode.
🛠️ Key Technical Highlights:
Hybrid Backbone: Interleaving Mamba-2 for sequence efficiency and Transformer layers for precision reasoning.
Latent MoE: Routing compressed tokens to 4x as many experts for the same inference cost.
Native NVFP4: Pretrained specifically for NVIDIA Blackwell to cut memory requirements and speed up inference.
API Control: Implementation of enable_thinking, reasoning_budget, and low_effort modes for granular control.
NVIDIA Technical Blog: nvda.ws/3PZORCq
NVIDIA Technical Report: research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf
Nemotron 3 Ultra on Hugging Face: huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4










