Uploaded March 2026 | Updated September 2026, 2 weeks ago
NVIDIA Nemotron 3 Super is a 120B total, 12B active-parameter hybrid Mamba-Transformer MoE designed for high-efficiency agentic reasoning.
Itβs a useful model for agentic and coding tasks - with its 1M context window making it a great option to power agent harnesses like OpenCode.
π οΈ Key Technical Highlights:
- Hybrid Backbone: Interleaving Mamba-2 for sequence efficiency and Transformer layers for precision reasoning.
- Latent MoE: Routing compressed tokens to 4x as many experts for the same inference cost.
- Native NVFP4: Pretrained specifically for NVIDIA Blackwell to cut memory requirements and speed up inference.
- API Control: Implementation of enable_thinking, reasoning_budget, and low_effort modes for granular control.
NVIDIA Technical Blog: nvda.ws/47oQ6jX
NVIDIA Tech Report: nvda.ws/4cxO98m
Nemotron 3 Super on Hugging Face: nvda.ws/3ORn5at
0:00 - Intro to Nemotron-3 Super (120B/12B)
0:55 - Live Demo: Reasoning & Riddles
1:28 - Python Tutorial: API Key & Setup
1:56 - Customizing "Thinking" & Reasoning Budgets
2:44 - Low-Effort vs. Deep Reasoning
3:12 - Nemotron-3 Super on Perplexity AI
3:57 - Building with OpenCode (HTML/CSS Landing Pages)
5:15 - The Neon Snake Game Reveal
5:30 - Resources & Wrap-up
NVIDIA Nemotron 3 Super is a 120B total, 12B active-parameter hybrid Mamba-Transformer MoE designed for high-efficiency agentic reasoning.
Itβs a useful model for agentic and coding tasks - with its 1M context window making it a great option to power agent harnesses like OpenCode.
π οΈ Key Technical Highlights:
- Hybrid Backbone: Interleaving Mamba-2 for sequence efficiency and Transformer layers for precision reasoning.
- Latent MoE: Routing compressed tokens to 4x as many experts for the same inference cost.
- Native NVFP4: Pretrained specifically for NVIDIA Blackwell to cut memory requirements and speed up inference.
- API Control: Implementation of enable_thinking, reasoning_budget, and low_effort modes for granular control.
NVIDIA Technical Blog: nvda.ws/47oQ6jX
NVIDIA Tech Report: nvda.ws/4cxO98m
Nemotron 3 Super on Hugging Face: nvda.ws/3ORn5at
0:00 - Intro to Nemotron-3 Super (120B/12B)
0:55 - Live Demo: Reasoning & Riddles
1:28 - Python Tutorial: API Key & Setup
1:56 - Customizing "Thinking" & Reasoning Budgets
2:44 - Low-Effort vs. Deep Reasoning
3:12 - Nemotron-3 Super on Perplexity AI
3:57 - Building with OpenCode (HTML/CSS Landing Pages)
5:15 - The Neon Snake Game Reveal
5:30 - Resources & Wrap-up










