Uploaded March 2026 | Updated September 2026, 2 weeks ago
Curious how NVIDIA’s recently released open model, Nemotron 3 Super, was built—and how to run serious agentic workloads on top of it? Nemotron 3 Super is an open, hybrid Mamba‑Transformer MoE model with a 120B‑total/12B‑active design, 1M‑token context, latent MoE, and multi‑token prediction, built for high‑throughput, long‑context reasoning in multi‑agent systems. This livestream is your chance to talk directly with the NVIDIA AI researchers behind Super and learn what’s new.
In this Nemotron Labs livestream, you can ask about:
- How the hybrid Mamba‑Transformer + latent MoE backbone works in practice, and what it means for throughput, memory efficiency, and long‑range reasoning.
- How the 1M‑token context window and multi‑environment RL alignment (via NeMo Gym and NeMo RL) help with multi‑document reasoning, long‑running agent memory, and tool‑using workflows.
- How native NVFP4 pretraining and the open training pipeline—weights, datasets, and recipes—let you customize and deploy Super efficiently on your own infrastructure.
- How to get Nemotron 3 Super running quickly using the public deployment and fine‑tuning cookbooks for vLLM, SGLang, TensorRT‑LLM, and NeMo‑based LoRA/SFT/RLVR workflows.
Don’t miss your chance to learn from NVIDIA AI research experts. Join the stream, bring your multi‑agent and long‑context challenges, and come ready with questions about deploying Nemotron 3 Super everywhere from cloud endpoints to DGX Spark.
Curious how NVIDIA’s recently released open model, Nemotron 3 Super, was built—and how to run serious agentic workloads on top of it? Nemotron 3 Super is an open, hybrid Mamba‑Transformer MoE model with a 120B‑total/12B‑active design, 1M‑token context, latent MoE, and multi‑token prediction, built for high‑throughput, long‑context reasoning in multi‑agent systems. This livestream is your chance to talk directly with the NVIDIA AI researchers behind Super and learn what’s new.
In this Nemotron Labs livestream, you can ask about:
- How the hybrid Mamba‑Transformer + latent MoE backbone works in practice, and what it means for throughput, memory efficiency, and long‑range reasoning.
- How the 1M‑token context window and multi‑environment RL alignment (via NeMo Gym and NeMo RL) help with multi‑document reasoning, long‑running agent memory, and tool‑using workflows.
- How native NVFP4 pretraining and the open training pipeline—weights, datasets, and recipes—let you customize and deploy Super efficiently on your own infrastructure.
- How to get Nemotron 3 Super running quickly using the public deployment and fine‑tuning cookbooks for vLLM, SGLang, TensorRT‑LLM, and NeMo‑based LoRA/SFT/RLVR workflows.
Don’t miss your chance to learn from NVIDIA AI research experts. Join the stream, bring your multi‑agent and long‑context challenges, and come ready with questions about deploying Nemotron 3 Super everywhere from cloud endpoints to DGX Spark.










