Uploaded June 2026 | Updated September 2026, 2 weeks ago
Livestream aired June 29, 2026
AI agents place new demands on inference infrastructure. Unlike a single chatbot response, an agentic workflow can involve many LLM calls, tool calls, long context windows, and repeated cache reuse across a task. NVIDIA Blackwell is designed to handle these production-scale agent workloads with high throughput, low latency, and improved energy efficiency.
This livestream explains how NVIDIA Blackwell helps developers scale AI agents in production, using AgentPerf results as one example of its performance on real-world coding-agent workloads. We’ll also cover how NVIDIA Dynamo adds software-level optimizations for routing, scheduling, and KV cache management.
What you’ll learn:
→ Why AI agents require different infrastructure than standard chat applications.
→ How NVIDIA Blackwell improves throughput and efficiency for concurrent agent workloads.
→ What AgentPerf results show about Blackwell performance on realistic agentic coding tasks.
→ How Dynamo optimizes inference with agent-aware routing, scheduling, and KV cache reuse.
→ What developers should consider when deploying AI agents at production scale.
Links:
nvidia.com/en-us/data-center/gb300-nvl72
developer.nvidia.com/blog/nvidia-achieves-leading-agentic-coding-performance-on-first-agentic-ai-benchmark
Livestream aired June 29, 2026
AI agents place new demands on inference infrastructure. Unlike a single chatbot response, an agentic workflow can involve many LLM calls, tool calls, long context windows, and repeated cache reuse across a task. NVIDIA Blackwell is designed to handle these production-scale agent workloads with high throughput, low latency, and improved energy efficiency.
This livestream explains how NVIDIA Blackwell helps developers scale AI agents in production, using AgentPerf results as one example of its performance on real-world coding-agent workloads. We’ll also cover how NVIDIA Dynamo adds software-level optimizations for routing, scheduling, and KV cache management.
What you’ll learn:
→ Why AI agents require different infrastructure than standard chat applications.
→ How NVIDIA Blackwell improves throughput and efficiency for concurrent agent workloads.
→ What AgentPerf results show about Blackwell performance on realistic agentic coding tasks.
→ How Dynamo optimizes inference with agent-aware routing, scheduling, and KV cache reuse.
→ What developers should consider when deploying AI agents at production scale.
Links:
nvidia.com/en-us/data-center/gb300-nvl72
developer.nvidia.com/blog/nvidia-achieves-leading-agentic-coding-performance-on-first-agentic-ai-benchmark










