Uploaded May 2025 | Updated September 2026, 4 hours ago
Explore how F5 VELOS enhances your AI initiatives by enabling high-throughput data ingestion and redundancy. This video discusses the importance of object-based storage systems like S3 for distributing data across multiple nodes and highlights the role of F5 Big IP in achieving high-performance data mobility. Discover strategies to accelerate your AI projects effectively.
⬇️⬇️⬇️ JOIN THE COMMUNITY! ⬇️⬇️⬇️
DevCentral is an online community of technical peers dedicated to learning, exchanging ideas, and solving problems - together.
Find all our platform links ⬇️ and follow our Community Evangelists! 👋
➡️ DEVCENTRAL: community.f5.com
➡️ YOUTUBE: youtube.com/devcentral
➡️ LINKEDIN: linkedin.com/showcase/f5-devcentral
➡️ TWITTER: twitter.com/devcentral
Your Community Evangelists:
👋 Jason Rahm: linkedin.com/in/jrahm | twitter.com/jasonrahm
👋 Buu Lam: linkedin.com/in/buulam | twitter.com/buulam
👋 Aubrey King: linkedin.com/in/aubreyking | twitter.com/aubreykingf5
👋 Chase Abbott: linkedin.com/in/chaseabbott1
Explore how F5 VELOS enhances your AI initiatives by enabling high-throughput data ingestion and redundancy. This video discusses the importance of object-based storage systems like S3 for distributing data across multiple nodes and highlights the role of F5 Big IP in achieving high-performance data mobility. Discover strategies to accelerate your AI projects effectively.
⬇️⬇️⬇️ JOIN THE COMMUNITY! ⬇️⬇️⬇️
DevCentral is an online community of technical peers dedicated to learning, exchanging ideas, and solving problems - together.
Find all our platform links ⬇️ and follow our Community Evangelists! 👋
➡️ DEVCENTRAL: community.f5.com
➡️ YOUTUBE: youtube.com/devcentral
➡️ LINKEDIN: linkedin.com/showcase/f5-devcentral
➡️ TWITTER: twitter.com/devcentral
Your Community Evangelists:
👋 Jason Rahm: linkedin.com/in/jrahm | twitter.com/jasonrahm
👋 Buu Lam: linkedin.com/in/buulam | twitter.com/buulam
👋 Aubrey King: linkedin.com/in/aubreyking | twitter.com/aubreykingf5
👋 Chase Abbott: linkedin.com/in/chaseabbott1



![Optimizing AI Inference using NGINX Gateway Fabric
Running AI inference workloads on Kubernetes presents unique challenges. Traditional load balancing methods designed for standard web APIs simply fall short when it comes to the nuanced resource management required for generative LLMs.
In this demo, we explore how the NGINX Gateway Fabric Inference Extension transforms your Kubernetes environment by acting as a dedicated Inference Gateway. Discover how to move beyond basic proxying by utilizing a powerful two-stage architecture—combining standard HTTPRoutes with the InferencePool CRD and Endpoint Picker. This allows NGINX to make intelligent, real-time routing decisions based on queue depth and GPU availability.
What you will see in this demo:
0:11 - The Challenges of AI Workloads: Why standard load balancing fails for GPUs.
0.51 - Two-Stage Orchestration: Understanding the InferencePool and Endpoint Picker.
2:24 - Model-Aware Routing: Inspecting traffic paths to offload lightweight tasks to CPUs and reserve premium GPUs for heavy text generation.
3:46 - Canary Testing: Validating new models safely using native 90/10 traffic splitting.
4:49 - Cost Optimization: Using custom HTTP headers to guarantee SLAs and reserve compute for mission-critical queries.
🔗 Helpful Links & Resources:
NGINX Gateway Fabric Documentation: [https://docs.nginx.com/nginx-gateway-fabric]
GitHub Repository / Lab Guides: [https://github.com/nginx/nginx-gateway-fabric]
Notes:
- Contributed by: Akash Ananthanarayanan
- Related Article: None
⬇️⬇️ JOIN THE COMMUNITY! ⬇️⬇️⬇️
DevCentral is an online community of technical peers dedicated to learning, exchanging ideas, and solving problems—together.
Find all our platform links ⬇️ and follow our Community Evangelists! 👋
➡️ DEVCENTRAL: https://community.f5.com
➡️ YOUTUBE: https://youtube.com/devcentral
➡️ LINKEDIN: https://www.linkedin.com/showcase/f5-devcentral/
➡️ X: https://x.com/devcentral
Your Community Evangelists:
👋 Jason Rahm: https://www.linkedin.com/in/jrahm/ | https://x.com/jasonrahm
👋 Buu Lam: https://www.linkedin.com/in/buulam/ | https://x.com/buulam
👋 Chase Abbott: https://www.linkedin.com/in/chaseabbott1 Optimizing AI Inference using NGINX Gateway Fabric](https://i.ytimg.com/vi/TReqi7V8yoI/mqdefault.jpg)






