[VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of Us @DevoxxForever
[VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of Us  @DevoxxForever
Uploaded May 2026 | Updated September 2026, 2 weeks ago
Not every organization operates with the hyperscale resources of Anthropic, Google, or OpenAI. For the majority of businesses integrating Large Language Models (LLMs) into their critical paths, the high costs and scarcity of GPU/TPU accelerators present a significant challenge. Striking the balance between performance, availability, scalability, and cost-efficiency is a must.

While Kubernetes is a ubiquitous runtime for modern workloads, deploying LLM inference effectively demands a specialized approach. This session dives deep into practical strategies for optimizing your Kubernetes clusters and LLM Inference workloads to run efficiently and cost effectively. We will explore:



– Container and Model Optimization

– Accelerator Management

– Data & Storage

– Network & Load Balancing

– Observability

Attendees will leave with practical techniques for maximizing cost/performance for LLM inference for their AI-powered applications on Kubernetes.
[VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of UsAccessibility powered by AI by Ramona DomenWho to blame? The AI, The Programmer, or The Prompt? by Makan SepehrifarSpring Boot in the Cloud: Advanced Optimization Deep Dive by Patrick BaumgartnerCompilers, User Interfaces & the Rest of Us by Matheus AlbuquerqueGraalVM 25: Whats New and Whats Next by Alina YurenkoTesting Challenges in the Age of AI by Konstantin PavlovThe record: migrate to immutability by Johan HuttingFriendly Borders: Graph algorithms reveal Eurovision voting patterns by Domagoj MarićCommit to Change: From Analyst to Java Developer by  Janne PloegerCoupling, Cohesion and Change, The Blueprint of Modern Software Design by  Boyen van GorpCoding a Multi-Agent Game Master with Spring AI by Alessandra Pasini and Arnaud Jean
Devoxx |

[VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of Us

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER