Uploaded March 2026 | Updated September 2026, 2 weeks ago
#VoxxedDaysCERN26
Managing and deploying AI models can often require extensive system configuration and complex software dependencies. RamaLama, a new open-source tool, aims to make working with AI models straightforward by leveraging container technology, making the process "boring"—predictable, reliable, and easy to manage. RamaLama integrates with container engines like Podman and Docker to deploy AI models within containers, eliminating the need for manual configuration and ensuring optimal setup for both CPU and GPU systems.
This talk will introduce RamaLama’s key features, including support for multiple AI model registries (Ollama, Hugging Face, and OCI), simplified commands for running models as chatbots or REST API services, and compatibility with alternative AI runtimes like llama.cpp and vllm. We’ll explore RamaLama’s unique capabilities, such as generating Podman quadlet files for edge deployments and Kubernetes YAML for scalable deployment, demonstrating how it allows developers to transition from local experimentation to production seamlessly. Join us to learn how RamaLama enables frictionless, containerized AI model deployment for developers and system administrators alike.
#VoxxedDaysCERN26
Managing and deploying AI models can often require extensive system configuration and complex software dependencies. RamaLama, a new open-source tool, aims to make working with AI models straightforward by leveraging container technology, making the process "boring"—predictable, reliable, and easy to manage. RamaLama integrates with container engines like Podman and Docker to deploy AI models within containers, eliminating the need for manual configuration and ensuring optimal setup for both CPU and GPU systems.
This talk will introduce RamaLama’s key features, including support for multiple AI model registries (Ollama, Hugging Face, and OCI), simplified commands for running models as chatbots or REST API services, and compatibility with alternative AI runtimes like llama.cpp and vllm. We’ll explore RamaLama’s unique capabilities, such as generating Podman quadlet files for edge deployments and Kubernetes YAML for scalable deployment, demonstrating how it allows developers to transition from local experimentation to production seamlessly. Join us to learn how RamaLama enables frictionless, containerized AI model deployment for developers and system administrators alike.


![[VDBUH2026] Magda Miu & Alin Miu - Engineering Leadership in the Age of AI
Engineering management isn’t a promotion, it’s a career shift from solving technical problems to navigating human complexity. In the AI era, this shift becomes even more challenging as leaders must make high-stakes decisions under uncertainty and guide teams through rapid change.
This talk is a condensed “management lab” for senior engineers and new managers, offering practical frameworks for building psychological safety, running effective feedback systems, coaching for autonomy, and making AI-aware decisions.
You will leave with actionable tools, not theory, to help you lead with clarity, confidence, and strategic impact from day one. [VDBUH2026] Magda Miu & Alin Miu - Engineering Leadership in the Age of AI](https://i.ytimg.com/vi/DAgVdXO4NFM/mqdefault.jpg)





![[VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of Us
Not every organization operates with the hyperscale resources of Anthropic, Google, or OpenAI. For the majority of businesses integrating Large Language Models (LLMs) into their critical paths, the high costs and scarcity of GPU/TPU accelerators present a significant challenge. Striking the balance between performance, availability, scalability, and cost-efficiency is a must.
While Kubernetes is a ubiquitous runtime for modern workloads, deploying LLM inference effectively demands a specialized approach. This session dives deep into practical strategies for optimizing your Kubernetes clusters and LLM Inference workloads to run efficiently and cost effectively. We will explore:
– Container and Model Optimization
– Accelerator Management
– Data & Storage
– Network & Load Balancing
– Observability
Attendees will leave with practical techniques for maximizing cost/performance for LLM inference for their AI-powered applications on Kubernetes. [VDBUH2026] Abdel Sghiouar - Optimizing LLM Inference for the Rest of Us](https://i.ytimg.com/vi/G58PbxBXC8c/mqdefault.jpg)

