Scaling LLM Inference: AWS Inferentia Meets Ray Serve on EKS | Ray Summit 2024 @anyscale
Scaling LLM Inference: AWS Inferentia Meets Ray Serve on EKS | Ray Summit 2024  @anyscale
Uploaded October 2024 | Updated September 2026, 3 weeks ago
The race for efficient, scalable AI inference is on, and AWS is at the forefront with innovative solutions. This session showcases how to achieve high-performance, cost-effective inference for large language models like Llama2 and Mistral-7B using Ray Serve and AWS Inferentia on Amazon EKS.

Vara Bonthu and Ratnopam Chakrabarti will guide you through the intricacies of building a scalable inference infrastructure that bypasses GPU availability constraints. They'll demonstrate how the synergy between Ray Serve, AWS Neuron SDK, and Karpenter autoscaler on Amazon EKS creates a powerful, flexible environment for AI workloads. Attendees will explore strategies for optimizing costs while maintaining high performance, opening new possibilities for deploying and scaling advanced language models in production environments.

--

Interested in more?
- Watch the full Day 1 Keynote: youtu.be/jwZHJthQvXo
- Watch the full Day 2 Keynote youtu.be/Lury2ad6KG8
- Check out the Ray Summmit Breakout sessions youtube.com/playlist?list=PLzTswPQNepXntmT8jr9WaNfqQ60QwW7-U&si=qPw-_SxT9lVmbRGE

--

🔗 Connect with us:
- Subscribe to our YouTube channel: youtube.com/@anyscale
- Twitter: https://x.com/anyscalecompute
- LinkedIn: linkedin.com/company/joinanyscale
- Website: anyscale.com
Scaling LLM Inference: AWS Inferentia Meets Ray Serve on EKS | Ray Summit 2024Inside NVIDIA Dynamo: Faster, Scalable AI Deployment | Ray Summit 2025Ray + Kubernetes: The Distributed OS for AI/ML | Ray on the Road – NYC 2025Ray Summit 2025 Keynote: Physical AI Turing Test with Jim Fan from NVIDIABen Horowitz - Historical Perspectives on AI and the Internet | Ray Summit 2023Ray Joins The Linux Foundation & PyTorch Sub-Foundation: Toward a Unified AI Compute StackHow Torc Robotics Scales Multimodal AI for Autonomous Driving with RayRay on Kubernetes: Powering Quant Research at Scale | Ray Summit 2024How the VAST AI Operating System Powers a Dynamic Data Plane for Ray | Ray Summit 2025Scaling Ray Train to 10K Kubernetes Nodes on GKE | Ray Summit 2024Motional’s Blueprint for High-Performance ML Systems in Autonomous Driving | Ray Summit 2025How Roblox Scaled Machine Learning by Leveraging Ray for Efficient Batch Inference | Ray Summit 2024
Anyscale |

Scaling LLM Inference: AWS Inferentia Meets Ray Serve on EKS | Ray Summit 2024

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER