Deploying Hugging Face models with Amazon SageMaker and AWS Inferentia2 @juliensimonfr
Deploying Hugging Face models with Amazon SageMaker and AWS Inferentia2  @juliensimonfr
Uploaded March 2024 | Updated September 2026, 2 weeks ago
In this video, I walk you through the simple process of deploying a Hugging Face large language model on AWS, with Amazon SageMaker and the AWS Inferentia2 accelerator.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

Notebook:
gitlab.com/juliensimon/huggingface-demos/-/blob/main/inferentia2/llm/deploy_zephyr_inf2.ipynb

Deep Dive: Hugging Face models on AWS AI Accelerators
youtu.be/66JUlAA8nOU

Blog posts:
huggingface.co/blog/how-to-generate
aws.amazon.com/blogs/machine-learning/elevating-the-generative-ai-experience-introducing-streaming-support-in-amazon-sagemaker-hosting
Deploying Hugging Face models with Amazon SageMaker and AWS Inferentia2Building a RAG chatbot with LangChain, Chroma, Hugging Face, and Arcee ConductorNo Cloud, No API Keys: Local Open-Source Coding with Trinity Mini, OpenCode, and MLXHugging Face / AWS roadshow - Day 2, MadridAI at the edge - live from Cisco Live in San Diego, CA!Arcee Agent, a 7B model for function calls and tools #ai #largelanguagemodels #chatbot #opensourceThe Stunning Speed of Local Small Language ModelsArcee Blitz 24B - A better Mistral Small 3 modelCreative Writing with AI: A New FrontierDeep Dive: Model Distillation with DistillKitIntroducing the Arcee AI Trinity ModelsArcee AI webinar: routing your function calling and reasoning queries with Arcee Conductor
Julien Simon |

Deploying Hugging Face models with Amazon SageMaker and AWS Inferentia2

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER