Uploaded March 2024 | Updated September 2026, 2 weeks ago
In this video, I walk you through the simple process of deploying a Hugging Face large language model on AWS, with Amazon SageMaker and the AWS Inferentia2 accelerator.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Notebook:
gitlab.com/juliensimon/huggingface-demos/-/blob/main/inferentia2/llm/deploy_zephyr_inf2.ipynb
Deep Dive: Hugging Face models on AWS AI Accelerators
youtu.be/66JUlAA8nOU
Blog posts:
huggingface.co/blog/how-to-generate
aws.amazon.com/blogs/machine-learning/elevating-the-generative-ai-experience-introducing-streaming-support-in-amazon-sagemaker-hosting
In this video, I walk you through the simple process of deploying a Hugging Face large language model on AWS, with Amazon SageMaker and the AWS Inferentia2 accelerator.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Notebook:
gitlab.com/juliensimon/huggingface-demos/-/blob/main/inferentia2/llm/deploy_zephyr_inf2.ipynb
Deep Dive: Hugging Face models on AWS AI Accelerators
youtu.be/66JUlAA8nOU
Blog posts:
huggingface.co/blog/how-to-generate
aws.amazon.com/blogs/machine-learning/elevating-the-generative-ai-experience-introducing-streaming-support-in-amazon-sagemaker-hosting










