Deploying SuperNova-Lite on Inferentia2: the best 8B model for $1 an hour! @juliensimonfr
Deploying SuperNova-Lite on Inferentia2: the best 8B model for $1 an hour!  @juliensimonfr
Uploaded September 2024 | Updated September 2026, 2 weeks ago
In this video, you will learn about Llama-3.1-SuperNova-Lite, the best open-source 8B model available today.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

Llama-3.1-SuperNova-Lite is an 8B parameter model developed by Arcee.ai, based on the Llama-3.1-8B-Instruct architecture. It is a distilled version of the larger Llama-3.1-405B-Instruct model, leveraging offline logits extracted from the 405B parameter variant. This 8B variation of Llama-3.1-SuperNova maintains high performance while offering exceptional instruction-following capabilities and domain-specific adaptability.

I'll show you how to compile on the fly and deploy SuperNova Lite on a SageMaker endpoint powered by an inf2.xlarge instance, the smallest Inferentia2 instance available at only $0.99 an hour!

00:00 Introduction
01:07 SuperNova and SuperNova-Lite
02:15 SuperNova-Lite, the number 1 8B model on the Hugging Face leaderboard
02:50 Developer resources to deploy Arcee models on AWS
03:30 Sample notebook to deploy SuperNova-Lite
04:30 The LMI container
07:30 Deploying the model and compiling in on the fly
13:11 Looking at the deployment log in CloudWatch
16:05 Running synchronous inference
18:30 Running streaming inference
20:40 Cleaning up and conclusion

* Blog post: blog.arcee.ai/meet-arcee-supernova-our-flagship-70b-model-alternative-to-openai
* Hugging Face leaderboard: huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard
* Model page: huggingface.co/arcee-ai/Llama-3.1-SuperNova-Lite
* Developer resources: github.com/arcee-ai/aws-samples
* Sample notebook: github.com/arcee-ai/aws-samples/blob/main/model_notebooks/sample-notebook-llama-supernova-lite-on-sagemaker-inf2.ipynb
* LMI container: docs.djl.ai/master/docs/serving/serving/docs/lmi/index.html
* LMI on NeuronX devices: docs.djl.ai/master/docs/serving/serving/docs/lmi/user_guides/tnx_user_guide.html

#ai #aws #slm #llm #openai #chatgpt #opensource #huggingface
Deploying SuperNova-Lite on Inferentia2: the best 8B model for $1 an hour!Parameter-efficient fine-tuning with QLoRA and Hugging FaceAWS User Group DubaiInterview BFM Business - Hugging Face (04/2023)What is Vibe Coding? Build based on intent!Deep Dive: Quantizing Large Language Models, part 2Arcee Llama Spark, a better Llama 3.1 #ai #largelanguagemodels #chatbot #opensourceUnderstanding AI Risk: Beyond the Surface in Enterprise!Unpacking the Complex World of Risk Management in AI – It’s Not What You Think!Unlocking the Secret to Impactful AI: Its All About Quality!Arcee AI live webinar - 18/09/2025Deep Dive: Optimizing LLM inference
Julien Simon |

Deploying SuperNova-Lite on Inferentia2: the best 8B model for $1 an hour!

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER