Uploaded November 2024 | Updated September 2026, 2 weeks ago
Live session on Twitch, 11/19/2024
We first discuss why AWS Graviton CPU instances are a great fit for AI inference, particularly for Small Language Models. To prove our point, we then run inference with Llama-3.1-SuperNova Lite 8B on a small Graviton4 instance, thanks to quantization and llama.cpp.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can become a channel member and enjoy exclusive perks: details at youtube.com/channel/UCVonoXm3SI_Q0ZNHd5JPawA/join
You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Model: huggingface.co/arcee-ai/Llama-3.1-SuperNova-Lite
Llama.cpp: github.com/ggerganov/llama.cpp
Live session on Twitch, 11/19/2024
We first discuss why AWS Graviton CPU instances are a great fit for AI inference, particularly for Small Language Models. To prove our point, we then run inference with Llama-3.1-SuperNova Lite 8B on a small Graviton4 instance, thanks to quantization and llama.cpp.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can become a channel member and enjoy exclusive perks: details at youtube.com/channel/UCVonoXm3SI_Q0ZNHd5JPawA/join
You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Model: huggingface.co/arcee-ai/Llama-3.1-SuperNova-Lite
Llama.cpp: github.com/ggerganov/llama.cpp










