Run performant and cost-effective GenAI Applications with AWS Graviton and Arcee AI @juliensimonfr
Run performant and cost-effective GenAI Applications with AWS Graviton and Arcee AI  @juliensimonfr
Uploaded November 2024 | Updated September 2026, 2 weeks ago
Live session on Twitch, 11/19/2024

We first discuss why AWS Graviton CPU instances are a great fit for AI inference, particularly for Small Language Models. To prove our point, we then run inference with Llama-3.1-SuperNova Lite 8B on a small Graviton4 instance, thanks to quantization and llama.cpp.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can become a channel member and enjoy exclusive perks: details at youtube.com/channel/UCVonoXm3SI_Q0ZNHd5JPawA/join
You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

Model: huggingface.co/arcee-ai/Llama-3.1-SuperNova-Lite
Llama.cpp: github.com/ggerganov/llama.cpp
Run performant and cost-effective GenAI Applications with AWS Graviton and Arcee AIHugging Face, the story so farArcee Nova 72B running on my Mac #ai #largelanguagemodels #opensourceThis model costs 15 cents per million tokens!How Synapse Medicine leverages Hugging Face to improve medication safetyThe AI Race: From Llama 2 to Llama 3 and Beyond!Arcee AI webinar: pick the right SLM/LLM for each query with Arcee ConductorTransformer training shootout, part 2: AWS Trainium vs. NVIDIA V100The Truth About MCP: Pros, Cons & Real-World Use CasesRun SLMs locally: Llama.cpp vs. MLX with 10B and 32B Arcee modelsEnterprise AI with the Hugging Face Enterprise HubRouting function calling queries to the best SLM/LLM with Arcee Conductor
Julien Simon |

Run performant and cost-effective GenAI Applications with AWS Graviton and Arcee AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER