Run SLMs locally: Llama.cpp vs. MLX with 10B and 32B Arcee models @juliensimonfr
Run SLMs locally: Llama.cpp vs. MLX with 10B and 32B Arcee models  @juliensimonfr
Uploaded February 2025 | Updated September 2026, 2 weeks ago
In this video, we run local inference on an Apple M3 MacBook with llama.cpp and MLX, two projects that optimize and accelerate small language models on CPU platforms. For this purpose, we use two new Arcee open-source models distilled from DeepSeek-v3: Virtuoso Lite 10B and Virtuoso Medium v2 32B.

First, we download the two models from the Hugging Face hub with the Hugging Face CLI. Then, we go through the step-by-step installation procedure for llama.cpp and MLX. Next, we optimize and quantize the models to 4-bit precision for maximum acceleration. Finally, we run inference and look at performance numbers. So, who's fastest? Watch and find out :)

If you’d like to understand how Arcee AI can help your organization build scalable and cost-efficient AI solutions, don't hesitate to contact sales@arcee.ai or book a demo at arcee.ai/book-a-demo.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

* Blog post: arcee.ai/blog/virtuoso-lite-virtuoso-medium-v2-distilling-deepseek-v3-into-10b-32b-small-language-models-slms
* Virtuoso Lite: huggingface.co/arcee-ai/Virtuoso-Lite
* Virtuoso Medium v2: huggingface.co/arcee-ai/Virtuoso-Medium-v2
* Llama.cpp: github.com/ggerganov/llama.cpp
* MLX for language models: github.com/ml-explore/mlx-examples

00:00 Introduction
00:40 A quick look at Virtuoso-Lite and Virtuoso-Medium-v2
03:10 Downloading the models from Hugging Face
04:30 Building llama.cpp from source
06:15 Converting Hugging Face models to GGUF
09:20 Quantizing models with llama.cpp
11:20 Running local inference with llama.cpp
15:00 Installing MLX
11:20 Running local inference with MLX
19:45 Conclusion
Run SLMs locally: Llama.cpp vs. MLX with 10B and 32B Arcee modelsEnterprise AI with the Hugging Face Enterprise HubRouting function calling queries to the best SLM/LLM with Arcee ConductorUnlocking Hidden Treasures: Your Data is More Valuable Than You Think! 💎Is AI the answer to every problem? 🤔 Think again!Retail AI at the edge - Cisco Live 2025 interviewArcee Orchestra - Build an Agentic Code Review WorkflowHow Witty Works leverages Hugging Face to scale inclusive languageSLM Inference on a Windows laptop 🤯 Intel Lunar Lake CPU/GPU/NPU + OpenVINOFine-tune Stable Diffusion with LoRA for as low as $1🚀 The AI Revolution is Happening Fast! Are You Keeping Up? 💡Build Mobile Apps with Replit in Minutes!
Julien Simon |

Run SLMs locally: Llama.cpp vs. MLX with 10B and 32B Arcee models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER