Uploaded February 2025 | Updated September 2026, 2 weeks ago
In this video, we run local inference on an Apple M3 MacBook with llama.cpp and MLX, two projects that optimize and accelerate small language models on CPU platforms. For this purpose, we use two new Arcee open-source models distilled from DeepSeek-v3: Virtuoso Lite 10B and Virtuoso Medium v2 32B.
First, we download the two models from the Hugging Face hub with the Hugging Face CLI. Then, we go through the step-by-step installation procedure for llama.cpp and MLX. Next, we optimize and quantize the models to 4-bit precision for maximum acceleration. Finally, we run inference and look at performance numbers. So, who's fastest? Watch and find out :)
If you’d like to understand how Arcee AI can help your organization build scalable and cost-efficient AI solutions, don't hesitate to contact sales@arcee.ai or book a demo at arcee.ai/book-a-demo.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
* Blog post: arcee.ai/blog/virtuoso-lite-virtuoso-medium-v2-distilling-deepseek-v3-into-10b-32b-small-language-models-slms
* Virtuoso Lite: huggingface.co/arcee-ai/Virtuoso-Lite
* Virtuoso Medium v2: huggingface.co/arcee-ai/Virtuoso-Medium-v2
* Llama.cpp: github.com/ggerganov/llama.cpp
* MLX for language models: github.com/ml-explore/mlx-examples
00:00 Introduction
00:40 A quick look at Virtuoso-Lite and Virtuoso-Medium-v2
03:10 Downloading the models from Hugging Face
04:30 Building llama.cpp from source
06:15 Converting Hugging Face models to GGUF
09:20 Quantizing models with llama.cpp
11:20 Running local inference with llama.cpp
15:00 Installing MLX
11:20 Running local inference with MLX
19:45 Conclusion
In this video, we run local inference on an Apple M3 MacBook with llama.cpp and MLX, two projects that optimize and accelerate small language models on CPU platforms. For this purpose, we use two new Arcee open-source models distilled from DeepSeek-v3: Virtuoso Lite 10B and Virtuoso Medium v2 32B.
First, we download the two models from the Hugging Face hub with the Hugging Face CLI. Then, we go through the step-by-step installation procedure for llama.cpp and MLX. Next, we optimize and quantize the models to 4-bit precision for maximum acceleration. Finally, we run inference and look at performance numbers. So, who's fastest? Watch and find out :)
If you’d like to understand how Arcee AI can help your organization build scalable and cost-efficient AI solutions, don't hesitate to contact sales@arcee.ai or book a demo at arcee.ai/book-a-demo.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
* Blog post: arcee.ai/blog/virtuoso-lite-virtuoso-medium-v2-distilling-deepseek-v3-into-10b-32b-small-language-models-slms
* Virtuoso Lite: huggingface.co/arcee-ai/Virtuoso-Lite
* Virtuoso Medium v2: huggingface.co/arcee-ai/Virtuoso-Medium-v2
* Llama.cpp: github.com/ggerganov/llama.cpp
* MLX for language models: github.com/ml-explore/mlx-examples
00:00 Introduction
00:40 A quick look at Virtuoso-Lite and Virtuoso-Medium-v2
03:10 Downloading the models from Hugging Face
04:30 Building llama.cpp from source
06:15 Converting Hugging Face models to GGUF
09:20 Quantizing models with llama.cpp
11:20 Running local inference with llama.cpp
15:00 Installing MLX
11:20 Running local inference with MLX
19:45 Conclusion






![How Witty Works leverages Hugging Face to scale inclusive language
During this webinar, Elena Nazarenko, Lead Data Scientist at Witty Works, Lukas Kahwe Smith, CTO & Co-Founder at Witty Works and Julien Simon, Chief Evangelist at Hugging Face, discuss how Witty Works leverages Hugging Face to scale inclusive language.
[No HD version, sorry]
- The impact of Transformers on text classification use cases
- How Witty Works leverages Hugging Face to scale inclusive language
- How to perform domain-adaptive pretraining on a transformer model
Speakers
Elena Nazarenko - Lead Data Scientist at Witty Works
Lukas Kahwe Smith - CTO & Co-Founder at Witty Works
Julien Simon - Chief Evangelist at Hugging Face
About Witty Works
Witty is a Digital Writing Assistant for Inclusive Language that enables organizations to detect their own bias, in writing and in behavior, and fix it. Because language builds culture.
About Hugging Face
Hugging Face is a wildly popular community-based repository for open-source ML technology. It is a platform that stores, serves and manages the latest and greatest in open-sources ML models, including enabling customers to fine-tune these models and deploy them at scale.
Hugging Face is one of the most used platforms and is empowering 10,000 companies to integrate artificial intelligence into their products or workflows. How Witty Works leverages Hugging Face to scale inclusive language](https://i.ytimg.com/vi/Z_S1gfRFtgA/mqdefault.jpg)



