Homunculus 12B and GLM-4-32B-Base-32K: 2 new Arcee AI research-oriented models @juliensimonfr
Homunculus 12B and GLM-4-32B-Base-32K: 2 new Arcee AI research-oriented models  @juliensimonfr
Uploaded July 2025 | Updated September 2026, 2 weeks ago
In this video, I introduce two new research-oriented models that Arcee AI recently released on Hugging Face.

Homunculus is a 12 billion-parameter instruction model distilled from Qwen3-235B onto the Mistral-Nemo backbone. It was purpose-built to preserve Qwen’s two-mode interaction style—/think (deliberate chain-of-thought) and /nothink (concise answers)—while running on a single consumer GPU, and even on CPU as demonstrated in the video.

GLM-4-32B-Base-32K is an enhanced version of THUDM's GLM-4-32B-Base-0414, specifically engineered to offer robust performance over an extended context window. While the original model's capabilities degraded after 8,192 tokens, this version maintains strong performance up to a 32,000-token context, making it ideal for tasks requiring long-context understanding and processing.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. You can also follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️

** Homunculus
- huggingface.co/arcee-ai/Homunculus
- huggingface.co/arcee-ai/Homunculus-GGUF

bin/llama-cli -m ~/models/homunculus/Homunculus-Q4_K_M.gguf --color -c 65535

"Looking at multi-head attention, group-query attention, multi-query attention, and multi-head latent attention, which method would optimize inference latency for a small language model with 32 attention layers running on a 64-core Intel CPU?"

** GLM-4-32B-Base-32K
- huggingface.co/arcee-ai/GLM-4-32B-Base-32K
- huggingface.co/bartowski/arcee-ai_GLM-4-32B-Base-32K-GGUF
- arcee.ai/blog/extending-afm-4-5b-to-64k-context-length

⭐️⭐️⭐️ While you're here, I’ve got a great deal for you! If you care about your online security, you need Proton Pass — the ultra-secure password manager from the creators of Proton Mail. GET 60% OFF at go.getproton.me/aff_c?offer_id=42&aff_id=13055&url_id=994 ⭐️⭐️⭐️
Homunculus 12B and GLM-4-32B-Base-32K: 2 new Arcee AI research-oriented modelsDeploy Hugging Face models on Google Cloud: from the hub to Inference EndpointsSLM in Action: Local Inference with Arcee Nova 72B and OllamaPhi-2 on Intel Meteor Lake - Physics questionVirtuoso Lite and Virtuoso Medium v2: distilling DeepSeek-V3 to 10B & 32BUnlock the Hidden Value of Your Data! 🌟Benchmarking TurboQuant with MLX on Apple SiliconUnlocking the Power of AI: How Chatbots Enhance Communication!Deep dive: model merging, part 2Arcee AI Drops Trinity Nano & Mini MoE Models!Unlock the Secret Goldmine of Data!SLM in Action: Arcee-Scribe, a 7.7B model for creative writing
Julien Simon |

Homunculus 12B and GLM-4-32B-Base-32K: 2 new Arcee AI research-oriented models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER