Deep Dive: How Three MoE Reasoning Models Actually Work — Trinity, DeepSeek R1, Kimi K2 - Part 2 @juliensimonfr
Deep Dive: How Three MoE Reasoning Models Actually Work — Trinity, DeepSeek R1, Kimi K2 - Part 2  @juliensimonfr
Uploaded May 2026 | Updated September 2026, 2 weeks ago
Three frontier open-weight MoE reasoning models — Trinity Large Thinking (~400B), DeepSeek R1 (671B), and Kimi K2 Thinking (1T) — are compared side by side. Architecture, training, and post-training, explained from first principles.

⭐️⭐️⭐️ More content on Substack at airealist.ai ⭐️⭐️⭐️

In Part 1 (youtu.be/2uQQ8nKNq1U), I broke down how these three models are actually built — not benchmarks, not vibes, but the engineering decisions and why they matter.

Now the question that actually matters: which one should you use? Benchmarks, costs, deployment, and practical details.

*** Models

Trinity Large Thinking: ~400B total, ~13B active, 256 experts, 512K context, Apache 2.0 huggingface.co/arcee-ai/Trinity-Large-Thinking NVFP4: huggingface.co/arcee-ai/Trinity-Large-Thinking-NVFP4

DeepSeek R1: 671B total, ~37B active, 256 experts, 128K context, MIT huggingface.co/deepseek-ai/DeepSeek-R1

Kimi K2 Thinking: 1T total, ~32B active, 384 experts, 256K context, Modified MIT huggingface.co/moonshotai/Kimi-K2-Instruct

*** Papers & blogs

Trinity architecture blog: arcee.ai/blog/trinity-large
Trinity tech report: arxiv.org/abs/2602.17004
OpenRouter (all three): openrouter.ai

#llm #moe #deepseek #reasoning #architecture #training #openweight #inference #arcee #kimi #trinity #mixtureofexperts
Deep Dive: How Three MoE Reasoning Models Actually Work — Trinity, DeepSeek R1, Kimi K2 - Part 2Arcee.ai and Small Language Models - Mobile World Congress, Las Vegas (10/2024)LLMs from the trenches - LLMs are not intelligent, there is no reasoningUnlocking the Future of Finance with AI! 🧠💰Discover the Power of Arcee-Lite: Language Modeling Redefined!3 production-ready models released by Arcee AI on Hugging FaceDeploying Llama3 with Inference Endpoints and AWS Inferentia2The Ease of Running Small Language Models LocallyDeploy Hugging Face models on Google Cloud: from the hub to Vertex AI🚀 Unveiling the Future: Meet Arcee Nova, the Ultimate LLM! 🌟Open Source AI with Hugging Face - Dallas AI  meetup (05/2024)Decoder-only inference: a step-by-step deep dive
Julien Simon |

Deep Dive: How Three MoE Reasoning Models Actually Work — Trinity, DeepSeek R1, Kimi K2 - Part 2

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER