Uploaded September 2025 | Updated September 2026, 2 hours ago
Efficiency breakthroughs don’t come often, but this one looks real!
Alibaba just dropped Qwen3-Next-80B-A3B, an 80B MoE model with only 3B active parameters at inference. The result? Really fast generation at 10% of the cost of a small model (eg. Qwen 32B).
What’s new here:
• Ultra-sparse MoE with 512 experts (Instruct + Thinking variants).
• Hybrid attention with a gated mechanism, giving speedups where most models slow down.
• Multi-Token Prediction to supercharge speculative decoding.
• Trained on 15T tokens at a fraction of Qwen3-32B’s cost.
• “Thinking” variant even outperforms Gemini-2.5-Flash-Thinking in benchmarks.
• 256K native context (scalable to 1M).
Why it matters: This model challenges the efficiency ceiling set by Mixtral and DeepSeek, while rivaling Qwen3-235B in reasoning with only 3B actrive parameters(!!). If these claims hold up outside benchmarks, we’re looking at a new efficiency baseline for frontier open LLMs.
Which model do you want me to cover next? Drop it in the comments and I’ll tag you.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#AIresearch #LLMs #ainews #short
Efficiency breakthroughs don’t come often, but this one looks real!
Alibaba just dropped Qwen3-Next-80B-A3B, an 80B MoE model with only 3B active parameters at inference. The result? Really fast generation at 10% of the cost of a small model (eg. Qwen 32B).
What’s new here:
• Ultra-sparse MoE with 512 experts (Instruct + Thinking variants).
• Hybrid attention with a gated mechanism, giving speedups where most models slow down.
• Multi-Token Prediction to supercharge speculative decoding.
• Trained on 15T tokens at a fraction of Qwen3-32B’s cost.
• “Thinking” variant even outperforms Gemini-2.5-Flash-Thinking in benchmarks.
• 256K native context (scalable to 1M).
Why it matters: This model challenges the efficiency ceiling set by Mixtral and DeepSeek, while rivaling Qwen3-235B in reasoning with only 3B actrive parameters(!!). If these claims hold up outside benchmarks, we’re looking at a new efficiency baseline for frontier open LLMs.
Which model do you want me to cover next? Drop it in the comments and I’ll tag you.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#AIresearch #LLMs #ainews #short










