Uploaded August 2026 | Updated September 2026, 2 weeks ago
Qwen 3.8 Flash Next vs Qwen 3.8 27B, a full teardown of the Qwen 4 architecture with every config difference, the real memory numbers, and an honest answer about what you can run at home.
Qwen shipped Flash Next as a preview of Qwen 4, and it is a very different model from the dense 27B most of us already have on disk. I open both configs side by side and go through the changes that matter: 512 experts replacing the feed forward layer, sparse attention that shrinks the KV cache from 64 KB to 24 KB per token, LayerNorms folded into what Qwen now calls the gated residual, and a 51B n-gram lookup table you can keep in RAM or on SSD instead of VRAM. I also show the MTP numbers I measured myself on the RTX 5090, where the 27B went from 77 to over 180 tokens per second. One thing I want to be straight about: Flash Next needs 360 GB in bf16, so I could not run it on my card and I am not going to pretend otherwise. Every Flash Next score you see here is Qwen's published number. This video is the architecture. The speed settings and the head to head coding tests are separate videos on the channel.
πΊ Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
β± Chapters:
0:00 Intro
0:53 Six months of Qwen on the same Qwen 3.5 architecture
2:13 125B total, 51B lookup table, 6B active
2:52 Why Flash Next will not load on an RTX 5090
3:42 What Flash Next keeps from Qwen 3.8 27B
4:51 Blocks and hidden size: 48 vs 64, 5120 to 2560
8:22 Mixture of Experts replaces the feed forward layer
12:34 How many experts fire per token
13:58 Where the 6B active parameters come from
17:01 Sparse attention and why the KV cache decides speed
19:01 KV cache per token: 24 KB vs 64 KB
21:22 The LayerNorm is gone, meet the gated residual
23:01 What the 51B n-gram lookup table actually does
25:41 MTP draft head: 77 to 180+ tokens per second
26:59 Qwen's benchmarks and where Flash Next wins
28:09 Memory at 262k context: 11 GB vs 66 GB
29:25 Which model should you run today
π Full written breakdown with all the numbers:
kgptalkie.com/tutorials/generative-ai/qwen-3-8-flash-next-vs-qwen-3-8-27b
πΊ Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention
youtube.com/watch?v=QkzEkfIzvBk&list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
π Go deeper, my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents
kgptalkie.com/langchain
If this cleared up what Qwen 4 is doing, hit like so more local AI folks find it, and tell me in the comments which part you want me to test the moment a smaller Qwen 4 model drops. I read every comment and I reply to every one of them. Subscribe and turn on the bell if you want the rest of these teardowns.
#Qwen #LocalLLM #Qwen4 #RTX5090 #GenerativeAI
Qwen 3.8 Flash Next vs Qwen 3.8 27B, a full teardown of the Qwen 4 architecture with every config difference, the real memory numbers, and an honest answer about what you can run at home.
Qwen shipped Flash Next as a preview of Qwen 4, and it is a very different model from the dense 27B most of us already have on disk. I open both configs side by side and go through the changes that matter: 512 experts replacing the feed forward layer, sparse attention that shrinks the KV cache from 64 KB to 24 KB per token, LayerNorms folded into what Qwen now calls the gated residual, and a 51B n-gram lookup table you can keep in RAM or on SSD instead of VRAM. I also show the MTP numbers I measured myself on the RTX 5090, where the 27B went from 77 to over 180 tokens per second. One thing I want to be straight about: Flash Next needs 360 GB in bf16, so I could not run it on my card and I am not going to pretend otherwise. Every Flash Next score you see here is Qwen's published number. This video is the architecture. The speed settings and the head to head coding tests are separate videos on the channel.
πΊ Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
β± Chapters:
0:00 Intro
0:53 Six months of Qwen on the same Qwen 3.5 architecture
2:13 125B total, 51B lookup table, 6B active
2:52 Why Flash Next will not load on an RTX 5090
3:42 What Flash Next keeps from Qwen 3.8 27B
4:51 Blocks and hidden size: 48 vs 64, 5120 to 2560
8:22 Mixture of Experts replaces the feed forward layer
12:34 How many experts fire per token
13:58 Where the 6B active parameters come from
17:01 Sparse attention and why the KV cache decides speed
19:01 KV cache per token: 24 KB vs 64 KB
21:22 The LayerNorm is gone, meet the gated residual
23:01 What the 51B n-gram lookup table actually does
25:41 MTP draft head: 77 to 180+ tokens per second
26:59 Qwen's benchmarks and where Flash Next wins
28:09 Memory at 262k context: 11 GB vs 66 GB
29:25 Which model should you run today
π Full written breakdown with all the numbers:
kgptalkie.com/tutorials/generative-ai/qwen-3-8-flash-next-vs-qwen-3-8-27b
πΊ Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention
youtube.com/watch?v=QkzEkfIzvBk&list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
π Go deeper, my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents
kgptalkie.com/langchain
If this cleared up what Qwen 4 is doing, hit like so more local AI folks find it, and tell me in the comments which part you want me to test the moment a smaller Qwen 4 model drops. I read every comment and I reply to every one of them. Subscribe and turn on the bell if you want the rest of these teardowns.
#Qwen #LocalLLM #Qwen4 #RTX5090 #GenerativeAI










