Uploaded August 2026 | Updated September 2026, 2 weeks ago
Qwen 3.5, Qwen 3.6 and Qwen 3.8 are the same architecture, so I ran 397 measured generations on one RTX 5090 to find out where the improvement is actually coming from.
Qwen 3.8 27B vs Qwen 3.6 27B vs Qwen 3.5 27B Teardown
kgptalkie.com/tutorials/generative-ai/qwen-3-8-27b-vs-qwen-3-6-27b-vs-qwen-3-5-27b
This is a teardown, not a vibe check. I open all three models and compare layers, hidden size, attention heads, vocabulary and parameter count, and they match. Then I show you what really changed: an MTP layer that was sitting dormant in 3.5 and 3.6, a vision encoder split out of the main weights, and a lot of post-training. Along the way you get the reasoning-token numbers, the timing per problem, why token per second figures disagree everywhere on the internet, and what actually happens when you fill a context window. Everything is reproducible from the setup I show at the start.
📺 Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
⏱ Chapters:
0:00 What changed between Qwen 3.5, 3.6 and 3.8
1:10 My setup: RTX 5090, Ollama, 4-bit, Windows 11
1:55 397 measured generations
2:46 There is no architectural difference
3:21 Same 64 layers, hidden size, heads and vocabulary
5:09 The two real changes: MTP layer and split vision encoder
5:27 Benchmark scores across the releases
6:41 Terminal-Bench 41 to 73: the agentic jump
7:50 How I verified it: 12 deterministic problems
9:02 Thinking tokens: 1K vs 5K to 7K
11:00 Time per problem: 23 minutes vs 3.5 minutes
13:27 Why the older models burn tokens
14:17 Qwen 3.5 had the answer at 12% and kept thinking
16:03 Reasoning effort: Ollama vs LM Studio
16:58 Why token per second numbers disagree everywhere
18:55 205 tok/s counting, 190 code, 176 JSON
19:49 How MTP accepts and rejects draft tokens
23:05 Base speed without MTP: 80 tok/s
24:01 Why Qwen 3.5 and 3.6 never enabled MTP
25:46 What happens when you fill the context window
27:38 16K, 48K, 80K: decode and prefill both drop
30:00 Where the gains really come from: post-training
🔗 Resources:
Full written teardown: kgptalkie.com/tutorials/generative-ai/qwen-3-8-27b-vs-qwen-3-6-27b-vs-qwen-3-5-27b
Qwen 3.8 on Ollama: ollama.com/library/qwen3.8
Setup: RTX 5090 32GB capped at 600W, Ollama on Windows 11, 4-bit quantization, 397 measured generations, 12 verified deterministic problems
📺 Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention: youtube.com/watch?v=QkzEkfIzvBk
🎓 Go deeper with my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents: kgptalkie.com/langchain
RTX 5090 image via Wikimedia Commons, CC BY 3.0
If this was useful, hit like and leave a comment, because I read every one and I reply. Subscribe and turn on the bell if you want the rest of the local model teardowns.
#Qwen3 #LocalLLM #Ollama #RTX5090
Qwen 3.5, Qwen 3.6 and Qwen 3.8 are the same architecture, so I ran 397 measured generations on one RTX 5090 to find out where the improvement is actually coming from.
Qwen 3.8 27B vs Qwen 3.6 27B vs Qwen 3.5 27B Teardown
kgptalkie.com/tutorials/generative-ai/qwen-3-8-27b-vs-qwen-3-6-27b-vs-qwen-3-5-27b
This is a teardown, not a vibe check. I open all three models and compare layers, hidden size, attention heads, vocabulary and parameter count, and they match. Then I show you what really changed: an MTP layer that was sitting dormant in 3.5 and 3.6, a vision encoder split out of the main weights, and a lot of post-training. Along the way you get the reasoning-token numbers, the timing per problem, why token per second figures disagree everywhere on the internet, and what actually happens when you fill a context window. Everything is reproducible from the setup I show at the start.
📺 Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
⏱ Chapters:
0:00 What changed between Qwen 3.5, 3.6 and 3.8
1:10 My setup: RTX 5090, Ollama, 4-bit, Windows 11
1:55 397 measured generations
2:46 There is no architectural difference
3:21 Same 64 layers, hidden size, heads and vocabulary
5:09 The two real changes: MTP layer and split vision encoder
5:27 Benchmark scores across the releases
6:41 Terminal-Bench 41 to 73: the agentic jump
7:50 How I verified it: 12 deterministic problems
9:02 Thinking tokens: 1K vs 5K to 7K
11:00 Time per problem: 23 minutes vs 3.5 minutes
13:27 Why the older models burn tokens
14:17 Qwen 3.5 had the answer at 12% and kept thinking
16:03 Reasoning effort: Ollama vs LM Studio
16:58 Why token per second numbers disagree everywhere
18:55 205 tok/s counting, 190 code, 176 JSON
19:49 How MTP accepts and rejects draft tokens
23:05 Base speed without MTP: 80 tok/s
24:01 Why Qwen 3.5 and 3.6 never enabled MTP
25:46 What happens when you fill the context window
27:38 16K, 48K, 80K: decode and prefill both drop
30:00 Where the gains really come from: post-training
🔗 Resources:
Full written teardown: kgptalkie.com/tutorials/generative-ai/qwen-3-8-27b-vs-qwen-3-6-27b-vs-qwen-3-5-27b
Qwen 3.8 on Ollama: ollama.com/library/qwen3.8
Setup: RTX 5090 32GB capped at 600W, Ollama on Windows 11, 4-bit quantization, 397 measured generations, 12 verified deterministic problems
📺 Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention: youtube.com/watch?v=QkzEkfIzvBk
🎓 Go deeper with my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents: kgptalkie.com/langchain
RTX 5090 image via Wikimedia Commons, CC BY 3.0
If this was useful, hit like and leave a comment, because I read every one and I reply. Subscribe and turn on the bell if you want the rest of the local model teardowns.
#Qwen3 #LocalLLM #Ollama #RTX5090










