Uploaded August 2026 | Updated September 2026, 2 weeks ago
IBM Granite 4.2 benchmarked against Qwen 3.8 27B, Gemma 4, and Ornith 1.5 on a single RTX 5090: we open the GGUF files, work out the KV cache cost per token, measure each model's real usable context, then grade coding, math, and tool calling with plain Python and real test runs.
Granite 4.2 moved back to full attention in every block, and this video measures what that choice costs. You will see why Granite 30B holds 58K of usable context on a 32 GB card while Qwen 3.8 27B reaches 246K at the same size, why Granite 4.2 8B still earns a place with a clean 9 of 9 on the coding test at 214 tokens per second, and how a draft head doubles Qwen's code speed. Made for developers who run local models and want numbers from real hardware, and everything here is measured in one session with the same server flags for all six models.
πΊ Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
β± Chapters:
0:00 Intro: Granite 4.2 vs Gemma 4 vs Qwen 3.8
0:15 Test Setup and Model Groups
1:44 Architecture: Where the Mamba Layers Went
6:23 KV Cache Cost per Token
9:52 Usable Context on a 32 GB Card
11:34 Generation Speed: Tokens per Second
12:58 Token Budget Test: 20 Problems
14:55 Coding Test: 99 Hidden Tests
16:01 Tool Calling Test
16:51 The Reasoning Effort Setting
17:02 Which Model to Run
π Resources:
Full write-up with every table: kgptalkie.com/tutorials/llm-benchmarking/granite-4-2-vs-gemma-4-vs-qwen-3-8-benchmark
πΊ Watch next:
Ornith 1.5 35B vs Qwen 3.8 27B vs Gemma 4: youtube.com/watch?v=zJpl8uml1_s
Qwen 3.8 27B Speed Settings (MTP, KV Cache, Flash Attention): youtube.com/watch?v=QkzEkfIzvBk
π Go deeper β my Udemy course:
Fine Tuning LLM with Hugging Face Transformers for NLP: kgptalkie.com/fine-tuning-llm
If this saved you a download, a like helps more people find the video. Tell me in the comments which model your card runs, and subscribe with the bell for a new local LLM benchmark every week.
#IBMGranite #LocalLLM #Qwen #Gemma
IBM Granite 4.2 benchmarked against Qwen 3.8 27B, Gemma 4, and Ornith 1.5 on a single RTX 5090: we open the GGUF files, work out the KV cache cost per token, measure each model's real usable context, then grade coding, math, and tool calling with plain Python and real test runs.
Granite 4.2 moved back to full attention in every block, and this video measures what that choice costs. You will see why Granite 30B holds 58K of usable context on a 32 GB card while Qwen 3.8 27B reaches 246K at the same size, why Granite 4.2 8B still earns a place with a clean 9 of 9 on the coding test at 214 tokens per second, and how a draft head doubles Qwen's code speed. Made for developers who run local models and want numbers from real hardware, and everything here is measured in one session with the same server flags for all six models.
πΊ Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
β± Chapters:
0:00 Intro: Granite 4.2 vs Gemma 4 vs Qwen 3.8
0:15 Test Setup and Model Groups
1:44 Architecture: Where the Mamba Layers Went
6:23 KV Cache Cost per Token
9:52 Usable Context on a 32 GB Card
11:34 Generation Speed: Tokens per Second
12:58 Token Budget Test: 20 Problems
14:55 Coding Test: 99 Hidden Tests
16:01 Tool Calling Test
16:51 The Reasoning Effort Setting
17:02 Which Model to Run
π Resources:
Full write-up with every table: kgptalkie.com/tutorials/llm-benchmarking/granite-4-2-vs-gemma-4-vs-qwen-3-8-benchmark
πΊ Watch next:
Ornith 1.5 35B vs Qwen 3.8 27B vs Gemma 4: youtube.com/watch?v=zJpl8uml1_s
Qwen 3.8 27B Speed Settings (MTP, KV Cache, Flash Attention): youtube.com/watch?v=QkzEkfIzvBk
π Go deeper β my Udemy course:
Fine Tuning LLM with Hugging Face Transformers for NLP: kgptalkie.com/fine-tuning-llm
If this saved you a download, a like helps more people find the video. Tell me in the comments which model your card runs, and subscribe with the bell for a new local LLM benchmark every week.
#IBMGranite #LocalLLM #Qwen #Gemma










