Uploaded August 2026 | Updated September 2026, 2 weeks ago
Ornith 1.5 9B and 35B benchmarked against Qwen 3.8 27B and Gemma 4 on a single RTX 5090, including the KV cache math that decides which of these models actually fits on your GPU.
I ran all four models on the same rig with llama-server, 4 bit weights and FP16 KV cache, and graded every answer with code instead of asking another model to judge. You will see generation speed, prompt processing speed, VRAM and KV cache per token, how speed holds up as the context window fills, a 128K needle test, parallel throughput at 8 concurrent requests, and tool call accuracy for agentic coding. I also opened the GGUF files and went through the tensors block by block, which is where the most interesting finding of this video came from. Everything you need is in this one video, no prior parts required.
📺 Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
⏱ Chapters:
0:00 Intro: Ornith 1.5 9B and 35B
0:58 Dense vs MoE and the 262K context window
1:54 Why I test on RTX 5090
3:19 Setup: llama-server, 4-bit, FP16 KV cache and code grading
4:12 Active parameters per token explained
5:52 Inside the GGUF: Ornith is built on Qwen 3.5
7:16 256 experts, 8 active per token
8:07 Attention on every 4th block and gated DeltaNet
9:50 40 blocks plus the MTP head
11:23 Why Gemma 4 uses more active parameters
13:00 Generation speed: 271 vs 192 vs 77 tokens per second
15:19 Gemma 4 MoE at 234 tokens per second
17:40 Prompt processing: Ornith 9B is the fastest
20:01 VRAM and KV cache per token
21:31 Why Qwen 3.8 27B cannot hold the full 262K context
23:26 Speed as the context window fills up
25:12 128K needle test: 75 seconds vs 23 seconds
26:17 MTP explained and why it slows Ornith down
29:55 Parallel throughput: 702 vs 673 tokens per second
32:08 Same task: time and tokens used
33:29 Tool call efficiency for agentic coding
33:57 Which model to pick for your use case
🔗 Resources:
Full write-up with all the KV cache numbers: kgptalkie.com/tutorials/llm-benchmarking/ornith-1-5-9b-vs-35b-a3b-benchmark
Ornith 1.5 35B-A3B GGUF: huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF
Ornith 1.5 9B GGUF: huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF
Runtime used: github.com/ggml-org/llama.cpp
📺 Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention: youtube.com/watch?v=QkzEkfIzvBk
🎓 Go deeper — my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents: kgptalkie.com/langchain
If this helped you pick a model, hit like, and tell me in the comments which card you are running and which model you settled on, so I know what to benchmark next. Subscribe and turn on the bell if you want the rest of the local LLM benchmarks.
#LocalLLM #Ornith15 #Qwen38 #Gemma4 #RTX5090
Ornith 1.5 9B and 35B benchmarked against Qwen 3.8 27B and Gemma 4 on a single RTX 5090, including the KV cache math that decides which of these models actually fits on your GPU.
I ran all four models on the same rig with llama-server, 4 bit weights and FP16 KV cache, and graded every answer with code instead of asking another model to judge. You will see generation speed, prompt processing speed, VRAM and KV cache per token, how speed holds up as the context window fills, a 128K needle test, parallel throughput at 8 concurrent requests, and tool call accuracy for agentic coding. I also opened the GGUF files and went through the tensors block by block, which is where the most interesting finding of this video came from. Everything you need is in this one video, no prior parts required.
📺 Full playlist: youtube.com/playlist?list=PLc2rvfiptPSReropGbvDFpB6dneNBwqhD
⏱ Chapters:
0:00 Intro: Ornith 1.5 9B and 35B
0:58 Dense vs MoE and the 262K context window
1:54 Why I test on RTX 5090
3:19 Setup: llama-server, 4-bit, FP16 KV cache and code grading
4:12 Active parameters per token explained
5:52 Inside the GGUF: Ornith is built on Qwen 3.5
7:16 256 experts, 8 active per token
8:07 Attention on every 4th block and gated DeltaNet
9:50 40 blocks plus the MTP head
11:23 Why Gemma 4 uses more active parameters
13:00 Generation speed: 271 vs 192 vs 77 tokens per second
15:19 Gemma 4 MoE at 234 tokens per second
17:40 Prompt processing: Ornith 9B is the fastest
20:01 VRAM and KV cache per token
21:31 Why Qwen 3.8 27B cannot hold the full 262K context
23:26 Speed as the context window fills up
25:12 128K needle test: 75 seconds vs 23 seconds
26:17 MTP explained and why it slows Ornith down
29:55 Parallel throughput: 702 vs 673 tokens per second
32:08 Same task: time and tokens used
33:29 Tool call efficiency for agentic coding
33:57 Which model to pick for your use case
🔗 Resources:
Full write-up with all the KV cache numbers: kgptalkie.com/tutorials/llm-benchmarking/ornith-1-5-9b-vs-35b-a3b-benchmark
Ornith 1.5 35B-A3B GGUF: huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF
Ornith 1.5 9B GGUF: huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF
Runtime used: github.com/ggml-org/llama.cpp
📺 Watch next:
Qwen 3.8 27B Speed Settings Explained: MTP, KV Cache and Flash Attention: youtube.com/watch?v=QkzEkfIzvBk
🎓 Go deeper — my Udemy course:
Master Langchain v1 and Ollama - Chatbot, RAG and AI Agents: kgptalkie.com/langchain
If this helped you pick a model, hit like, and tell me in the comments which card you are running and which model you settled on, so I know what to benchmark next. Subscribe and turn on the bell if you want the rest of the local LLM benchmarks.
#LocalLLM #Ornith15 #Qwen38 #Gemma4 #RTX5090

