DeepSeek V4 Flash Fully Local — 32 tok/s on a Single Chip @fahdmirza
DeepSeek V4 Flash Fully Local — 32 tok/s on a Single Chip  @fahdmirza
Uploaded July 2026 | Updated September 2026, 2 weeks ago
Running a 284 billion parameter DeepSeek V4 Flash model completely locally on a single AMD Ryzen AI MAX+ 395 with 128GB unified memory, hitting 32 tokens per second using the Lucebox inference engine with DSpark speculative decoding.

🔥 Get 50% Discount on any A6000 or A5000 GPU rental, use following link and coupon:

bit.ly/fahd-mirza
Coupon code: FahdMirza

🔥 Buy Me a Coffee to support the channel: ko-fi.com/fahdmirza

#lucebox #dflash #halostrix #deepseekv4flash

PLEASE FOLLOW ME:
â–¶ LinkedIn: / fahdmirza
â–¶ YouTube: / @fahdmirza
â–¶ Blog: fahdmirza.com

RESOURCES:

â–¶ fahdmirza.com

All rights reserved © Fahd Mirza
DeepSeek V4 Flash Fully Local — 32 tok/s on a Single ChipNanowhale-100m: Fascinating Implemention of DeepSeek-V4 ArchitectureOrnith-1.5-9B: Great on Paper, Struggled in My Tests LocallyGemma4 12B in Quantization-Aware Training (QAT) with Ollama - Full TestingNVIDIA Nemotron 3 Embed 1B Cuts Agent Token Costs by 31%Thomson Reuters Built Their Own AI Lawyer: Run Thomson-1 LocallyvLLM + PegaFlow: KV Cache That Survives Restarts (Hands-On)Qwen3.8-27B Obliterated: Dont Use this Model in ProductionLocateAnything: NVIDIA’s New AI Sees EVERYTHING: Run LocallyOh Baby! Qwen3.8-27B Coming - Lets Test Qwen3.8-Max NowLuce Spark: Run a 35B Model Under 16GB VRAM LocallyQwen3.8-27B Ridge: Smarter Quantization, Full Power on 12GB
Fahd Mirza |

DeepSeek V4 Flash Fully Local — 32 tok/s on a Single Chip

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER