Uploaded August 2026 | Updated September 2026, 3 weeks ago
Run 360GB models on 8GB RAM using Colibri technology. This approach challenges NVIDIA dominance in local LLM inference.
The demand for massive hardware to support large language models is changing. While tools like Lama.cpp started the movement, Colibri technology introduces a new method for handling model components. We look at how data processing is optimized to bypass the typical reliance on high-end GPUs, making 360GB models accessible on standard consumer CPUs.
By visualizing the internal model processing, you can see exactly where the efficiency gains originate. This breakdown is intended for anyone building local AI infrastructure on a budget who needs to run large language models without shelling out for enterprise-grade hardware. It is a clear look at a potential path forward for local LLM inference that prioritizes accessibility over expensive component requirements.
Subscribe for weekly local AI infrastructure breakdowns, and tell us in the comments if you want a direct performance test against Lama.cpp.
CHAPTERS
00:00 The claim
00:46 What it claims
04:04 Why I was skeptical
06:42 The setup
08:32 Run one — colibrì
11:11 Run two — llama.cpp
13:26 The dig
15:24 Run two, properly
19:05 Pushing harder
22:12 The verdict
THE SETUP
Model:
DeepSeek-V4-Flash-0731, 284B total / 13B active, 43 layers, 256 routed experts , plus 1 shared, top-6 routing
Box:
Ryzen 5 5600X (12 threads, no AVX-512),
61 GB RAM, RTX 3060 12 GB, NVMe
Weights
colibrì: official HF checkpoint, 166.9 GB, no conversion
llama.cpp: unsloth UD-Q8_K_XL, 161.9 GB
Size-matched within 3.1%, both verified byte-exact against the HF API.
THE FLAG
llama.cpp will not load this model at all without -nr (--no-repack). Without it, it tries to build a 147 GB repacked buffer in RAM and dies.
That flag is not in the docs folder, it's in the argument parser source and a couple of tool READMEs. llama-bench doesn't accept it at all, which is why every measurement here was done by hand.
LINKS
colibrì — github.com/JustVugg/colibri
llama.cpp — github.com/ggml-org/llama.cpp
The open work on this, in llama.cpp:
github.com/ggml-org/llama.cpp/pull/24524
github.com/ggml-org/llama.cpp/pull/25294
github.com/ggml-org/llama.cpp/pull/25932
github.com/ggml-org/llama.cpp/pull/26003
github.com/ikawrakow/ik_llama.cpp/pull/2101
Other people building in this space:
github.com/lyogavin/airllm
github.com/Helldez/BigMoeOnEdge
github.com/FareedKhan-dev/kimi-k3-in-c
Every number in this video came off my own box. If you get different results on yours I genuinely want to hear about it — the whole point is that this stuff is testable.
What's the biggest model you've got running at home, and what did it take to get there?
Comment below.
#colibri #llamacpp #deepseek #homelab #localai
Run 360GB models on 8GB RAM using Colibri technology. This approach challenges NVIDIA dominance in local LLM inference.
The demand for massive hardware to support large language models is changing. While tools like Lama.cpp started the movement, Colibri technology introduces a new method for handling model components. We look at how data processing is optimized to bypass the typical reliance on high-end GPUs, making 360GB models accessible on standard consumer CPUs.
By visualizing the internal model processing, you can see exactly where the efficiency gains originate. This breakdown is intended for anyone building local AI infrastructure on a budget who needs to run large language models without shelling out for enterprise-grade hardware. It is a clear look at a potential path forward for local LLM inference that prioritizes accessibility over expensive component requirements.
Subscribe for weekly local AI infrastructure breakdowns, and tell us in the comments if you want a direct performance test against Lama.cpp.
CHAPTERS
00:00 The claim
00:46 What it claims
04:04 Why I was skeptical
06:42 The setup
08:32 Run one — colibrì
11:11 Run two — llama.cpp
13:26 The dig
15:24 Run two, properly
19:05 Pushing harder
22:12 The verdict
THE SETUP
Model:
DeepSeek-V4-Flash-0731, 284B total / 13B active, 43 layers, 256 routed experts , plus 1 shared, top-6 routing
Box:
Ryzen 5 5600X (12 threads, no AVX-512),
61 GB RAM, RTX 3060 12 GB, NVMe
Weights
colibrì: official HF checkpoint, 166.9 GB, no conversion
llama.cpp: unsloth UD-Q8_K_XL, 161.9 GB
Size-matched within 3.1%, both verified byte-exact against the HF API.
THE FLAG
llama.cpp will not load this model at all without -nr (--no-repack). Without it, it tries to build a 147 GB repacked buffer in RAM and dies.
That flag is not in the docs folder, it's in the argument parser source and a couple of tool READMEs. llama-bench doesn't accept it at all, which is why every measurement here was done by hand.
LINKS
colibrì — github.com/JustVugg/colibri
llama.cpp — github.com/ggml-org/llama.cpp
The open work on this, in llama.cpp:
github.com/ggml-org/llama.cpp/pull/24524
github.com/ggml-org/llama.cpp/pull/25294
github.com/ggml-org/llama.cpp/pull/25932
github.com/ggml-org/llama.cpp/pull/26003
github.com/ikawrakow/ik_llama.cpp/pull/2101
Other people building in this space:
github.com/lyogavin/airllm
github.com/Helldez/BigMoeOnEdge
github.com/FareedKhan-dev/kimi-k3-in-c
Every number in this video came off my own box. If you get different results on yours I genuinely want to hear about it — the whole point is that this stuff is testable.
What's the biggest model you've got running at home, and what did it take to get there?
Comment below.
#colibri #llamacpp #deepseek #homelab #localai

![Can a 3.5GB model replace my 35B daily driver? (Bonsai 27B)
PrismML compressed Qwens 27B dense model down to 3.5 GB with true 1-bit weights. Not quantization: the weights were trained as −1/+1 from day one. So I put both Bonsai builds up against my actual daily driver, a 21 GB Qwen 3.6 35B MoE, plus Gemma 4 12B as the fair-size control, and ran about thirty tests on one RTX 3060.
Same landing-page brief. Same broken server over SSH. Same sampling settings, same KV cache quant, same effort level, every model.
What came out of it surprised me more than once. The 21 GB model is more than twice as fast as a 7 GB one, and the reason has nothing to do with file size. The only model that could not reliably finish a long task was not the smallest one. And whether Bonsai makes sense on your machine comes down to a number almost nobody quotes on a spec sheet.
If you run this on a Mac or a small card, tell me what you get in the comments.
CHAPTERS
0:00 Two models, one broken server
0:39 What Bonsai actually is (1-bit vs quantization)
3:24 Test 1 — can it design a webpage?
6:34 Test 2 — SSH debugging, everyone passed
9:26 Test 3 — the trap that survives a restart
13:29 Where it breaks (Gemmas loop)
15:35 Why the bigger model is faster (A3B vs dense)
17:54 The number nobody quotes (total footprint)
20:06 The verdict — who should run what
THE MODELS
Bonsai 27B (binary Q1_0, 3.5 GB / ternary Q2_0, 7 GB): [https://huggingface.co/prism-ml]
Qwen 3.6 35B-A3B: [https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF]
llama.cpp Metal Q2_0 support (merged): github.com/ggml-org/llama.cpp/pull/25419
llama.cpp CUDA Q2_0 support (open PR I built from): github.com/ggml-org/llama.cpp/pull/25707
THE RIG:
RTX 3060 12 GB
llama.cpp built from PR #25707
KV cache q8_0 (all models)
reasoning effort medium (all models)
MoE runs with 26 expert layers offloaded to system RAM
MY RESULTS AT A GLANCE
Generation speed at depth (tg, tokens/sec):
MoE 47.9 · Bonsai Q1 34.7 · Bonsai Q2 21.6
Total memory footprint:
MoE ~21 GB (10.5 VRAM + 10.6 RAM)
ternary 8.8 GB
binary 5.3 GB
Server-repair task, wall clock:
MoE ~56s
ternary ~114s
binary ~134s
Design test:
ternary and the MoE land in the same class.
Binary and Gemma sit a tier below.
#localai #llm #bonsai #quantization #1bit #qwen #llamacpp #selfhosted #ai #rtx3060 Can a 3.5GB model replace my 35B daily driver? (Bonsai 27B)](https://i.ytimg.com/vi/rBLWDJrXCp0/mqdefault.jpg)






