Colibri: Run GLM 5.2 on a 32GB Laptop With Disk Streaming

TL;DR
A solo developer built a 1,300-line C inference engine that runs the 744B GLM 5.2 model on consumer hardware by streaming routed experts from disk.
Last updated: August 14, 2026
Update (August 14, 2026): GLM-5.3 is out, but its weights are not - Z.ai promises the checkpoint roughly two weeks after launch. Since 5.3 shares 5.2's base model and MoE architecture, Colibri's disk-streaming and expert-offloading approach should apply to it directly once the weights land. Everything below still describes the only GLM flagship you can run locally today.
A developer with a 12-core laptop and 32GB of RAM got GLM 5.2 running locally. Not a quantized 7B parameter model - the full 744B Mixture-of-Experts flagship. The project, Colibri, hit the Hacker News front page with 730+ points and 180 comments. The HN thread reflects equal parts admiration for the hacker spirit and practical questions about when this approach makes sense.
If disk streaming sounds too slow for daily use, the more common path to running GLM 5.2 locally is quantization with Unsloth, which trades some accuracy for speed instead of trading speed for full precision.
The Core Insight#
GLM 5.2's MoE architecture activates only ~40B parameters per token out of its 744B total. Of those, only ~11GB changes from token to token (the routed experts). The rest - attention layers, shared experts, embeddings (~17B parameters) - stays constant.
Colibri exploits this by:
- Keeping the dense part resident in RAM at int4 quantization (~9.9GB)
- Storing routed experts on disk (~370GB at int4, ~19MB per expert)
- Streaming experts on demand with per-layer LRU caching and OS page cache as a free L2
The engine is a single C file - c/glm.c at ~1,300 lines. No BLAS, no Python at runtime, no GPU required.
Performance Numbers#
The author is upfront: this is slow. Initial reports mention 0.1 tokens per second. But that was never the point. From the HN submission:
"The important thing was the journey to reach this goal. I just wanted it to work at all costs, even slowly."
Community benchmarks show better results on faster hardware. Users with NVMe drives and more RAM report usable speeds, though still far below cloud API performance.
What HN Is Saying#
The hacker spirit resonates:
The top comment simply states: "This is the hacker spirit." The author replied: "Thank you so much, it's true! It all started with this spirit!"
This energy pervades the thread. Multiple commenters compared Colibri to antirez's ds4 project (the creator of Redis working on a similar disk-streaming approach for GLM 5.2). The author confirmed inspiration: "Antirez is the number one!"
SSD wear concerns:
Several asked about disk lifespan. The README addresses this directly with an SSD wear warning. The author clarified that heavy writes are limited to the KV cache, while expert reads dominate. Architectures with unified memory (like Apple Silicon) can keep the KV cache in RAM entirely.
Practical alternatives:
A pragmatic commenter noted: "For most projects the more practical solution is to use clouds offering GLM 5.2 for free. 1 token per minute is minuscule compared to their rate limits for free usage."
This is true. For production use, cloud APIs are faster and cheaper - see our roundup of free and cheap ways to access GLM 5.2 if that's the goal. But that misses what Colibri demonstrates: that the architectural constraints of MoE models enable approaches previously thought impossible.
Agent integration:
Users asked whether Colibri can plug into coding agents like Claude Code or Pi. The author confirmed work is underway: "We're working on it right now with a pull request that will also arrive for opencode!"
This would enable fully local agentic workflows, though at reduced speed.
Hardware pricing concerns:
A subthread lamented current RAM and SSD prices. One commenter's shopping cart went "from $399 to $475" for basic DDR5. Another observed that affordable local inference is getting harder as hardware costs rise. Yet others pointed out that used or budget hardware can hit the necessary specs for under $600.
Technical Details#
The architecture breaks down like this:
| Component | Size (int4) | Location |
|---|---|---|
| Dense part (attention, shared experts, embeddings) | ~9.9GB | RAM (resident) |
| Routed experts (21,504 total) | ~370GB | Disk (streamed) |
| Per-expert size | ~19MB | - |
The 21,504 routed experts come from 75 MoE layers with 256 experts each, plus the MTP head. At runtime, only a small subset is active per token, and the LRU cache keeps hot experts in memory.
No external dependencies means the project compiles anywhere with a C compiler. The tradeoff is reimplementing functionality that libraries would provide, but for a research/hobby project, the simplicity has value.
Related Work#
Colibri isn't the only project exploring disk-based inference:
- antirez's ds4: The Redis creator has a GLM 5.2 branch using similar SSD streaming techniques. Reports suggest usable speeds on a 128GB M5 MacBook Pro.
- llama.cpp: The standard for local inference, though it typically expects models to fit in memory or uses mmap for slower streaming. Since October 2026 it can also keep recently used experts in a GPU cache while the rest stay in RAM; our llama.cpp MoE offload guide covers when that beats the older
--n-cpu-moesplit. It is the engine underneath Ollama, which is the friendlier entry point for most people. - ExLlamaV2: GPU-focused but exploring similar expert-level offloading strategies.
The common thread is exploiting MoE sparsity. When only ~5% of parameters are active per token, you don't need the entire model in fast memory. For the economics behind why sparsity matters so much for open-weight models, see GLM 5.2's cost math versus other open-weight coding models.
When This Makes Sense#
Colibri fills a specific niche:
- Learning and experimentation: Understanding how MoE inference actually works at the systems level
- Fully offline operation: No network dependency, no API costs, no data leaving your machine
- Proof of concept: Demonstrating that consumer hardware can run frontier models
It does not make sense for:
- Production workloads requiring speed
- Cost optimization (cloud APIs are cheaper per token)
- General coding assistance (too slow for interactive use)
The author is clear-eyed about this: "I don't have that hardware so I can't test it on hardware that is more powerful than my computer."
What's Next#
The project is actively developed. The author is working on:
- OpenCode integration for agentic workflows
- Performance improvements to reduce streaming overhead
- Community contributions (the README welcomes participation)
For developers interested in low-level LLM inference, Colibri offers a readable codebase. At 1,300 lines of C, you can understand the entire system in an afternoon. That's rare for ML inference code.
The project embodies a principle that resonates with HN's audience: software doesn't have to be practical to be valuable. Sometimes you build something just to prove it can be done.
FAQ#
Can I run GLM 5.2 on a laptop without a GPU?#
Yes, with Colibri. The engine keeps the ~9.9GB dense part of the model resident in RAM and streams the routed experts from disk on demand, so it runs on a CPU-only 32GB machine. The tradeoff is speed - initial reports show around 0.1 tokens per second, far below what a GPU or cloud API delivers.
How much disk space does Colibri need for GLM 5.2?#
About 370GB at int4 quantization for the routed experts, plus the ~9.9GB dense part that stays in RAM. The routed experts are split across 21,504 individual expert files (~19MB each) so only the ones needed for a given token get read.
Is disk-streamed local inference practical for daily coding work?#
Not yet. At the speeds Colibri and similar projects (like antirez's ds4) currently achieve, interactive coding assistance is impractical. It fills a different niche: offline experimentation, learning how MoE inference works, and proving that consumer hardware can technically run frontier-scale models. For actual coding work, free and cheap cloud access to GLM 5.2 or a quantized local deployment are the practical options, and our best local models hub covers the smaller models that run well on a laptop today.
Will SSD wear be a problem running Colibri long-term?#
The project's README addresses this directly - heavy disk writes are limited to the KV cache, while the bulk of I/O is expert reads, which wear SSDs far less than writes. Machines with unified memory (like Apple Silicon) can keep the KV cache in RAM entirely, avoiding the write concern altogether.
Continue Reading#
- GLM 5.2 Local Deployment with Unsloth Quantization - the more practical route to running GLM 5.2 on your own hardware
- GLM 5.2 Free and Cheap Access in 2026 - cloud alternatives when local inference is too slow
- GLM 5.2 Cost Math for Open-Weight Coding Models - why MoE sparsity matters for pricing, not just hardware
- GLM 5.2 vs DeepSeek v4 vs Qwen3: Open-Weights Coding Showdown - how GLM 5.2 stacks up against other open-weight models
- GLM 5.2 in 9 Minutes - a fast primer on the model Colibri is running
- GLM 5.2 Matches Human Bookkeeper Accuracy on UK VAT Returns - With Some Caveats
Sources#
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on local and open-weight models
GLM-5.2 Local Deployment: Running Z.ai's 744B Model on Consumer Hardware
Unsloth's dynamic quantization makes GLM-5.2 runnable on a 256GB Mac or a 24GB GPU with CPU offloading. Here is the hardware math, the quantization tradeoffs, and what the HN community learned from actually running it.
7 min readRunning Gemma 4 26B at 5 Tokens/Sec on a 13-Year-Old Xeon With No GPU
A developer got Google's Gemma 4 26B running on 2013 Xeon hardware for under $300. The fix for a silent MoE bug is now upstream - here's what it means for local inference.
6 min readGLM 5.2 Matches Human Bookkeeper Accuracy on UK VAT Returns - With Some Caveats
A new benchmark shows GLM 5.2 processing 59 transactions and producing VAT returns off by only 7 pence - at $2.73 versus typical accounting fees of $1,000+. Here is what the benchmark actually tested, where the model failed, and why the HN discussion focused on liability.
7 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








