
Colibri: Running GLM 5.2 on a 32GB Laptop with Disk Streaming and Expert Offloading
A solo developer built a 1,300-line C inference engine that runs the 744B GLM 5.2 model on consumer hardware by streaming routed experts from disk. Here's how it works.











