ConfirmedFramework or tooling · Tier 2
Colibrì runs the 744B GLM-5.2 on 16–25 GB of RAM with no GPU, slowly
So what? A tiny open-source C engine streams a frontier-size MoE model's experts from disk, so it runs offline with no GPU and no per-token bill. Needs ~372 GB of NVMe and runs at about 1 token every 10–20 s on a 25 GB machine: great for private batch jobs, not for chat.