Where to Run GLM-5.2 Free and Cheap: Providers Compared

GLM-5.2
6 parts- 1Where to Run GLM-5.2 Free and Cheap: Providers ComparedCurrent
- 2GLM-5.2 Cost Math: When Open-Weights Coding Models Actually Save You Money
- 3GLM-5.2 vs DeepSeek V4 vs Qwen3: The Open-Weights Coding Model Showdown (2026)
- 4GPT-5.5 Has a 3x Higher Hallucination Rate Than MIT-Licensed GLM-5.2
- 5GLM-5.2 Local Deployment: Running Z.ai's 744B Model on Consumer Hardware
- 6GLM 5.2 and the AI Margin Collapse Thesis
TL;DR
GLM-5.2 is an MIT-licensed open-weights coding model. Cheapest routes: OpenRouter hosts from about $0.36 per million input tokens, Z.ai's API, or self-host.
Last updated: October 7, 2026 (meta description revised; prices as checked September 28, 2026)
The cheapest hosted way to run GLM-5.2 is a third-party host on OpenRouter, where several providers list it well below Z.ai's own $1.40 input and $4.40 output per million tokens. There is no permanent free hosted tier. The weights are MIT-licensed on Hugging Face, so self-hosting has no per-token cost, but it needs datacenter-class hardware. GLM-5.3 is the newer model in the family.
Official sources#
| Source | What it covers |
|---|---|
| Hugging Face: zai-org/GLM-5.2 | Open weights, MIT license, supported serving frameworks |
| Z.ai pricing docs | First-party per-token API prices |
| Z.ai devpack overview | GLM Coding Plan tiers, credits, supported models |
| OpenRouter: z-ai/glm-5.2 | Live multi-provider prices and quantization |
| Ollama: glm-5.2 | Ollama listing for the model |
Prices are per million tokens and were checked on September 28, 2026. Pricing pages are the source of truth and they move, so treat the numbers as a snapshot.
What GLM-5.2 is#
GLM-5.2 is Z.ai's (formerly Zhipu AI) open-weights coding model, released in June 2026. It is a mixture-of-experts model of roughly 753B total parameters with a 1M-token context window, released under an MIT license. Its Hugging Face card now points to GLM-5.3 as the newer version, which shares the same base model. GLM-5.3 changed the license to a custom one, so GLM-5.2 is still the pick if you need MIT terms. For benchmark and cost breakdowns, see the GLM-5.2 cost math.
Cheapest paid routes#
Because the weights are open, any inference shop can serve GLM-5.2 and compete on price. OpenRouter's endpoint list shows 31 endpoints. A sample with stated quantization:
| Provider (via OpenRouter) | Input ($/1M) | Output ($/1M) | Quantization |
|---|---|---|---|
| Inceptron | 0.36 | 2.40 | fp4 |
| Baidu | 0.36 | 1.13 | fp8 |
| DeepInfra | 0.56 | 1.80 | fp4 |
| Novita | 0.65 | 2.04 | fp8 |
| CoreWeave | 0.76 | 2.42 | fp4 |
| Z.ai (first party) | 1.40 | 4.40 | fp8 |
- OpenRouter is a router, not a host. It sends your request to a provider that meets your price and speed constraints, so you get failover and price competition without managing keys for each host. See the OpenRouter profile and Model Routers and the Optionality Advantage.
- Quantization matters. The cheapest routes often serve fp4 or unstated quantization. For coding agents the quality gap can be small but real, so test your own task before optimizing on price.
- Context limits differ by host. Some providers cap context well below 1M, so check the endpoint before sending long agent sessions.
Direct from Z.ai: API vs the Coding Plan#
The per-token API charges $1.40 input, $0.26 cached input, and $4.40 output per million tokens. That suits variable or bursty usage.
The GLM Coding Plan has Lite, Pro, and Max tiers, starting at $18 per month for Lite. Z.ai's docs list GLM-5.3 and GLM-5.3-Flash as the supported models and say requests for older models are routed to current versions, so the Coding Plan is no longer a GLM-5.2 route. Use the API or a third-party host if you want to pin GLM-5.2. The plan wins for daily coding with predictable cost.
Local and self-host#
The weights are on Hugging Face under MIT, so you can run GLM-5.2 with no per-token cost if you have the hardware.
- Ollama. The model is listed as
glm-5.2on Ollama. At roughly 753B total parameters this is a datacenter-class model, not a laptop one. For genuinely local coding on modest hardware, see the best local coding LLMs and the best local models hub. The disk-streaming approach in Colibri is one option if you want the full model on small hardware. - Self-host at scale. The Hugging Face card lists serving paths for vLLM, SGLang, Transformers, and KTransformers. This only pays off at high, steady token volume. Below that, a hosted route is cheaper and far less operational work.
Which route should you pick?#
- Cheapest production tokens: an OpenRouter host such as DeepInfra or Novita, after testing quality at its quantization.
- Daily agentic coding in one tool: the GLM Coding Plan, which now serves GLM-5.3.
- MIT license and full control: self-host the weights.
- Truly offline: local weights, with serious GPU memory, or a smaller model.
FAQ#
Is GLM-5.2 free?#
The weights are free under an MIT license, so self-hosting has no per-token cost. Hosted access is paid: OpenRouter has no free GLM-5.2 endpoint in its list, and prices start around $0.36 per million input tokens on the cheapest hosts.
What is the cheapest way to use GLM-5.2?#
A third-party host on OpenRouter. Baidu and Inceptron list input at $0.36 per million, and DeepInfra at $0.56, against $1.40 at Z.ai. Self-hosting is cheaper only at high, steady volume.
Can I run GLM-5.2 with Claude Code, Cursor, or OpenCode?#
Yes, through any provider that exposes an OpenAI-compatible or Anthropic-compatible endpoint. Model-agnostic agents like OpenCode let you point at OpenRouter or Z.ai and swap models freely. Z.ai's Coding Plan now routes to GLM-5.3, so use the API or a third-party host to stay on 5.2.
Can I run GLM-5.2 locally on my laptop?#
Not practically. At roughly 753B total parameters, even a 4-bit quant needs high-RAM multi-GPU hardware. For consumer hardware, use a smaller dense model.
What license is GLM-5.2 under?#
MIT, per the Hugging Face model card. GLM-5.3 uses a custom glm-5.3 license instead.
Continue Reading#
- Where to Run GLM-5.3 Free and Cheap - the newer edition of this guide
- Where to Access AI Models in 2026 - access routes, free tiers, and prices for every major model
- GLM-5.2 Cost Math for Open-Weight Coding Models - the worked cost-per-task numbers
- GLM-5.2 in 9 Minutes - a fast primer on the model itself
- The Best Local Coding LLMs of 2026 - smaller models for laptop-class inference
- Colibri: Running GLM 5.2 on a 32GB Laptop - the disk-streaming approach to self-hosting
Get the next comparison like this in your inbox
One email a week on glm and the rest of the AI dev stack. Free.
Read next on local and open-weight models
Where to Run GLM-5.3 Free and Cheap: Flash, Batch and Provider Prices (2026)
GLM-5.3 routes compared, verified October 1, 2026: OpenRouter hosts list it from $0.12 per million input tokens, GLM-5.3-Flash costs $0.15 input and $0.50 output from Z.ai, batch drops to $0.45/$2.00, and the Coding Plan starts at $18 per month.
10 min readThe Router Era: Why Not Owning a Frontier Model Became an Advantage
No single model wins every task anymore, and the companies that never trained one - Factory, Devin, Perplexity, Cursor, OpenCode - are turning that into a moat. This is how model routing works, why open weights and neoclouds make it cheap, and the honest counter-argument.
11 min readGLM 5.2 in 9 Minutes: The Open-Weight Rival to GPT-5.5
A companion guide to the GLM 5.2 video: an open-weight model positioned against GPT-5.5, walked through with benchmarks, pricing, and a live OpenCode demo. Here is what the video covers and where to go deeper.
6 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.





