
GLM-5.2
6 partsTL;DR
A data-rich, source-cited comparison of the open-weights coding models that matter in 2026: GLM-5.2, DeepSeek V4, Qwen3, and the new Kimi K3 frontier entrant. Benchmark table, per-token pricing, context windows, self-host footprint, and a clear pick-X-if decision matrix.
Direct answer
A data-rich, source-cited comparison of the open-weights coding models that matter in 2026: GLM-5.2, DeepSeek V4, Qwen3, and the new Kimi K3 frontier entrant. Benchmark table, per-token pricing, context windows, self-host footprint, and a clear pick-X-if decision matrix.
Best for
Developers comparing real tool tradeoffs before choosing a stack.
Covers
Verdict, tradeoffs, pricing signals, workflow fit, and related alternatives.
Last updated: July 31, 2026
The interesting fight in coding models is no longer open versus closed. It is open versus open. Three labs now ship weights you can download, self-host, and route to per-token at a fraction of frontier pricing, and all three post real software-engineering benchmark numbers: Z.ai's GLM-5.2, DeepSeek V4, and Alibaba's Qwen3 line. If your question is "which open-weights model should I point my coding agent at," this is the comparison built to answer it.
This piece is part of our model-economics beat. For the single-model deep dives, see our GLM-5.2 cost math, our DeepSeek V4 economics breakdown, and the self-hosting break-even math. For the layer that decides when to use which, read model routing recipes and why the orchestration layer is the next big play.
Every figure below is attributed to a primary or named source, with verification dates. Prices and benchmarks move fast on open weights because anyone can host them - verify against the live vendor page before you commit a production budget.
deepseek-chat / deepseek-reasoner names; the live API is now deepseek-v4-pro and deepseek-v4-flash, with V4 Flash 0731 shipping on July 31 as a re-post-trained agent workhorse at the same $0.14/$0.28 rates (DeepSeek change log).Each of these is a family, not a single model. To keep the comparison fair, we anchor on the strongest open-weights variant from each lab, because the whole point of this category is downloadable weights:
deepseek-v4-pro API name.That last point about Qwen is the single most important caveat in this post: in mid-2026, Qwen's very best coding model is closed. The open-weights Qwen you can actually self-host is a much smaller MoE - which, as the numbers show, punches far above its size. Kimi K3 inverts that story: it is the frontier model that did come out, and the trade-off is a datacenter-sized footprint (see the self-host section).
| GLM-5.2 | DeepSeek V4 Pro | Qwen3.6-35B-A3B | Kimi K3 | |
|---|---|---|---|---|
| Vendor | Z.ai (Zhipu) | DeepSeek | Alibaba | Moonshot |
| Released | Jun 16, 2026 | Apr 2026 | Apr 16, 2026 | Jul 27, 2026 |
| Total params | 753B (MoE) | 1.6T (MoE) | 35B (MoE) | 2.8T (MoE) |
| Active params | 40B | 49B | 3B | 104B |
| Context window | 1M | 1M | large (VRAM-bound when self-hosted) | 1M |
| Max output | ~131K | 384K | - | not published |
| License | MIT | MIT | Apache 2.0 | Custom (free under thresholds) |
| Self-host class | multi-GPU server | multi-GPU server | single 24GB GPU | ~1.5TB VRAM (datacenter) |
Sources for this table are in the Sources section. The standout structural facts: DeepSeek V4 Pro is a large 1.6T total-parameter MoE, GLM-5.2 sits in the middle, Kimi K3 is the largest open release ever at 2.8T total with 104B active, and Qwen3.6-35B-A3B is two orders of magnitude smaller in total parameters and activates only 3B per token - which is why it is the only one of the four that runs on a single consumer-class GPU.
Coding benchmarks for open-weights models are reported by a mix of vendors, aggregators, and third-party evaluators using different scaffolds. Treat them as a cluster, not a leaderboard, and never compare a SWE-bench Verified number against a SWE-bench Pro number - they are different, harder tests.
| Benchmark | GLM-5.2 | DeepSeek V4 Pro | Qwen3.6-35B-A3B | Kimi K3 |
|---|---|---|---|---|
| SWE-bench Verified | not separately reported | ~80.6% (V4 Pro-Max) | 73.4 | not reported |
| SWE-bench Pro | 62.1 | 55.4 (unverified scaffold) | not reported | not reported |
| Terminal-Bench 2.x | 81.0 (v2.1) | 67.9 (v2.0) | not reported | 88.3 (v2.1) |
| AA Intelligence Index | 51 | 44 | not reported | not reported |
| ProgramBench | not reported | not reported | not reported | 77.8 |
| SWE-Marathon | not reported | not reported | not reported | 42.0 |
| MCPMark-Verified | not reported | not reported | not reported | 94.5 |
A few honest reads of this table:
The practical takeaway: on raw quality, Kimi K3 is now the frontier-substitute pick of the category, DeepSeek V4 Pro and GLM-5.2 hold the middle, and Qwen3.6-35B-A3B remains the efficiency play that trades a few points of capability for a dramatically smaller footprint.
From the archive
Jun 17, 2026 • 7 min read
Jun 17, 2026 • 7 min read
Jun 17, 2026 • 11 min read
Jun 17, 2026 • 8 min read
Open-weights pricing is a moving target because the original lab is just one of many hosts. Here are the official first-party API list prices, verified July 31, 2026.
| Model | Input ($/MTok) | Cached input ($/MTok) | Output ($/MTok) |
|---|---|---|---|
| GLM-5.2 (Z.ai list) | $1.40 | - | $4.40 |
| GLM-5.2 (provider median) | ~$0.55 | - | ~$1.85 |
| DeepSeek V4 Pro | $0.435 | $0.003625 | $0.87 |
| DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 |
| Kimi K3 | $3.00 | $0.30 | $15.00 |
| Qwen3.6 Plus (API) | $0.325 (35% off) | - | $1.95 (35% off) |
GLM-5.2 list pricing is $1.40 / $4.40, with a provider median closer to $0.55 / $1.85 because the open weights let third parties compete on hosting (Artificial Analysis). The slide continues: pricing has dropped roughly 45% since the June 16 launch, and the cheapest blended hosts sit near $0.72-0.80 per million tokens on a 3:1 input-output mix, with Wafer the fastest at over 200 tokens/sec. DeepSeek's live pricing page lists V4 Pro at $0.435 / $0.87 with a near-free cache hit of $0.003625, and V4 Flash at $0.14 / $0.28 - unchanged across the July 24 rename and the July 31 re-post-training (api-docs.deepseek.com/quick_start/pricing).
Kimi K3 lists at $3.00 / $15.00 with a $0.30 cache-read rate across Moonshot, Together, Fireworks, Modal, SiliconFlow, and OpenRouter - priced like a frontier model, because it benchmarks like one. The interesting wrinkle: K3's 2.8T weights are open under a custom license, so expect the same third-party hosting price competition that drove GLM-5.2 down once dedicated providers spin up.
One important correction worth flagging: a lot of April launch coverage - and our own earlier DeepSeek V4 economics post - quoted V4 Pro at $1.74 / $3.48. The live DeepSeek pricing page lists $0.435 / $0.87, roughly a quarter of that. The live page is the source of truth; if you are modeling spend, use it, not the launch articles.
Qwen3.6-35B-A3B itself is open weights, so its "price" depends entirely on who hosts it or your own hardware. The Qwen3.6 Plus API line above ($0.325 / $1.95 at a 35% promo, per OpenRouter) is included as a managed-API reference point for the Qwen family, not as the price of the 35B open-weights model specifically.
Per-token rates do not tell you what work costs. Model one agentic coding task: roughly 40,000 input tokens (repo context, files, tool results) and 8,000 output tokens (plan plus diffs), no caching.
| Model (first-party list) | Cost per task |
|---|---|
| DeepSeek V4 Pro ($0.435 / $0.87) | ~$0.024 |
| GLM-5.2 (provider median) | ~$0.037 |
| GLM-5.2 ($1.40 / $4.40) | ~$0.091 |
| Kimi K3 ($3.00 / $15.00) | ~$0.24 |
At list prices, DeepSeek V4 Pro is the cheapest of the high-quality options on this input-heavy profile, and its near-zero cache-hit rate makes repeated-context agent loops cheaper still. GLM-5.2 closes most of the gap if you shop the open-weights hosting market for the lower provider-median rate. Kimi K3 costs about 10x V4 Pro per task - the honest premium for frontier-level agent scores from open weights. The caveat that survives every pricing table: cheap per token only becomes cheap per task if the model lands the work in as few attempts as a pricier model, which is exactly why the benchmark cluster above matters.
All three flagships are built for repo-scale work.
If your workload is "drop a whole repo in and ask for a coordinated change," any of the four handles the input side. DeepSeek's 384K output ceiling is the one to reach for when the model has to write a lot, not just read a lot.
This is the category that separates "open weights in principle" from "open weights you will actually run."
If "no per-token bill, runs on hardware I already own" is your hard requirement, the decision is effectively made for you: Qwen3.6-35B-A3B is the open-weights coder that fits on a single GPU, and it is the only one of the four in that class.
Pick Kimi K3 if you want the top of the open-weights category, period - Terminal-Bench 2.1 at 88.3 and agent benchmarks that match closed frontier models. You are consuming it via a managed API (Moonshot, Together, Fireworks, Modal, SiliconFlow, or OpenRouter at $3/$15) and your business is under the license's revenue thresholds. Budget for a frontier-model token bill: K3 costs about 10x DeepSeek V4 Pro per task.
Pick GLM-5.2 if you want the strongest reported SWE-bench Pro and Terminal-Bench numbers among the established open-weights models, you are consuming it via API or a managed host, and you value an MIT license with a wide ecosystem of coding-tool integrations. The price slide since launch (roughly 45%, with blended third-party rates near $0.72-0.80 per million tokens) makes it the value pick at the frontier-substitute tier.
Pick DeepSeek V4 Pro if unit cost is your primary axis and you want frontier-substitute quality. At $0.435 / $0.87 with a near-free cache hit and a 384K output ceiling, it is the cheapest high-quality option for input-heavy, cache-friendly agent loops. Route to V4 Flash ($0.14 / $0.28) for the bounded, high-volume inner-loop steps where you do not need Pro-level reasoning - the 0731 re-post-training made it a credible whole-loop agent.
Pick Qwen3.6-35B-A3B if you must self-host on modest hardware, want zero per-token cost, or care about latency and data residency. Scoring 73.4 on SWE-bench Verified from a 3B-active model that fits on a 24GB GPU is the best capability-per-footprint deal in open weights right now. Choose the managed Qwen3.7-Max API only if you need Qwen's absolute top coding quality and can accept that it is closed-weights.
Route across all four if you are running real volume. The honest answer for most production setups is not "pick one" but "tier them": cheap open-weights for the easy majority, a frontier-substitute for the hard reasoning, with a failover chain. That is precisely the pattern in our model routing recipes, and the reason the orchestration layer is where the margin is moving. K3 as the frontier rung, GLM-5.2 or V4 Pro as the middle, V4 Flash or Qwen3.6-35B-A3B as the cheap floor is a coherent July 2026 stack.
At first-party API list prices verified July 31, 2026, DeepSeek V4 Pro is the cheapest high-quality option at $0.435 input / $0.87 output per million tokens, with a near-free cache-hit input price of $0.003625. DeepSeek V4 Flash is cheaper still at $0.14 / $0.28 for lighter work. GLM-5.2 lists at $1.40 / $4.40 on Z.ai but has a provider median closer to $0.55 / $1.85 because third parties host the open weights. Kimi K3 is the premium option of the category at $3.00 / $15.00.
It depends on the benchmark, and you should read them as a cluster. Kimi K3 is the new category leader: Terminal-Bench 2.1 at 88.3, the top open-model score, plus ProgramBench 77.8 and MCPMark-Verified 94.5. Among the established three, GLM-5.2 leads the SWE-bench Pro (62.1) and Terminal-Bench 2.1 (81.0) figures it reports, DeepSeek V4 Pro-Max posts the strongest SWE-bench Verified number (around 80.6%), and Qwen3.6-35B-A3B scores 73.4 on SWE-bench Verified with only 3B active parameters.
Qwen3.6-35B-A3B is the only one of the four that runs on a single consumer-class GPU - about 21GB VRAM at Q4_K_M, fitting a 24GB card. GLM-5.2 (753B), DeepSeek V4 Pro (1.6T), and Kimi K3 (2.8T, ~1.5TB VRAM at native MXFP4) require server-class or datacenter deployments, so most teams consume them via API rather than self-host.
Yes, with a license. Moonshot released the 2.8T K3 weights on Hugging Face on July 27, 2026 under a custom Kimi K3 License: free for most use, but a separate commercial agreement is required for model-as-a-service businesses above $20M aggregate revenue over any 12 consecutive months, and prominent "Kimi K3" branding applies at very large scale. The full inference stack (MoonEP expert parallelism, AgentEnv eval environment) is also open sourced.
No. As of mid-2026, Alibaba's strongest coding model, Qwen3.7-Max, is API-only on DashScope ($2.50 / $7.50 per million tokens) with no published weights. The best Qwen you can self-host is the smaller Qwen3.6-35B-A3B under Apache 2.0.
GLM-5.2 and DeepSeek V4 are released under the MIT license; Qwen3.6-35B-A3B is under Apache 2.0; Kimi K3 uses a custom license that is free under revenue thresholds and requires a commercial agreement for large model-as-a-service businesses. All four permit commercial use and self-hosting (K3 with the license caveat).
| Source | Link | Last verified |
|---|---|---|
| Artificial Analysis: GLM-5.2 | artificialanalysis.ai/models/glm-5-2 | July 31, 2026 |
| llm-stats: GLM-5.2 | llm-stats.com/models/glm-5.2 | July 31, 2026 |
| DeepSeek API pricing | api-docs.deepseek.com/quick_start/pricing | July 31, 2026 |
| DeepSeek change log | api-docs.deepseek.com/updates | July 31, 2026 |
| Hugging Face: Kimi K3 | huggingface.co/moonshotai/Kimi-K3 | July 31, 2026 |
| Kimi K3 access and pricing guide | kimi.com/blog/kimi-k3 | July 31, 2026 |
| OpenRouter: K3 | openrouter.ai/moonshotai/kimi-k3-20260715 | July 31, 2026 |
| Qwen3.6-35B-A3B announcement | qwen.ai/blog?id=qwen3.6-35b-a3b | July 31, 2026 |
| Will It Run AI: Qwen3.6 VRAM | willitrunai.com/blog/qwen-3-6-vram-requirements | July 31, 2026 |
Figures verified July 31, 2026. Benchmark scores are reported by different evaluators on different scaffolds; treat them as a cluster, not an exact ranking, and re-verify against the live vendor pages before making a production decision.
Read next
GLM-5.2 ships under an MIT license, so it is hosted everywhere - and a few places run it for free or nearly free right now. Here is every way to access Z.ai's open-weights coding model, from OpenCode Go referral credits and Devin to the cheapest per-token routes on OpenRouter, Fireworks, and DeepInfra, plus local Ollama.
10 min readZ.ai's GLM-5.2 lands as a 753B open-weights coding model that beats GPT-5.5 on SWE-bench Pro for roughly one-sixth the per-token cost. Here is the real cost math, a worked cost-per-task example, and a when-to-use-which decision guide.
9 min readChoosing a local coding LLM in 2026 means balancing benchmark performance, hardware cost, and the compliance pressure to keep code off third-party servers. Here is what to run and on what hardware.
8 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.
Fastest inference for open-source models. 200+ models via unified API. Ranks #1 on speed benchmarks for DeepSeek, Qwen,...
View ToolDeepSeek's open-weights frontier family, previewed April 24, 2026. V4-Pro is 1.6T total / 49B active params; V4-Flash is...
View ToolOpen-source terminal agent runtime with approval modes, rollback snapshots, MCP servers, LSP diagnostics, and a headless...
View ToolOpen-source reasoning models from China. DeepSeek-R1 rivals o1 on math and code benchmarks. V3 for general use. Fully op...
View ToolInstall Ollama and LM Studio, pull your first model, and run AI locally for coding, chat, and automation - with zero cloud dependency.
Getting StartedDeep comparison of the top AI agent frameworks - LangGraph, CrewAI, Mastra, CopilotKit, AutoGen, and Claude Code.
AI AgentsClickable PR link in the footer with review state color coding.
Claude Code
Exploring Google's Advanced Gemma 2 AI Models and Exciting Updates In this video, I delve into Google's newly released Gemma 2 AI models, including the 9 billion and 27 billion parameter versions....

DeepSeek V4: 1M Context, 10x KV Cache Savings, and Ultra-Low Pricing DeepSeek released V4, highlighting major long-context efficiency gains: at a 1M-token context, V4 Pro uses 27% of FLOPs and 10% of...

Learn more about Kimi K2 here; https://www.kimi.com/?utm_campaign=TR_LG3aFs5j&utm_content=&utm_medium=Youtube&utm_source=CH_kEMBez3l&utm_term= In this video, I dive deep into why Kimi K2,...

GLM-5.2 ships under an MIT license, so it is hosted everywhere - and a few places run it for free or nearly free right n...

Z.ai's GLM-5.2 lands as a 753B open-weights coding model that beats GPT-5.5 on SWE-bench Pro for roughly one-sixth the p...

Choosing a local coding LLM in 2026 means balancing benchmark performance, hardware cost, and the compliance pressure to...

Moonshot AI released the full Kimi K3 weights on HuggingFace today - 2.8T parameters, 1M context, native MXFP4 quantizat...

Compare every verified Kimi K3 access route, including Moonshot, Together, Fireworks, Baseten, Modal, Vercel AI Gateway,...

DeepSeek re-post-trained V4 Flash into an agent workhorse: Terminal Bench 82.7, DeepSWE 54.4, native Responses API, and...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.