Tencent Hy3: A 295B Open MoE That Punches Above Its Weight

TL;DR
Tencent's Hy3 ships 295B parameters but activates only 21B per token, matching flagship performance at flash-tier pricing under Apache 2.0.
Tencent released Hy3, the full production version of their Hunyuan 3 model series, on July 6, 2026. The model ships under Apache 2.0 with weights on Hugging Face and ModelScope, aiming squarely at developers who want frontier-adjacent capability without frontier pricing. The successor arrived August 28, 2026: Hy4 preview is 770B total with 49B active, a 1M context window, and an OpenRouter price of $0.834/$2.501 - the step-change generation this post's comparisons now feed into.
The Architecture#
Hy3 is a 295B-parameter Mixture-of-Experts model with 192 experts using top-8 routing. Only 21B parameters activate per token (plus 3.8B for the MTP layer), so inference compute stays low despite the headline parameter count. Context length is 256K tokens.
For comparison, DeepSeek V4 Flash sits at 284B total parameters with about 13B active. The two models occupy similar hardware requirements, which makes their performance delta meaningful.
What sets Hy3 apart from the April preview:
- Hallucination rate dropped from 12.5% to 5.4%
- Commonsense errors fell from 25.4% to 12.7%
- The model integrated feedback from over 50 internal Tencent product teams
What HN Is Saying#
The Hacker News thread focused heavily on practical comparisons rather than benchmark tables.
DeepSeek V4 vs Hy3: Several commenters tested both models head-to-head. One noted that "GLM 5.2 is pretty close to gpt-5.4 base, and much better than it when it comes to design stuff" while Hy3 slots in below GLM 5.2 but trades favorably against DeepSeek V4 Flash on many tasks.
Another commenter running both locally wrote: "DS4 Flash can currently run reasonably well on systems with 96GB+ RAM, I wonder if Hy3 can compete there." The answer depends heavily on quantization tolerance - DeepSeek V4's architecture handles aggressive quantization (down to 2-bit) better than most models due to its FP4 native MoE parameters.
Local inference reality: A practical assessment from the thread: "I've found DS4 Flash to be very temperamental via Claude Code. The speed is great, but it often builds a completely wrong mental model and charges off down the wrong path... Hy3 isn't as fast, but so far it seems to stay on track much more reliably."
KV cache differences: Hy3 lacks DeepSeek V4's aggressive KV cache optimizations. One commenter running both on DGX Sparks reported: "Whereas I can run DS4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 quantized to FP4, there is only room for 130K tokens of KV cache."
Coding benchmarks: The skeptics pointed to DeepSWE scores - Hy3 at 28% vs GPT-5.4 xhigh at 52%. One commenter suspected "a lot of contaminated benchmarks in the blog post about Hy3, needs real testing though I have a distinct feeling it's benchmaxxed like a lot of Chinese models."
Pricing and Availability#
Hy3 is free on OpenRouter until July 21, 2026. After that, expect pricing similar to DeepSeek V4 Flash tier - roughly $0.10-0.30 per million input tokens.
The model is also available on:
- Hugging Face and ModelScope (weights under Apache 2.0)
- Hermes, Kilo, Cline, OpenClaw, OpenCode, and Cherry Studio
- Tencent's own Hunyuan API
For local deployment, Tencent recommends H20-3e or equivalent GPUs with large memory capacity to serve the full 295B parameters across 8 GPUs.
When to Use Hy3#
Based on the HN discussion and Tencent's benchmarks, Hy3 fits specific workflows:
Good fit:
- Agentic tasks where reliability matters more than raw speed
- Long-context reasoning (256K window)
- Workflows where Apache 2.0 licensing is required
- Cost-sensitive production with OpenRouter's promotional pricing
Less ideal:
- Deep coding tasks (coding benchmarks lag behind GLM 5.2 and frontier models)
- Extremely long sessions requiring large KV caches
- Cases where you need aggressive quantization to fit in memory
The Bigger Picture#
Hy3 represents the continued compression of "frontier-tier" capability into open-weight models. A year ago, you needed API access to GPT-4 or Claude to get this level of performance. Now a 295B MoE with 21B active parameters - runnable on high-end consumer hardware - delivers comparable results on many tasks.
The practical question for developers is whether to build on these open models or stick with the API providers. Open models give you full control over inference, no rate limits, and no surprise deprecations. The tradeoff is operational complexity and the need to track new releases manually.
For now, the free tier on OpenRouter makes Hy3 worth testing. If your agentic workflows need a model that stays on track better than DeepSeek V4 Flash, this is a legitimate option.
Continue Reading#
- Gleam Moves to Tangled: What the ATProto Code Forge Means for Developers
- GLM 5.2 Outperforms Claude Code on Semgrep's IDOR Vulnerability Benchmarks
- Godot Bans AI-Authored Code Contributions - What It Means for Open Source
Sources#
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next
Tencent Hy4 Preview: 770B Open MoE, Agentic Benchmarks, and Where It Fits in 2026
Tencent's Hy4 preview ships 770B total parameters with 49B active under Apache 2.0 - a 1M-context text MoE with DeepSeek-style sparse attention, posted Terminal-Bench 85.4 and DeepSWE 64.3, and an OpenRouter price of $0.834/$2.501. Verified against the model card and the live OpenRouter page on August 31, 2026.
9 min readInkling-Small: Thinking Machines Ships a 12B-Active Open Model That Beats Its Big Sibling on Agent Work
Inkling-Small is a 276B-parameter MoE with 12B active per token, Apache 2.0, and open weights. It beats the 975B Inkling on SWEBench Verified (80.2), HLE (31.6), and tool use at a quarter of the size and a third of the output price.
8 min readEcho Claims Fable-Level Results at One-Third the Cost Using Open-Weight Models
A new multi-model orchestration system routes requests across open-weight models to match frontier performance at reduced inference cost. Here is what we know.
7 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







