Topic
All blog posts, tools, and guides about MoE from Developers Digest.
6 resources - 4 posts, 2 tools

Tencent's Hy4 preview ships 770B total parameters with 49B active under Apache 2.0 - a 1M-context text MoE with DeepSeek-style sparse attention, posted Terminal-Bench 85.4 and DeepSWE 64.3, and an OpenRouter price of $0.834/$2.501. Verified against the model card and the live OpenRouter page on August 31, 2026.

Tencent's Hy3 ships 295B parameters but activates only 21B per token, matching flagship performance at flash-tier pricing under Apache 2.0.

A practical walkthrough of Nemotron 3 Super: latent mixture of experts, hybrid Mamba transformer architecture, 1M context, reasoning modes, and the code you actually need to run it on NVIDIA hardware.

NVIDIA's Nemotron 3 Super combines latent mixture of experts with hybrid Mamba architecture - 120B total parameters, 12B active per token, 1M context, and up to 4x more experts at the same cost.
Alibaba's open-weight coding model, released in 2025. 480B total parameters, 35B active (MoE). Native 256K context, extends to 1M. Apache 2.0 license. Built for agentic coding.
AI ModelsDeepSeek's open-weights frontier family, previewed April 24, 2026. V4-Pro is 1.6T total / 49B active params; V4-Flash is 284B / 13B. 1M context standard. Weights on Hugging Face.
AI ModelsKeep exploring

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.
Explore 819 topics
Browse All Topics