
Kimi K3
6 partsTL;DR
Compare every verified Kimi K3 access route, including Moonshot, Together, Fireworks, Baseten, Modal, Vercel AI Gateway, Cloudflare, RunPod, SiliconFlow, OpenRouter, and OpenCode Go.
Direct answer
Compare every verified Kimi K3 access route, including Moonshot, Together, Fireworks, Baseten, Modal, Vercel AI Gateway, Cloudflare, RunPod, SiliconFlow, OpenRouter, and OpenCode Go.
Best for
Developers comparing real tool tradeoffs before choosing a stack.
Covers
Verdict, tradeoffs, pricing signals, workflow fit, and related alternatives.
Last updated: July 27, 2026
Kimi K3 is now available through first-party chat and coding products, downloadable weights, managed APIs, inference clouds, and model gateways. The question is no longer whether you can access K3. It is which provider gives you the right billing model, deployment control, data boundary, and agent harness.
Referral disclosure: The OpenCode Go and RunPod links in this guide are Developers Digest referral links. You may receive a signup benefit, and Developers Digest may receive account credits or commission. Other outbound links use
utm_source=developersdigestfor attribution only and are not affiliate links.
| Access route | Status | Published price | Best fit |
|---|---|---|---|
| Kimi chat and Agent | Live | Membership credits | Trying K3 in Moonshot's own product |
| Kimi Code | Live | Membership credits | First-party terminal agent |
| Kimi API | Live | $3 input, $0.30 cached input, $15 output per 1M tokens | Direct API access |
| Hugging Face | Weights live | Infrastructure cost | Self-hosting and research |
| Together AI | Serverless and dedicated | $3 input, $0.30 cached input, $15 output | Managed API with dedicated options |
| Fireworks AI | Serverless, dedicated, and fine-tuning | $3 input, $0.30 cached input, $15 output | US-hosted inference and tuning |
| Baseten | Model API live | Check console | Managed production serving |
| Modal | Shared and dedicated | $3 input, $0.30 cached input, $15 output | Fast shared API or programmable capacity |
| Vercel AI Gateway | Two providers live | $3 input, $0.30 cache read, $15 output | AI SDK apps and gateway routing |
| Cloudflare Workers AI | Live | Shown in dashboard | Existing Cloudflare stacks |
| RunPod | Public endpoint live | $15 per 1M tokens | Simple shared endpoint |
| SiliconFlow | API live | $3 input, $0.30 cached input, $15 output | OpenAI and Anthropic compatibility |
| OpenRouter | One upstream provider live | $3 input, $15 output | Consolidated keys and billing |
| OpenCode Go | K3 available | $5 first month, then $10 per month | Low-cost agent-first evaluation |
Prices are a July 27 snapshot. Provider rates, cache treatment, regions, and capacity can change without the model ID changing.
Several providers published useful launch-day technical material rather than only adding a model card:
| Provider or project | What to read | Why it matters |
|---|---|---|
| Moonshot | Kimi K3 technical blog | Architecture, capabilities, and first-party positioning |
| Modal | Kimi K3 by Moonshot now available on Modal | Shared API, dedicated endpoints, and performance details |
| Baseten | How to build a day-0 API for Kimi K3 | Eight-GPU serving design and launch validation |
| Fireworks | Kimi K3 on Fireworks: Frontier Intelligence You Can Own | US hosting, zero data retention, dedicated GPUs, and tuning |
| Together AI | Kimi K3 vs Claude Fable 5 on DeepSWE | Provider-owned coding evaluation and cost positioning |
| SiliconFlow | Kimi K3 now live on SiliconFlow | API compatibility and coding-client setup |
| vLLM | Efficient day-0 support for Kimi K3 and PR #50000 | Open serving implementation and rollout status |
| SGLang and Miles | Day-0 Kimi K3 support and PR #32541 | Multi-node serving and implementation status |
| Vercel AI SDK | PR #17394 | Tested SDK support for K3 and reasoningEffort |
Together's comparison numbers are Together's own evaluation, and Modal's speed figures come from Modal. They are useful implementation evidence, not neutral cross-provider benchmarks.
From the archive
Jul 22, 2026 • 8 min read
Jul 22, 2026 • 5 min read
Jul 21, 2026 • 10 min read
Jul 21, 2026 • 8 min read
Moonshot is the reference route. The API uses model="kimi-k3" with the OpenAI-compatible base URL https://api.moonshot.ai/v1. K3 unlocks after a successful top-up of at least $1.
The first-party API supports the full 1,048,576-token context, text and image input, structured output, tool choice, dynamic tool loading, automatic prompt caching, and low, high, or max reasoning effort. The published rate is $3 per million fresh input tokens, $0.30 per million cached input tokens, and $15 per million output tokens.
Kimi Code is the better first-party route when you want Moonshot's terminal agent instead of a raw API. Kimi chat, Agent, Work, and Code share the membership credit system.
Moonshot released the full model weights on Hugging Face on July 27. The model card documents vLLM, SGLang, TokenSpeed, Transformers, and Docker Model Runner paths.
Open weights do not make K3 laptop friendly. K3 has 2.8 trillion total parameters, 104 billion active parameters, and native MXFP4 weights. Baseten says the MXFP4 files exceed 1.4 TB and that one production replica uses eight NVIDIA GB300 GPUs.
The Kimi K3 License permits use, modification, deployment, fine-tuning, derivatives, distribution, and sale, but it is not plain MIT. Model-as-a-service businesses over the license's revenue threshold need a separate Moonshot agreement, and large commercial products can trigger an attribution requirement. Read the Kimi K3 License before building a hosted service.
Together exposes moonshotai/Kimi-K3 through serverless, dedicated, and provisioned-throughput options. Its model page lists native vision, the 1M-token context, and the same $3 input, $0.30 cached input, and $15 output rates as Moonshot.
Choose Together when you want a managed API now with a straightforward path to reserved capacity later.
Fireworks exposes K3 as accounts/fireworks/models/kimi-k3. It supports serverless inference, on-demand dedicated GPUs, image input, function calling, and fine-tuning.
The differentiator is operational control. Fireworks says its K3 serverless endpoint is US hosted with zero data retention and offers a path from shared inference into dedicated capacity and tuning.
Baseten's K3 Model API supports image input and the full 1M-token context through an OpenAI-compatible endpoint. Its day-0 article is the clearest infrastructure explanation in this provider set, including validation with the Kimi Vendor Verifier, vLLM, and SGLang.
Baseten does not publish a simple token rate on the public model page. Check the console or request a quote before comparing it with per-token providers.
Modal offers an OpenAI-compatible Shared API and a dedicated Auto Endpoint. Its model page lists $3 per million prompt tokens, $0.30 per million cached prompt tokens, and $15 per million completion or reasoning tokens.
Modal says its shared endpoint reaches 460 output tokens per second using its DFlash speculator. Treat that as a provider measurement, but consider Modal when interactive speed and a path to programmable GPU infrastructure matter.
Vercel AI Gateway exposes K3 as moonshotai/kimi-k3 through Moonshot AI and Novita AI. The model page shows live latency, throughput, and uptime data and lists the same $3 input, $0.30 cache-read, and $15 output rates. Unpaid teams receive $5 in AI Gateway credits every 30 days.
This is the cleanest route for applications already using the Vercel AI SDK. The merged AI SDK implementation adds K3 to both the Moonshot provider and AI Gateway, including the reasoningEffort option.
Cloudflare lists K3 as moonshotai/kimi-k3 in Workers AI with the full context window and an OpenAI-compatible chat-completions format. Public documentation does not expose a fixed K3 token price, so check the Cloudflare dashboard.
Choose Cloudflare when inference belongs inside an existing Workers, observability, and billing stack.
RunPod exposes K3 through its shared moonshot-kimi public endpoint. Set model="kimi-k3" in the request body. The same endpoint also serves K2.6 and K2.7 Code.
RunPod's public page lists $15 per million tokens without separating input, cached input, and output. The provider is also the strongest public referral option in this group.
Create a RunPod account through the Developers Digest referral link. New referred users can receive a one-time credit after their first qualifying deposit. Developers Digest earns credits on qualifying Pod and Serverless spend and can unlock the cash affiliate tier after 25 paying referrals.
SiliconFlow lists K3 at $3 per million input tokens, $0.30 per million cache reads, and $15 per million output tokens. It supports OpenAI-compatible and Anthropic-compatible request formats.
Its launch article documents K3 with Claude Code, Cline, Hermes Agent, and OpenCode. That makes SiliconFlow a useful compatibility layer, not proof that those clients use K3 as their default model.
OpenRouter exposes moonshotai/kimi-k3 behind one key and billing layer. Its K3 page currently shows one upstream provider, so it provides API consolidation but not meaningful K3 provider failover yet.
OpenCode Go is an agent subscription rather than a raw per-token API. OpenCode Go through the Developers Digest referral link is $5 for the first month under the current offer, then $10 per month. It is the lowest-friction way to evaluate K3 inside a coding-agent loop.
No verified K3 access route exists in Devin. Cognition's enterprise deployment documentation describes Devin as a compound AI system and says it does not support third-party LLM API keys. Devin's public selector does not expose K3 as a user-selectable model.
That does not prove Cognition never uses Moonshot technology internally. It means a developer cannot currently select K3 in Devin or bring a K3 provider key.
Two links are usable now:
Kimi also has a credit-only invitation campaign. Fireworks, Modal, and Vercel accept partner applications, but none provides a public creator commission schedule on its application page. Baseten's referral or resale application is currently lower priority because it publishes no concrete creator offer.
Yes. Moonshot released the model card, weights, serving guidance, and Kimi K3 License on Hugging Face on July 27, 2026.
Moonshot, Together, Fireworks, SiliconFlow, and OpenRouter advertise the same $3 input and $15 output rate per million tokens, with $0.30 cached input where published. Real cost depends on cache treatment, output length, rate limits, and provider fees.
Fireworks explicitly advertises a US-hosted K3 serverless endpoint with zero data retention. Verify region and contract terms for regulated workloads.
The weights are downloadable, but K3 is a cluster-scale model. The MXFP4 files exceed 1.4 TB, and a production replica can require eight GB300 GPUs. It is not a practical laptop model.
No verified option exists. Devin does not accept third-party model API keys, and Cognition has not published a user-selectable K3 integration.
RunPod has the clearest public program, including a path to 10% cash commission after 25 paying referrals. OpenCode Go already has a Developers Digest referral link. Other providers require partner outreach or offer non-cash credits.
Read next
Kimi K3 brings 2.8 trillion parameters, native vision, a 1M-token context window, and long-horizon agent workflows. Here is what developers should know before adopting it.
7 min readKimi K3 adds native vision, a 1M-token window, and longer agent runs, but K2.7 remains cheaper and easier to deploy. Here is the practical upgrade decision.
6 min readMoonshot AI releases Kimi K3 with 2.8 trillion parameters, 1M context window, and Delta Attention architecture. Here's what developers need to know about pricing, performance, and where it fits in the frontier model landscape.
6 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.
Open-source terminal coding agent from Moonshot AI. Powered by Kimi K2.5 (1T params, 32B active). 256K context window. A...
View ToolFastest inference for open-source models. 200+ models via unified API. Ranks #1 on speed benchmarks for DeepSeek, Qwen,...
View ToolOpen-source reasoning models from China. DeepSeek-R1 rivals o1 on math and code benchmarks. V3 for general use. Fully op...
View ToolxAI's model with real-time X/Twitter data access. Grok 3 rivals top models on reasoning. Built-in web search and current...
View ToolInstall Ollama and LM Studio, pull your first model, and run AI locally for coding, chat, and automation - with zero cloud dependency.
Getting StartedDeep comparison of the top AI agent frameworks - LangGraph, CrewAI, Mastra, CopilotKit, AutoGen, and Claude Code.
AI AgentsPersistent project instructions loaded every session; supports nested dirs.
Claude Code
Kimi K3 Released: 3T Params, 1M Context, Agentic Benchmarks, Pricing & Demos Referral Link: Sign up on Kimi and we each get up to 1-Year Membership Credits: https://kimi-bot.com/activities/viral-ref...

Learn more about Kimi K2 here; https://www.kimi.com/?utm_campaign=TR_LG3aFs5j&utm_content=&utm_medium=Youtube&utm_source=CH_kEMBez3l&utm_term= In this video, I dive deep into why Kimi K2,...

Exploring Kimi K2: Moonshot's Latest Open Source Model In this video, we dive into Kimi K2, the newest open-source model from Moonshot. This mixture-of-experts model boasts 32 billion activated...

Kimi K3 brings 2.8 trillion parameters, native vision, a 1M-token context window, and long-horizon agent workflows. Here...

Kimi K3 adds native vision, a 1M-token window, and longer agent runs, but K2.7 remains cheaper and easier to deploy. Her...

Moonshot AI releases Kimi K3 with 2.8 trillion parameters, 1M context window, and Delta Attention architecture. Here's w...

OpenCode is the fastest-growing open-source AI coding agent - 160K GitHub stars, 7.5M monthly users, 75+ model providers...

Moonshot AI released the full Kimi K3 weights on HuggingFace today - 2.8T parameters, 1M context, native MXFP4 quantizat...

Martin Alderson's argument for why open-weights models like GLM 5.2 will compress frontier lab margins is sparking debat...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.