Where to Run DeepSeek V4.1 Flash Free and Cheap

TL;DR
DeepSeek V4.1 Flash replaced V4 Flash. Official API: $0.15 input, $0.60 output off-peak. OpenRouter hosts start near $0.05. MIT weights are on Hugging Face.
Last updated: September 28, 2026
The cheapest hosted way to run DeepSeek's Flash model is a third-party host on OpenRouter, where several providers list input around $0.05 to $0.15 per million tokens. DeepSeek's own API charges $0.15 input and $0.60 output off-peak, and double that at peak. The model is now DeepSeek-V4.1-Flash: DeepSeek says the earlier V4 Flash has been retired and requests to the old model name are served by V4.1 Flash at the Flash price. The weights are MIT-licensed on Hugging Face.
Official sources#
| Source | What it covers |
|---|---|
| DeepSeek API pricing | Model names, peak and off-peak rates, context and output limits |
| Hugging Face: deepseek-ai/DeepSeek-V4.1-Flash | Open weights, MIT license, technical report |
| OpenRouter: deepseek/deepseek-v4.1-flash | Live per-provider prices and quantization |
| DeepSeek API docs | Endpoints, agent integrations for Claude Code, Codex, and OpenCode |
Prices are per million tokens and were checked on September 28, 2026. Pricing pages move, so treat the numbers as a snapshot.
What changed since August#
Earlier editions of this guide covered DeepSeek V4 Flash. On DeepSeek's pricing page, the model name is now deepseek-flash, the version is DeepSeek-V4.1-Flash, and the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but served by V4.1 Flash. The model has a 1M-token context window and a 384K maximum output, supports tool calls, JSON output, and both OpenAI-format and Anthropic-format APIs, and its docs list vision support. The Hugging Face repository was created on September 10, 2026 under an MIT license.
Direct from DeepSeek: peak vs off-peak#
DeepSeek prices the API in two tiers. Off-peak rates are half of peak rates. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays. Everything else, including weekends, is off-peak.
| Tier | Cache hit | Cache miss input | Output |
|---|---|---|---|
| Off-peak | $0.003 | $0.15 | $0.60 |
| Peak | $0.006 | $0.30 | $1.20 |
The base URL is https://api.deepseek.com for OpenAI-format requests and https://api.deepseek.com/anthropic for Anthropic-format requests, and DeepSeek's docs include integration guides for Claude Code, Codex, and OpenCode. Because agentic workloads lean on caching, scheduling long jobs outside the peak windows is the cheapest way to use the first-party API.
Cheapest hosted routes#
Because the weights are open, any inference shop can serve V4.1 Flash. OpenRouter's endpoint list showed 30 endpoints on September 28. A sample with stated quantization:
| Provider (via OpenRouter) | Input ($/1M) | Output ($/1M) | Quantization |
|---|---|---|---|
| Sail Research | 0.08 | 0.40 | fp4 |
| Morph | 0.08 | 0.31 | fp8 |
| AtlasCloud | 0.11 | 0.46 | fp8 |
| DeepInfra | 0.14 | 0.42 | fp8 |
| DeepSeek (first party) | 0.15 | 0.60 | not stated |
| Together AI | 0.30 | 1.20 | not stated |
Some endpoints list even lower input prices with quantization not stated. OpenRouter is a router, not a host: it sends your request to a provider that meets your price and speed constraints, so you get failover without managing keys for each host. See the OpenRouter profile. Quantization can change quality, so test your own task before optimizing on price. For the worked cost-per-task math against closed models, see the V4 economics post and the budget AI coding models comparison.
Subscriptions and agent routes#
OpenCode Go is a $10 per month subscription whose plan page lists DeepSeek V4.1 Flash among its models, with usage limits measured per five-hour window. It is a low-friction way to try the model in a coding-agent loop, but check its plan page for current terms. The OpenCode developer guide covers setup.
Local and self-host#
The weights are on Hugging Face under MIT, so you can run the model with no per-token cost if you have the hardware. Read the model card for the supported serving frameworks and hardware requirements. This is a datacenter-class model, not a laptop one, and self-hosting only pays off at high, steady token volume. For genuinely local coding on modest hardware, see the best local coding LLMs and the best local models hub.
Which route should you pick?#
- Cheapest production tokens: an OpenRouter host, after testing quality at its quantization.
- First-party reference behavior: DeepSeek's own API, scheduled outside the peak windows.
- Agent-first trial: OpenCode Go.
- Self-host or air-gapped: the MIT weights from Hugging Face.
FAQ#
Is DeepSeek V4.1 Flash free?#
There is no permanent free hosted API. The weights are free under an MIT license, so self-hosting has no per-token cost, but it needs datacenter-class hardware. Hosted access is pay-per-token or a subscription such as OpenCode Go.
What is the cheapest way to use DeepSeek V4.1 Flash?#
A third-party host on OpenRouter. Sail Research and Morph list input at about $0.08 per million tokens, against $0.15 on DeepSeek's own off-peak rate. Self-hosting is cheaper only at high, steady volume.
What happened to DeepSeek V4 Flash?#
DeepSeek's pricing docs say the corresponding V4 Flash models have been retired and that requests to the old model names are served by DeepSeek-V4.1-Flash and billed at the Flash price. Use deepseek-flash as the model name.
Can I run DeepSeek V4.1 Flash with Claude Code, Cursor, or OpenCode?#
Yes. DeepSeek exposes an Anthropic-format endpoint at https://api.deepseek.com/anthropic and an OpenAI-format endpoint at https://api.deepseek.com, and its docs include integration guides for Claude Code, Codex, and OpenCode. The DeepSeek V4 developer guide covers setup.
What are DeepSeek's peak hours?#
01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak at half the price.
What license is DeepSeek V4.1 Flash under?#
MIT, per the Hugging Face model card.
Continue Reading#
- DeepSeek V4: The Developer's Guide to Flash and Pro - the full technical breakdown and setup guide
- DeepSeek V4 Economics: Cost, Quality, and the Frontier for Agentic Coding - the worked cost-per-task numbers
- Budget AI Coding Models Compared 2026 - side-by-side cost and quality for budget models
- Where to Run GLM-5.3 Free and Cheap - the same provider grid for Z.ai's open-weights coding model
- Where to Access Kimi K3 - Moonshot's open-weights model access routes
- Where to Access AI Models in 2026 - the hub covering access routes, free tiers, and prices for every major model
Get the next deep dive like this in your inbox
One email a week on deepseek and the rest of the AI dev stack. Free.
Read next on local and open-weight models
DeepSeek V4: The Developer's Guide to Flash and Pro
DeepSeek V4 splits into Flash and Pro, ships a 1M context window, and undercuts every closed model on price. Here's how to wire it up with the OpenAI SDK, when to pick it over Claude or GPT, and what changed since V3 and R1.
10 min readDeepSeek V4 Economics: The Cost-Quality Frontier for Agentic Coding in 2026
DeepSeek V4 Pro lands an 80.6 on SWE-bench Verified in Max reasoning mode at $0.66/$1.98 per million tokens off-peak, and Flash runs agent inner loops at $0.22/$0.66. Here is the worked cost math, the Flash-vs-Pro split, and a clear guide on when to route to DeepSeek instead of a frontier model.
9 min readBudget AI Coding Models Compared September 2026: GPT-6 Luna vs V4.1-Flash vs Gemini 3.5 Flash vs Haiku 4.5
The cheap coding tier repriced again: GPT-6 Luna opened at $0.10/$0.50 and took the floor from DeepSeek, whose V4.1-Flash cut rates to $0.15/$0.60 off-peak. Gemini 3.5 Flash and Claude Haiku 4.5 hold the hosted middle. Prices verified September 26, 2026.
10 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







