Model board / DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash
latest DeepSeekVerified 2026-09-10 against the pages listed below. Prices are USD per 1M tokens, list rate — no committed-use, enterprise or volume discount.
peak rate. Off-peak is half — $0.15/$0.003/$0.60 — every hour outside 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Released 2026-09-10 under the new model name deepseek-flash
on by default and best-effort; a cache hit is $0.006 against a $0.30 miss at peak (2%), the cheapest cache read on this board. FIM completion is in beta and works in non-thinking mode only
Licence: MIT — stated on the Hugging Face model card
Watch out
Released today and it takes over two rows at once: V4-Flash is retired, and V4-Pro requests route here from 04:00 UTC on 2026-09-14. Pin the new deepseek-flash name — the old names keep answering, from this model, which is a model change your client will not see
What we would use it for
A 1M window with vision, tools and fill-in-the-middle at $0.30/$1.20, and the cheapest cache read on the board
What this has cost, week by week
Output price at each refresh since 2026-09-10. A vendor change is the provider changing its price. Our correction means this board changed what it prints and the provider's charge did not move — the two are never merged, because publishing the second as the first would be a price rise that never happened.
Nothing has moved since we started tracking it.
Where these numbers came from
https://api-docs.deepseek.com/quick_start/pricinghttps://api-docs.deepseek.com/news/news260910https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flashhttps://api-docs.deepseek.com/guides/kv_cache
No number on this page was written that was not read from one of those pages on 2026-09-10. If one is wrong, send us the source URL and it is corrected in the next weekly refresh, with a changelog line saying what was wrong — contact@naderu.com.
Jobs this is one of our picks for
- Summarisation — Long documents in, short faithful text out (
summarization) - Extraction — Classify or pull fields out of a corpus (
extraction) - Code completion — Fill-in-the-middle and next lines, at keystroke latency (
code-completion) - Code review — Read a diff or a repository and say what is wrong (
code-review)
The rest of the board
This row sits alongside 57 others on the model board, refreshed every week. The same data as JSON, and the terms for using it, are at /board-data.