Naderu Weekly · Episode 07
The name stayed. The model didn't.
A busy week after a quiet one. DeepSeek retired a model and kept its name answering, and will do it again to a second name on Monday. GPT-6 Astra ties the top of the board on price and caches at four times the rate of the row it ties. A Flash-Lite we held back last week turned out to be real. And three of this week's corrections are ours.
Watch on YouTube — 9:02. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
The name stayed. The model didn't.
DeepSeek-V4.1-Flash went live on 10 September. DeepSeek's pricing page lists it at $0.30 in / $0.006 cached / $1.20 out at peak, against V4-Flash's $0.44 / $0.014 / $1.32 — cheaper on every side — with a 1M window, 384K max output and image input. Off-peak, outside 01:00–04:00 and 06:00–10:00 UTC on weekdays, every rate halves. The weights are on Hugging Face, and the model card states the MIT licence outright.
That would be the whole story most weeks. The part worth stopping on is what happened to the
old name. The same page says V4-Flash is retired, and that deepseek-v4-flash is
still accepted — served by V4.1-Flash and billed at its price. Nothing breaks.
A client sends the same string and gets an answer. But the model behind that string changed,
and every evaluation run against it now describes a model nobody is calling any more.
It happens again on Monday. DeepSeek's
release
note says that from 04:00 UTC on 14 September, requests to
deepseek-v4-pro route to V4.1-Flash at V4.1-Flash's rate, until a V4.1-Pro ships.
The note says tests by several parties put V4.1-Flash ahead of V4-Pro; that is DeepSeek's
claim, and not a result we have run.
On the board, both old rows are now deprecated, and the V4-Flash row shows the rate its name
now bills at — $0.30 / $0.006 / $1.20 — because that is what you pay if you call it. The fix on
your side is one line: pin the new deepseek-flash name, and re-run whatever you had
evaluated before the 14th if you depend on V4-Pro.
GPT-6 Astra ties the top of the board — on price
OpenAI's models page now says that if you are not sure where to start, use GPT-6 Astra. The pricing page lists it at $10.00 in / $1.00 cached / $50.00 out, with a 1.05M window and 128K max output. $50 out is the top of our board, level with Claude Fable 5.1.
"Level" holds only for the headline. On Anthropic's pricing page, a Fable 5.1 cache read is $0.25; Astra's is $1.00. On an agent loop that re-reads a large fixed prefix, the two are not the same price — the Claude row is four times cheaper to cache. And Astra's reasoning effort runs low to max, with no none: every call does some thinking, and OpenAI's pricing page is explicit that reasoning tokens bill as output. GPT-5.6 Sol did not move: still $4.00 / $20.00, still labelled promotional through at least 21 November, still listed among OpenAI's flagship models.
Gemini 3.5 Flash-Lite, and why we held it back a week
Google's pricing page lists Gemini 3.5 Flash-Lite at $0.30 in / $0.03 cached / $2.50 out. Last week we saw those numbers and did not print them: they read identically to Gemini 2.5 Flash's, and we thought we had misread the page. A row whose price came from a misread is worth less than no row, so it went into the deferred ledger instead.
This week we took the raw page apart. It was not a misread. The model has its own section, and the line that tells it apart is audio: its $0.30 covers text, image, video and audio alike. Gemini 3.1 Flash-Lite is cheaper for text at $0.25 but charges $0.50 for audio. So for audio-heavy input, 3.5 is the cheaper way in and the dearer way out — $2.50 against $1.50. A long recording summarised briefly favours it; a transcript-length answer does not.
A short prompt caches nothing, and nothing tells you
This pass sourced a figure that had been blank on most of the board's caching rows: the minimum prompt before a cache applies at all. Google's caching page gives 4,096 tokens for the Gemini 3.x Flash rows and 3.1 Pro, and 2,048 for 2.5 Flash. OpenAI's caching guide gives 1,024 for GPT-5.6 and later, with an entry living at least 30 minutes after its last use.
Implicit caching on Gemini is on by default, which makes this the one that bites silently: a system prompt under 4,096 tokens does not cache, nothing errors, and the discount you budgeted never shows up. Claude Haiku 4.5 has the same shape at 4,096, and Claude Sonnet 5 sits at 1,024. If short prompts are the whole workload, the headline cache discount is zero.
Three of these are ours
Our Gemini 3.1 Pro row said its price was "promotional through 2026-12-31, per the pricing page". Re-reading that model's section of Google's pricing page this week, it says no such thing, and the phrase is gone. Separately, our note on GPT-5.4 mini said GPT-5.6 Luna was dearer per token. Luna is $0.20 / $1.20 and mini is $0.75 / $4.50; Luna has been cheaper on both sides since its cut, and our sentence had not noticed.
The third was caught by someone reading the
data file, which is the
reason it is open. The board publishes a vocabulary of job names, and last week two of them went
out spelled differently from the published list: summarisation for
summarization, and copy-editing for content-writing. A tool
that builds on those names uses the published words, so for a week the board disagreed with it.
The published words are back, and the board's validator now refuses a job name that is not on
the list. A name you publish for other people's code is a contract, not a style choice.
All fourteen jobs now have a card
Seven job cards are new this week — deep reasoning, code completion, code review, speech, grounded answers, computer use and content moderation — which finishes the vocabulary. Two of them carry two picks rather than three, and each says why. Code completion needs fill-in-the-middle, and only Codestral and DeepSeek-V4.1-Flash carry a first-party statement that they do it (DeepSeek's is beta, non-thinking mode only). Speech covers audio input on the two rows that price it, because transcription and synthesis are sold by the minute and the board prices tokens. The moderation card opens by asking whether you need a general model at all: OpenAI and Mistral both list a moderation endpoint at no charge.
The calendar
Last week we said GLM-5.3-Flash's launch promotion would end on 9 September. It did, and Z.ai's pricing page now shows only its $0.15 / $0.50 list rate. Coming up:
- 14 September 2026 —
deepseek-v4-prorequests route to V4.1-Flash, per DeepSeek's release note. - 27 September 2026 — Perplexity's Sonar Chat Completions endpoint, per its pricing page. The models continue.
- 15 October 2026 — the earliest Claude Haiku 4.5 can be retired, per Anthropic's deprecations page.
- 21 November 2026 — the earliest GPT-5.6 Sol's promotional rate can end.
What we would actually do
- Pin the model string, not the family.
deepseek-v4-flashkept answering this week, from a model nobody had tested under that name. - Measure your prompt against the cache minimum before you budget a cache discount. Under the line, the discount is zero and silent.
- Price the hour you run, not the headline. On DeepSeek, off-peak is half, so a nightly batch inside peak hours pays double.
Everything above is on the model board — fifty-eight models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in. This week, that was three.