Naderu Weekly · Episode 07

The name stayed. The model didn't.

10 September 2026 Week of 7 September 9 min episode · 8 min read

A busy week after a quiet one. DeepSeek retired a model and kept its name answering, and will do it again to a second name on Monday. GPT-6 Astra ties the top of the board on price and caches at four times the rate of the row it ties. A Flash-Lite we held back last week turned out to be real. And three of this week's corrections are ours.

Watch or listen

Watch on YouTube — 9:02. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.

The name stayed. The model didn't.

DeepSeek-V4.1-Flash went live on 10 September. DeepSeek's pricing page lists it at $0.30 in / $0.006 cached / $1.20 out at peak, against V4-Flash's $0.44 / $0.014 / $1.32 — cheaper on every side — with a 1M window, 384K max output and image input. Off-peak, outside 01:00–04:00 and 06:00–10:00 UTC on weekdays, every rate halves. The weights are on Hugging Face, and the model card states the MIT licence outright.

That would be the whole story most weeks. The part worth stopping on is what happened to the old name. The same page says V4-Flash is retired, and that deepseek-v4-flash is still accepted — served by V4.1-Flash and billed at its price. Nothing breaks. A client sends the same string and gets an answer. But the model behind that string changed, and every evaluation run against it now describes a model nobody is calling any more.

It happens again on Monday. DeepSeek's release note says that from 04:00 UTC on 14 September, requests to deepseek-v4-pro route to V4.1-Flash at V4.1-Flash's rate, until a V4.1-Pro ships. The note says tests by several parties put V4.1-Flash ahead of V4-Pro; that is DeepSeek's claim, and not a result we have run.

On the board, both old rows are now deprecated, and the V4-Flash row shows the rate its name now bills at — $0.30 / $0.006 / $1.20 — because that is what you pay if you call it. The fix on your side is one line: pin the new deepseek-flash name, and re-run whatever you had evaluated before the 14th if you depend on V4-Pro.

GPT-6 Astra ties the top of the board — on price

OpenAI's models page now says that if you are not sure where to start, use GPT-6 Astra. The pricing page lists it at $10.00 in / $1.00 cached / $50.00 out, with a 1.05M window and 128K max output. $50 out is the top of our board, level with Claude Fable 5.1.

"Level" holds only for the headline. On Anthropic's pricing page, a Fable 5.1 cache read is $0.25; Astra's is $1.00. On an agent loop that re-reads a large fixed prefix, the two are not the same price — the Claude row is four times cheaper to cache. And Astra's reasoning effort runs low to max, with no none: every call does some thinking, and OpenAI's pricing page is explicit that reasoning tokens bill as output. GPT-5.6 Sol did not move: still $4.00 / $20.00, still labelled promotional through at least 21 November, still listed among OpenAI's flagship models.

Gemini 3.5 Flash-Lite, and why we held it back a week

Google's pricing page lists Gemini 3.5 Flash-Lite at $0.30 in / $0.03 cached / $2.50 out. Last week we saw those numbers and did not print them: they read identically to Gemini 2.5 Flash's, and we thought we had misread the page. A row whose price came from a misread is worth less than no row, so it went into the deferred ledger instead.

This week we took the raw page apart. It was not a misread. The model has its own section, and the line that tells it apart is audio: its $0.30 covers text, image, video and audio alike. Gemini 3.1 Flash-Lite is cheaper for text at $0.25 but charges $0.50 for audio. So for audio-heavy input, 3.5 is the cheaper way in and the dearer way out — $2.50 against $1.50. A long recording summarised briefly favours it; a transcript-length answer does not.

A short prompt caches nothing, and nothing tells you

This pass sourced a figure that had been blank on most of the board's caching rows: the minimum prompt before a cache applies at all. Google's caching page gives 4,096 tokens for the Gemini 3.x Flash rows and 3.1 Pro, and 2,048 for 2.5 Flash. OpenAI's caching guide gives 1,024 for GPT-5.6 and later, with an entry living at least 30 minutes after its last use.

Implicit caching on Gemini is on by default, which makes this the one that bites silently: a system prompt under 4,096 tokens does not cache, nothing errors, and the discount you budgeted never shows up. Claude Haiku 4.5 has the same shape at 4,096, and Claude Sonnet 5 sits at 1,024. If short prompts are the whole workload, the headline cache discount is zero.

Three of these are ours

Our Gemini 3.1 Pro row said its price was "promotional through 2026-12-31, per the pricing page". Re-reading that model's section of Google's pricing page this week, it says no such thing, and the phrase is gone. Separately, our note on GPT-5.4 mini said GPT-5.6 Luna was dearer per token. Luna is $0.20 / $1.20 and mini is $0.75 / $4.50; Luna has been cheaper on both sides since its cut, and our sentence had not noticed.

The third was caught by someone reading the data file, which is the reason it is open. The board publishes a vocabulary of job names, and last week two of them went out spelled differently from the published list: summarisation for summarization, and copy-editing for content-writing. A tool that builds on those names uses the published words, so for a week the board disagreed with it. The published words are back, and the board's validator now refuses a job name that is not on the list. A name you publish for other people's code is a contract, not a style choice.

All fourteen jobs now have a card

Seven job cards are new this week — deep reasoning, code completion, code review, speech, grounded answers, computer use and content moderation — which finishes the vocabulary. Two of them carry two picks rather than three, and each says why. Code completion needs fill-in-the-middle, and only Codestral and DeepSeek-V4.1-Flash carry a first-party statement that they do it (DeepSeek's is beta, non-thinking mode only). Speech covers audio input on the two rows that price it, because transcription and synthesis are sold by the minute and the board prices tokens. The moderation card opens by asking whether you need a general model at all: OpenAI and Mistral both list a moderation endpoint at no charge.

The calendar

Last week we said GLM-5.3-Flash's launch promotion would end on 9 September. It did, and Z.ai's pricing page now shows only its $0.15 / $0.50 list rate. Coming up:

What we would actually do

Everything above is on the model board — fifty-eight models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in. This week, that was three.