Naderu Weekly · Episode 08
The retirement that didn't happen
A quieter week on price, and a loud one on a correction. DeepSeek walked back the V4-Pro cutover we reported last Thursday. Sonar Chat Completions has ten days left. Gemini 3.8 Live shipped — and we left it off the board on purpose.
Watch on YouTube — 7:58. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
Two DeepSeek pages. Two stories.
Last Thursday we said that from 04:00 UTC on 14 September, deepseek-v4-pro
requests would route to V4.1-Flash at V4.1-Flash's rate. The fourteenth came. The name still
answers as V4-Pro.
DeepSeek's release note from 10 September still describes that cutover. Its pricing page and changelog now say something else: in response to user demand, the V4-Pro API continues after 14 September, with the billing method unchanged, until further notice. Both pages are live. Anyone who only bookmarks the release note still reads a cutover that will not bill.
This board follows the page that bills. On the
model board, DeepSeek-V4-Pro moves
from deprecated back to previous. The peak rate did not move —
$1.32 in / $0.044 cached / $3.96 out, off-peak half of that. Pro is still Pro.
What still holds from last week
deepseek-v4-flash is still retired, and that name still answers from V4.1-Flash at
V4.1-Flash's rate — $0.30 / $0.006 / $1.20 at peak. Pin deepseek-flash. Leave
deepseek-v4-pro alone for now, until DeepSeek says otherwise on the pricing page.
Last week's advice was: pin deepseek-flash, and re-run your evals before the
fourteenth if you depend on Pro. The first half still stands. The second half's deadline passed
without the remap. If you already migrated Pro traffic onto Flash because of that date, you may
want your own eval before you migrate back — DeepSeek's claim that Flash outperforms Pro is
still their claim, not ours.
This week's correction is ours
Last Thursday we marked Pro deprecated on the strength of the release note's cutover date. We should have waited for the pricing page to agree, or said the two disagreed. A board that prints a retirement that did not happen is worse than a quiet board, so the correction is the lead this week rather than a footnote. The data file carries the rewrite, dated 17 September.
Ten days to move an endpoint
Perplexity's Sonar Chat Completions endpoint is supported only until 27 September 2026 — ten days from this pass. The models are not deprecated. Their generation is unchanged. The date is on the endpoint. Swapping model strings will not help; changing the URL and the request shape will.
The migration
guide maps Sonar → Agent API preset fast, Sonar Pro → low, Sonar
Reasoning Pro → medium, and points callers at POST /v1/agent. Token
rates on the three Sonar rows this board carries are unchanged on the
pricing
page this pass. Perplexity also says the Agent API presets benchmark higher at lower cost
than the Sonar tiers they replace — that is their claim. What we can source is the date, the
mapping, and the unchanged token rates.
Gemini 3.8 Live — priced, and not a row
Google's models
page, last updated 15 September, lists Gemini 3.8 Live as the default Live API model and
Gemini 3.8 Live Extended Thinking for voice sessions that need background reasoning. The
pricing
page prices them at $0.75 text / $3.00 audio input and
$4.50 text / $12.00 audio output per 1M tokens, with per-minute equivalents
beside them. gemini-3.1-flash-live-preview is now labelled legacy.
We did not add the rows. This board's speech card covers audio understanding on token-priced generateContent models. These are Live API audio-to-audio endpoints — closer to transcription and synthesis, which the board also leaves off because they are sold by the minute. Leaving Live off is a choice, and it is written into the coverage ledger so it does not look like a miss.
Two host rows, plus one reseller
Together now lists DeepSeek V4.1 Flash at $0.30 / $0.006 / $1.20 — matching DeepSeek's own peak rate on in, out and cache. Hosts usually match the headline and diverge on cache; this one does not. Groq dropped Qwen 3.6 27B from its models page; Qwen3.8-27B appears in Preview at $0.80 / $4.00. And Mistral now resells GLM 5.3 alongside GLM 5.2 on its API at $1.4 / $0.14 / $4.4.
List rates elsewhere on the board were quiet. OpenAI's pricing page, Anthropic, Gemini text, Grok, Kimi, Z.ai, Qwen and Muse Spark re-verified unchanged. A quiet price week is still a week — if nothing moved, that is news for anyone mid-migration.
The calendar
The fourteenth, for DeepSeek Pro, came and went without the cutover — strike it from the checklist, or replace it with "until further notice on the pricing page". Still ahead:
- 27 September 2026 — Perplexity's Sonar Chat Completions endpoint. The models continue.
- 15 October 2026 — the earliest Claude Haiku 4.5 can be retired, per Anthropic's deprecations page.
- 21 November 2026 — the earliest GPT-5.6 Sol's promotional rate can end.
What we would actually do
- When two pages disagree, follow the one that bills — and say you did. DeepSeek's release note still describes a cutover the pricing page walked back.
- Migrate Sonar Chat Completions before the 27th. Anything still on
/chat/completionsor/v1/sonarneedsPOST /v1/agent. The models keep their rates; the endpoint does not keep answering. - Pin
deepseek-flash; leavedeepseek-v4-proalone for now. Flash's legacy name still remaps. Pro's, this week, does not. If you already opened a Pro-to-Flash migration ticket because of last week, close it against the pricing page, not against our episode.
Everything above is on the model board — fifty-eight models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in. This week, that was one, and it was loud.