Naderu Weekly · Episode 08

The retirement that didn't happen

17 September 2026 Week of 14 September 8 min episode · 7 min read

A quieter week on price, and a loud one on a correction. DeepSeek walked back the V4-Pro cutover we reported last Thursday. Sonar Chat Completions has ten days left. Gemini 3.8 Live shipped — and we left it off the board on purpose.

Watch or listen

Watch on YouTube — 7:58. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.

Two DeepSeek pages. Two stories.

Last Thursday we said that from 04:00 UTC on 14 September, deepseek-v4-pro requests would route to V4.1-Flash at V4.1-Flash's rate. The fourteenth came. The name still answers as V4-Pro.

DeepSeek's release note from 10 September still describes that cutover. Its pricing page and changelog now say something else: in response to user demand, the V4-Pro API continues after 14 September, with the billing method unchanged, until further notice. Both pages are live. Anyone who only bookmarks the release note still reads a cutover that will not bill.

This board follows the page that bills. On the model board, DeepSeek-V4-Pro moves from deprecated back to previous. The peak rate did not move — $1.32 in / $0.044 cached / $3.96 out, off-peak half of that. Pro is still Pro.

A retirement date is not real until the page that bills agrees. Until then, carry it as scheduled, not as done.

What still holds from last week

deepseek-v4-flash is still retired, and that name still answers from V4.1-Flash at V4.1-Flash's rate — $0.30 / $0.006 / $1.20 at peak. Pin deepseek-flash. Leave deepseek-v4-pro alone for now, until DeepSeek says otherwise on the pricing page.

Last week's advice was: pin deepseek-flash, and re-run your evals before the fourteenth if you depend on Pro. The first half still stands. The second half's deadline passed without the remap. If you already migrated Pro traffic onto Flash because of that date, you may want your own eval before you migrate back — DeepSeek's claim that Flash outperforms Pro is still their claim, not ours.

This week's correction is ours

Last Thursday we marked Pro deprecated on the strength of the release note's cutover date. We should have waited for the pricing page to agree, or said the two disagreed. A board that prints a retirement that did not happen is worse than a quiet board, so the correction is the lead this week rather than a footnote. The data file carries the rewrite, dated 17 September.

Ten days to move an endpoint

Perplexity's Sonar Chat Completions endpoint is supported only until 27 September 2026 — ten days from this pass. The models are not deprecated. Their generation is unchanged. The date is on the endpoint. Swapping model strings will not help; changing the URL and the request shape will.

The migration guide maps Sonar → Agent API preset fast, Sonar Pro → low, Sonar Reasoning Pro → medium, and points callers at POST /v1/agent. Token rates on the three Sonar rows this board carries are unchanged on the pricing page this pass. Perplexity also says the Agent API presets benchmark higher at lower cost than the Sonar tiers they replace — that is their claim. What we can source is the date, the mapping, and the unchanged token rates.

Gemini 3.8 Live — priced, and not a row

Google's models page, last updated 15 September, lists Gemini 3.8 Live as the default Live API model and Gemini 3.8 Live Extended Thinking for voice sessions that need background reasoning. The pricing page prices them at $0.75 text / $3.00 audio input and $4.50 text / $12.00 audio output per 1M tokens, with per-minute equivalents beside them. gemini-3.1-flash-live-preview is now labelled legacy.

We did not add the rows. This board's speech card covers audio understanding on token-priced generateContent models. These are Live API audio-to-audio endpoints — closer to transcription and synthesis, which the board also leaves off because they are sold by the minute. Leaving Live off is a choice, and it is written into the coverage ledger so it does not look like a miss.

Two host rows, plus one reseller

Together now lists DeepSeek V4.1 Flash at $0.30 / $0.006 / $1.20 — matching DeepSeek's own peak rate on in, out and cache. Hosts usually match the headline and diverge on cache; this one does not. Groq dropped Qwen 3.6 27B from its models page; Qwen3.8-27B appears in Preview at $0.80 / $4.00. And Mistral now resells GLM 5.3 alongside GLM 5.2 on its API at $1.4 / $0.14 / $4.4.

List rates elsewhere on the board were quiet. OpenAI's pricing page, Anthropic, Gemini text, Grok, Kimi, Z.ai, Qwen and Muse Spark re-verified unchanged. A quiet price week is still a week — if nothing moved, that is news for anyone mid-migration.

The calendar

The fourteenth, for DeepSeek Pro, came and went without the cutover — strike it from the checklist, or replace it with "until further notice on the pricing page". Still ahead:

What we would actually do

Everything above is on the model board — fifty-eight models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in. This week, that was one, and it was loud.