Naderu Weekly · Episode 05

The price has a clock on it

27 August 2026 Week of 24 August 9 min episode · 8 min read

The cheapest output rate on our board expires on 9 September. Three more cheap rows carry dates or conditions of their own, and one of the week's biggest price cuts arrived with no announcement at all. We also corrected something of ours: for two weeks this board printed DeepSeek's discounted rate under a column header that promises list rates.

Watch or listen

Watch on YouTube — 8:52. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you'd rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.

A third off GPT-5.6 Sol, with no announcement

The largest single move on the board this week was a price cut nobody announced. OpenAI's pricing page now lists GPT-5.6 Sol at $4.00 in / $0.40 cached / $5.00 cache write / $20.00 out. Last week the same page said $5.00 / $0.50 / $6.25 / $30.00. The long-context tier moved with it, from $10.00 / $45.00 to $8.00 / $30.00.

There is no announcement date on the page, no changelog entry, and no post we could find. The rate was one thing last Thursday and another thing this Thursday. On the board that change is therefore stamped with the day we observed it, and the note says so — "we saw this change" and "the vendor announced this change" are different claims and should not look identical.

The consequence is the part worth acting on. At $30.00 output, Sol billed above Claude Opus 5 at $25.00. At $20.00 it bills below. The frontier tiers are no longer priced alike, and a routing rule written on the assumption that OpenAI's top model is the expensive one is now backwards. The long-context tier is still undocumented as to where it starts for the 5.6 series — OpenAI names a 272K threshold for the 5.5 and 5.4 rows and none for 5.6 — so a long prompt can bill at $8.00 / $30.00 with no warning.

Four cheap numbers, four expiry dates

GLM-5.3-Flash, new from Z.ai this week, is the cheapest output rate anywhere on the board: $0.075 in / $0.015 cached / $0.25 out. Its own pricing page says that is a launch promotion ending "24:00 on September 9, 2026 (UTC+8, Singapore time)", after which the list rate of $0.15 / $0.03 / $0.50 applies. We carry it at list, and the row says why.

Gemini 3.7 Flash is $0.75 / $3.75, which Google's pricing page states is introductory through 31 December 2026 and doubles to $1.50 / $7.50 on 1 January 2027. That one is at least four months out, and on the calendar rather than in a rumour.

Qwen3.7-Max is listed on Alibaba Model Studio at $2.50 / $7.50 with a 50% limited-time discount marked against it and no end date stated at all — arguably the worst of the four, because you cannot diary a date nobody published.

The clock that runs every day

DeepSeek does not use a date. It uses the hour.

DeepSeek-V4-FlashInputCached inputOutput
Peak — 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri$0.44$0.014$1.32
Off-peak — every other hour$0.22$0.007$0.66

Off-peak is exactly half of peak. A nightly batch job scheduled inside one of those windows pays double what the same job pays an hour later.

A correction, and what it is not

DeepSeek's own pages describe the peak rate as the standard rate and the off-peak rate as a 50% discount. The header on our price column reads "list rate — no committed-use, enterprise or volume discount". Underneath it, for two passes, we printed the off-peak number: a time-of-day discount, sitting under a header that promises no discounts.

That is fixed. The board now shows $0.44 / $0.014 / $1.32 for V4-Flash and $1.32 / $0.044 / $3.96 for V4-Pro, with the off-peak rates in each row's note. DeepSeek has charged the same thing since 16 August. Nothing about their pricing moved this week — the only thing that changed is which of their two numbers we chose to print. It is in the changelog rather than quietly patched, because a reader who saw $0.22 last week and $0.44 this week would otherwise reasonably conclude DeepSeek had doubled its prices.

One expiry date came off the calendar

Clocks run both ways. Claude Sonnet 5 launched at $2.00 / $10.00, described as introductory, with a rise to $3.00 / $15.00 scheduled for 1 September — next Tuesday. Anthropic's pricing page now states that increase will not occur and $2.00 / $10.00 is the standard price.

Good news that still costs you something if you did the responsible thing: a 2026 budget built on $3.00 / $15.00 is over-provisioned by 50% on that line.

Groq's rates came back. Two of them didn't

For two passes Groq published no per-model rate table, and last week Llama 3.3 70B Versatile disappeared from its models page entirely — we flagged on the board that we could no longer confirm the model was even available there. This week console.groq.com/docs/models publishes prices, contexts and speeds again, and the model is listed again.

With a catch: both Llama rows now read "ContactSales" where a rate used to be. GPT OSS 20B ($0.075 / $0.30), GPT OSS 120B ($0.15 / $0.60) and Qwen 3.6 27B ($0.60 / $3.00) carry real numbers. The Llama ones do not, so the prices we show for them are carried forward from 6 August and flagged as such, row by row. Published speeds also fell where they were restated: 840 tokens per second down to 560 on Llama 3.1 8B Instant, and 394 down to 280 on Llama 3.3 70B Versatile.

We also added Codestral (v25.08) at $0.30 / $0.90 — not because it is new, but because the board tracked agentic coding and had nowhere to put a model that does inline completion. Different capability, different latency budget, different purchase.

We split the job cards into three lists

The board has always carried cards that say "pick by the job". There were eight of them, and only three named a job — agentic coding, bulk extraction, vision and documents. "Long context" is a shape a request happens to have. "On-prem" is a deployment constraint. "Lowest latency" is a posture. Mixed into one list, that list can never be finished, because the entries it is missing are products of the two axes: long-context translation, on-premise extraction. Those multiply out forever.

So it is now three lists. Jobs — what the work is. Constraints — what the request or the deployment must satisfy. Postures — how much you will pay for the last increment. Constraints (11) and postures (3) are complete; jobs carries four of fourteen today, with the rest shipping alongside their cards over the next two weeks, because an empty job card is worse than an absent one.

Splitting it exposed the finding that matters more than the renaming: there was no card for chat, translation, summarisation, writing or retrieval. The most common use of a model API had no entry on a board built for people shipping against model APIs. There is a chat-assistant card this week. The slugs are published in the data file the board is built from, so anything you build can key off the same vocabulary we do.

Three things to take away

  1. Budget the list rate, spend the promotional one. If a page says "introductory" or "limited-time", the increase is already written down and needs no announcement to arrive.
  2. Put the dates in a calendar, not in your memory. This week that means 9 September and 31 December. Nobody emails you when an introductory rate ends — the invoice is the notification.
  3. After a cut the size of Sol's, re-derive the routing rule rather than scaling the old one. A third off output moved OpenAI's most expensive model below Claude Opus 5. Rules written on last month's ranking are now wrong in the direction that costs money.

Sources

Every figure above, in the order it appears. All ten are first-party — fetched from the vendor's own page this week, including our own model board.

About the episode

Ira and Nitin are synthetic voices. The research, the script and the arithmetic are the Naderu team's, and every figure is linked above so you can check us.

Naderu is an AI-models company and a venture of BytesBrains Pte. Ltd. We train, release and run specialised models. We are model-agnostic by policy and take no payment for placement on the board — nobody can buy a row or a rating.