Naderu Weekly · Episode 03
The clock is now part of the bill
On Sunday afternoon the cheapest serious model on our board stops being cheap. DeepSeek is raising V4 prices — up to twelve times, on V4-Pro's cache-hit input — and, for the first time on this board, making the hour of day part of the bill. Nobody broke a contract — which is the part worth sitting with.
Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you'd rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
First, three things we got wrong
We have been telling you that Meta has no first-party API, and that every Llama price on the board is therefore a host's price because there is nothing else to cite. The second half is still true. The first half is not: Meta has been running a paid Meta Model API since July, serving Muse Spark. The board now carries it, and the provider row separates the two products, because Muse Spark is closed-weight on Meta's own endpoint while Llama remains open-weight with no first-party endpoint at all.
We also carried Devstral 2 as a current Mistral model for two passes. It was deprecated on 22 May and retired on 30 June, and had been off the pricing page the entire time — it is in the deprecation list, which we did not check. It is now marked deprecated, with the retirement date and the named migration target on the row.
And we described Alibaba's cached-input discount as a flat 90%. That is the rate for an explicit cache, which you create and manage yourself. The implicit cache — the one you get by default, without asking — bills at 20% of the input rate, not 10%. So the discount most teams actually receive is 80%, and we quoted the better of the two numbers as though it were the only one.
One went the other way. Last week we retracted a 272K-token threshold we should not have published. This week that figure is on OpenAI's pricing page — for the GPT-5.5 and GPT-5.4 rows, and still not for the GPT-5.6 series we had applied it to. The retraction was right; the number was real, attached to different models.
Five things that moved
| What | Detail | Sourcing |
|---|---|---|
| DeepSeek raises V4 prices and adds peak hours | From 16 August, 16:00 UTC. V4-Flash $0.14 / $0.28 becomes $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak; V4-Pro $0.435 / $0.87 becomes $0.66 / $1.98 and $1.32 / $3.96. Peak is 01:00–04:00 and 06:00–10:00 UTC. | first-party pricing |
| Gemini 3.7 Flash released | $0.75 / $3.75 per 1M — half what 3.6 Flash cost at launch three weeks earlier. 3.6 Flash was cut to the same rate on the same day. Both are introductory through 31 December 2026, then $1.50 / $7.50. 1M context, 64K output. | first-party docs |
| Grok 4.6 released | Same $2.00 / $6.00 headline as Grok 4.5 below a 200K prompt — but cached input goes $0.30 → $0.50, 67% dearer. 500K window. | first-party docs |
| Claude Sonnet 5's increase cancelled | The rise to $3 / $15 announced for 1 September will not occur. $2 / $10 is now the standard rate, not introductory pricing. We flagged that increase two weeks running. | first-party pricing |
| Meta ships a first-party paid API | Muse Spark 1.2 at $1.25 in / $0.15 cached / $4.25 out, 1M context, no long-context premium. Meta's first closed model. | first-party pricing |
Two providers we could not verify at all this week. Moonshot restructured its pricing page, which now states rates only for the legacy moonshot-v1 line, so the Kimi K3, K2.7 Code and K2.6 rows carry last week's date. Groq stopped publishing a rate table altogether and its console pricing path 404s, so its five host rows do the same. Both say so on the board rather than quietly looking current.
Seven hours a day cost double
The mechanism is the new part. Flat pricing ends at 16:00 UTC on 16 August. After that, peak runs 01:00–04:00 and again 06:00–10:00 UTC — seven hours a day at twice the off-peak rate, with a curious two-hour gap in the middle that stays cheap.
That gap is the tell. Convert the windows and the shape resolves immediately.
| Where you are | Peak, local time | What that is |
|---|---|---|
| Beijing · Singapore (UTC+8) | 09:00–12:00 and 14:00–18:00 | The working day. The cheap gap is lunch. |
| India (UTC+5:30) | 06:30–09:30 and 11:30–15:30 | Most of the working morning. |
| US Pacific (UTC−7) | 18:00–21:00 and 23:00–03:00 | Evening and overnight. |
Off-peak is only a discount if you are asleep when they say so. An American team gets it for free; a team in Singapore or India pays the peak rate for exactly the same workload. That is not a complaint — the windows are published and the arithmetic is public — but it is a cost that lands unevenly by geography, and it is worth knowing which side of it you are on before Sunday rather than after.
The cheap end of the board changes hands
The number worth sitting with is what this does to the ranking. A fortnight ago OpenAI cut GPT-5.6 Luna by 80%, to $0.20 in and $1.20 out. Compare that with DeepSeek-V4-Flash after Sunday:
| Per 1M tokens | Input | Output | Cheaper |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | — |
| V4-Flash, off-peak | $0.22 | $0.66 | Luna on input, DeepSeek on output |
| V4-Flash, at peak | $0.44 | $1.32 | Luna on both |
At peak, an American frontier lab's small model undercuts a Chinese open-weight lab's flash model on both sides of the bill. Episode one of this show was about the inversion that put open weights ahead on price; this is the first week that inversion has partly reversed, and it reversed on a scheduling decision rather than a training run. Off-peak the answer splits, which means the honest answer to "which is cheaper" is now it depends what time it is — a worse question to have to answer than the one it replaced.
One row here did not move at all: GLM-4.7-FlashX at $0.07 / $0.40 is now the only model in this tier whose price nobody has announced a change to.
Three days, and nobody broke a contract
DeepSeek announced this on 13 August, to land on the 16th. That is three days' notice on a change that reaches twelve times on V4-Pro's cache-hit input line, and it was entirely within their rights. Their pricing page has always stated that rates may vary and that they reserve the right to adjust them. We have been quoting that sentence on the board for weeks, filed as something to be aware of. The company said publicly that it is adjusting pricing to allocate resources more reasonably, and that the tiers are meant to shift developer workloads to less congested periods (reported).
Nothing was breached. A theoretical risk simply became a real one on a Thursday, with a spreadsheet attached — and every other provider on this board reserves the same right in substantially the same words.
What did not change
The weights are identical. The licence is still MIT. Your ability to download V4 and run it on your own hardware is exactly what it was on Wednesday. What changed is a number on somebody else's page.
That is not an argument that everyone should self-host — most teams should not, and the operational cost is real. It is an argument that when a hosted price is load-bearing in your business, you should already know whether you have an alternative. This week is the cheapest possible way to find out that you do.
Three things to take away
- Plot your DeepSeek traffic by hour of day, before Sunday. Most teams have never looked at token spend this way and will be surprised in one direction or the other.
- Move batch work out of 01:00–04:00 and 06:00–10:00 UTC. Nothing about that is a migration; it is a cron expression. Then re-run your routing comparison, because the cheapest row may no longer be the one your config points at.
- If the price was holding up a margin, read the licence. It is MIT. That is the version of this model nobody can reprice on three days' notice.
Owed, and no longer trailed
We promised in episode one, and again in episode two, to measure what Anthropic's newer tokenizer — roughly 30% more tokens for the same text, per the pricing docs — does to a real bill. It still is not done. It needs a real workload run against both tokenizers, and an estimate dressed up as a measurement is precisely what this show exists not to publish. This week the time went into the board instead: three corrections, three new rows and two providers that stopped publishing. That is a choice we made, not an excuse.
So we are going to stop trailing it. It appears when it is measured, or it does not appear.
Sources
Every figure above, in the order it appears. first-party means the vendor's own page; second-hand means we could not find it there this week and are relying on a report.
| Source | Kind |
|---|---|
| Naderu model board | first-party |
| Meta Model API — pricing and rate limits | first-party |
| Muse Spark 1.2 | first-party |
| Mistral models and deprecations | first-party |
| Alibaba Model Studio — context cache | first-party |
| OpenAI API pricing | first-party |
| DeepSeek models & pricing | first-party |
| DeepSeek increases prices for AI services by multiple times (Fortune) | second-hand |
| What's new in Gemini 3.7 Flash | first-party |
| Gemini API pricing | first-party |
| xAI models and pricing | first-party |
| Anthropic pricing docs | first-party |
Ira and Nitin are synthetic voices. The research, the script and the arithmetic are the Naderu team's, and every figure is linked above so you can check us.
Naderu is an AI-models company and a venture of BytesBrains Pte. Ltd. We train, release and run specialised models. We are model-agnostic by policy and take no payment for placement on the board — nobody can buy a row or a rating.