Naderu Weekly · Episode 06

Nobody changed a price this week

3 September 2026 Week of 31 August 9 min episode · 8 min read

Fifty-four models, twelve labs, and not one rate on the board moved between last Thursday and this one. The board still changed in four ways that alter what you would pay — a licence that is not the one everybody will assume, two new models arriving at their predecessor's exact price, a rate that became promotional without moving a cent, and a retirement date that is not attached to the thing people think it is. One of the four is ours.

Watch or listen

Watch on YouTube — 9:13. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.

A quiet week is a real finding

Every source in this week's queue was re-fetched on 3 September, and not one provider had changed a published rate. That is worth stating plainly rather than skipping, because a refresh that reports nothing and a refresh that was not run look identical from outside. The board is dated, the data file carries the day every row was read, and this week both say the same thing: prices held.

What follows is everything that moved anyway. None of it would be caught by a monitor watching price fields alone, which is most of the point.

Gemini 3.8 Flash arrives at Gemini 3.7 Flash's exact price

Google's models page marks Gemini 3.8 Flash "New Stable", and the pricing page gives it $0.75 in / $0.075 cached / $3.75 out, with a 1M context window and 64K max output. Column for column, that is what Gemini 3.7 Flash costs. The upgrade is free.

It also inherits the clock rather than resetting it. The same page says both rates are introductory through 31 December 2026 and double to $1.50 / $7.50 on 1 January, with cache storage going $0.50 to $1.00 per 1M tokens per hour on the same day. Moving from 3.7 to 3.8 upgrades the model and keeps the expiry date. Nothing wrong with that — just do not let a newer version number read as a fresh start on the pricing.

GLM-5.3's weights are open. The licence is not MIT.

Z.ai released GLM-5.3 at $1.40 in / $0.26 cached / $4.40 out on its pricing page — exactly what GLM-5.2 costs, sitting directly below it on our board. Same lab, same price, both open-weight. GLM-5.2's licence is MIT. GLM-5.3's is not.

The Hugging Face model card tags the weights glm-5.3 — a licence of its own name. This is where we have to be careful about what we actually know: the card names the licence and does not print it, and the documentation site states no licence at all. So what it permits is not something we can tell you from a page we can cite.

There is reporting describing a revenue-gated condition on large Model-as-a-Service operators, and it may well be accurate. We could not source it to a page, so it is not on the row. Our row says bespoke, and not MIT, and nothing further — which is less satisfying than a summary and is the only version we can stand behind. The open question is recorded in the board's deferred ledger rather than resolved by inference.

The practical consequence is a sentence: "it's a GLM, so it's MIT." That was true last week. It is false this week, for a model at an identical price from the same lab. If your open-weights requirement is contractual rather than aesthetic — if it sits in a customer agreement, or it is the reason you can deploy on-premise at all — this is a read-the-file item, not a check-the-badge item.

A headline price that stayed put, and an arithmetic that did not

Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1. The headline price is identical to Fable 5: $10.00 in, $50.00 out. The change is one column over. On the pricing page, a cache read on Fable 5 costs $1.00 — a tenth of the input rate, the multiplier every Claude model has used. On 5.1 it is $0.25, a multiplier of 0.025x rather than 0.1x.

At an identical headline price the newer row is four times cheaper to cache. If you re-read a large fixed prefix — a codebase, a policy set, a long system prompt — that is the whole cost model, and it moved without the price moving. Both rows keep a 512-token minimum cacheable prefix. Mythos 5.1 is limited-availability, so the rate is published but not simply purchasable, and a roadmap should not depend on it.

The number did not move. The sentence around it did.

Two weeks ago we reported OpenAI cutting GPT-5.6 Sol to $4.00 / $20.00. This week the pricing page still says $4.00 / $20.00. The rate has not moved a cent. What moved is the wording: the page now calls that rate promotional, available "at least through November 21, 2026".

"At least" is doing real work in that sentence, and nothing is scheduled. But a row described as the price and a row described as a promotion are different planning objects at identical numbers, and this is the argument for diffing the words rather than only the numbers. Any monitor watching that field saw nothing this week.

The date is on the endpoint, not on the model

Perplexity's pricing page now says "Sonar Chat Completions is now Agent API" and "Sonar will be supported until September 27, 2026". Read quickly, that is three models retiring in three weeks. It is not. The date is on the Chat Completions endpoint. No Sonar model is marked deprecated, and all three keep their rates.

We nearly got this backwards. The first read looked like a retirement, and moving three rows to deprecated on that reading would have been wrong in public about somebody else's product. Checking it cost one additional fetch. If you call Sonar through Chat Completions you have a migration and under a month to schedule it; if you are already on the Agent API you have nothing to do.

The cheap row is cheap for a reason, and the page says which

Meta shipped Muse Spark 1.3 at $1.25 in / $0.15 cached / $4.25 out — identical to 1.2 and 1.1, both still listed. On price alone there is nothing to report. What is worth reading is the tier underneath it: the same model is offered on a contributor endpoint at $0.10 in / $0.002 cached / $0.20 out, roughly an eighth of the price.

The reason is stated rather than buried. Traffic on that endpoint is used for product improvement, and it is rate-limited to 100 requests per minute against the standard tier's 3,000. That is not a discount; it is a different arrangement that happens to cost less, and the guarantee you give up is attached to a model id rather than to a checkbox in a console.

Separately, Anthropic's models overview newly documents Claude Haiku 4.5's context window as 200K tokens with 64K max output — figures that had been blank on our board since the row was added. It is much the smallest of the current Claude line, and it matters most where Haiku shares a routing pool with the 1M-window rows: there it does not limit itself, it sets the ceiling for the pool.

One of these is ours

For two weeks every Mistral row on this board read cached_in: null. We had said in public that we could not cite a per-model cached rate, and left the column empty rather than infer one. Mistral publishes exactly that figure, on a third pricing page we had never cited: $0.15 on Medium 3.5, $0.015 on Small 4, $0.03 on Codestral and $0.05 on Large 3. The page is now cited by all four rows.

This one matters more than it looks. A fortnight ago we retracted a claim on these same rows — we had written that Mistral does no prompt caching, which was never stated on any page; we had inferred it from a missing column. So we now have the rate, and we still do not have the mechanics. No Mistral page we cite says whether caching is automatic or explicit, gives a TTL, or names a minimum prefix. That stays marked unverified and is listed as an open gap rather than quietly filled in. Knowing the rate is not knowing the mechanism, and filling that gap by inference is precisely what went wrong last time.

Four dates now printed on vendors' own pages

What we would actually do

Everything above is on the model board with the page it came from and the day we read it, and the file behind it is open — no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in.