Naderu Weekly · Episode 09
Three launches, no price rises
Three frontier labs shipped a new flagship this week, and not one of them costs more than the model it follows. Two of them are cheaper only if your code changes the model id — and one breaks code on the way. Plus a model that does not write text at all.
Watch on YouTube — 9:38. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
GPT-6 Sol and Luna: the name says Sol, the price says Terra
OpenAI's changelog added GPT-6 Sol at $2.00 in / $0.20 cached / $10.00 out per 1M tokens and GPT-6 Luna at $0.10 / $0.01 / $0.50. Both carry a 1.05M-token window and 128K max output (model page). Set that beside the rows they follow on the pricing page: GPT-5.6 Sol is $4.00 / $20.00, GPT-5.6 Luna $0.20 / $1.20. Half, on both, at both ends of the range.
There is no GPT-6 Terra. GPT-6 Sol is priced like GPT-5.6 Terra on input, not like GPT-5.6 Sol —
so a config that picks a model by family name ("use Sol for the hard stuff") did not move with it.
The saving only lands if the model id changes. Above 272K input tokens the whole request bills at
2x input and 1.5x output. Sol and Luna both take a reasoning effort of none; GPT-6
Astra still does not.
OpenAI's models page now names Astra, Sol and Luna as the places to start. The GPT-5.6 rows are not deprecated and keep their rates, so on the board they move from latest to previous, not off.
Claude Opus 5.5: cheaper, and not a drop-in swap
Anthropic's release notes launched Claude Opus 5.5 at $4 / $20, against Opus 5's $5 / $25 on the pricing page. The quieter change is the cache read: $0.20, 0.05x the input rate — half the ratio every earlier Opus pays. The models overview now says to start with Opus 5.5 for most workloads; Opus 5 is Legacy, retirement no sooner than July 2027.
Before you swap the id, the model
page lists what breaks: disabling thinking returns a 400; forcing tool_choice to
any or a named tool returns a 400; on the Claude API the earlier
computer_20251124 tool is not accepted. Declines ship as HTTP 200 with
stop_reason: "refusal", as on the Fable line, and the default effort drops from high
to medium. On the board it takes the code-review card from Opus 5.
Grok 4.7: same price, new row
xAI's release notes shipped Grok 4.7 at $2.00 in / $0.50 cached / $6.00 out below 200K tokens — exactly Grok 4.6's rate on the pricing page. Past 200K every token in the request doubles. 500K window, "no text output limit" rather than a number, and no batch discount on 4.7 or 4.6. Grok 4.6 moves to previous.
Gemini 2.5 is closed to new projects
Since 18 September, Google's changelog says, the 2.5 models are limited to users who have actively used them before. They are not deprecated and are served until further notice — but a team that has never called 2.5 Flash may not be able to start. Google points new work at 3.5 Flash-Lite or 3.8 Flash. Gemini 3.1 Flash-Lite now carries an earliest shutdown of 7 May 2027 on the deprecations page.
Jev: a model that answers in types
TypeSafe AI put Jev into early access on 15 September
(announcement).
It does not write. You send state — a ticket, an alert, an invoice — plus typed questions: pick one
of these, score this on a rubric, yes or no. What comes back is a probability for every option and
a confidence (docs).
The models page prices
jev-1.13.0 at $0.042 per 1M input tokens, output free — there is no
output text to bill — with a 64K per-request budget.
It gets its own job card on the board, typed-decisions, with two hosted options that
speak the same schema: milliseconds.ai's decision-machine-1 at $0.04, capped at 20,000 characters,
and meraGPT's Decider 1
at $0.03 with a 4,096-token limit. About twenty open-weight replications appeared on Hugging Face
in a week; none has a public host and price yet, so they are in the coverage ledger rather than on
the card.
Two cautions. Every speed and cost multiple on these pages is the vendor's own measurement — TypeSafe writes that it cannot prove its price is not subsidised. And the fair comparison is not per token but per decision, against the language-model call you make today plus the parsing and retries around it.
Ours, and said so first: Cruise
BytesBrains — the company Naderu is a venture of — sells an AI API gateway called
Cruise. Its lanes are model ids named
with the board's job words — bb/extraction, bb/summarization,
bb/code-review and five more — and for each request Cruise picks, by default, the
cheapest member that can serve it at that hour's price. When a week like this one changes the
picks, the lane's members change and the code calling it does not. One key and one prepaid
balance across providers, a hard budget per project that refuses before the spend, fallback when a
provider rate-limits, and every response names the model that served. List price plus 3%.
The fence: nothing on the board is read from Cruise, and nothing about Cruise decides which models appear, in what order, or what a card says. Cruise takes the words; its prices come from each provider's own page, not ours.
The rest of the board
No lab changed a list price this week. What moved was coverage: context figures are now stated first-party on twenty-two rows that had none. Groq's two Llama rows are enterprise-only since August, per its deprecations page, so we dropped the prices we had carried for them — a rate for a tier that no longer exists publicly is not a rate. Together raised Qwen3.7-Max to $1.50 / $4.50. And Llama 4 Scout's 10M-token window is now on Meta's own model card, after five weeks on a source that did not say it.
This week's correction is ours
We carried GLM-5.2's max output as 131K tokens, from a secondary source. Z.ai's own GLM-5.2 guide says 128K. Z.ai changed nothing; we did. It is logged as a board correction in the data file, so the history does not show a move that never happened.
The calendar
- 27 September 2026 — Perplexity's Sonar Chat Completions endpoint (migration guide). Three days.
- 15 October 2026 — the earliest Claude Haiku 4.5 can be retired, per Anthropic's deprecations page.
- 20 October 2026 — Kimi's built-in web search is expected to be deprecated (pricing page).
- 21 November 2026 — the earliest GPT-5.6 Sol's promotional rate can end.
- 1 January 2027 — Gemini 3.8 Flash doubles, per Google's pricing page.
What we would actually do
- Pin the version, not the family name. GPT-6 Sol is half GPT-5.6 Sol, and only the id change collects it.
- Move Opus 5 to 5.5 on a test, not a find-and-replace. Disabling thinking and forcing a tool now return errors, and the default effort drops to medium.
- Ask whether the job needs text at all. For routing, triage, guardrails and rubric scoring, price a Jev-class model per decision against the language-model call you make today.
Everything above is on the model board — sixty-five models this week and a new job card — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in.