Naderu Weekly · Episode 10
Same price, new rules
Two labs shipped a new model this week, and neither moved a price by a cent. Both changed what your code is allowed to send. Perplexity's Sonar now bills search per call, and Alibaba shipped a model in Jev's class.
Watch on YouTube — 9:44. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
GPT-6.1 Sol: half the cache read, and no effort none
OpenAI's changelog added GPT-6.1 Sol on 29 September at $2.00 in / $0.10 cached / $10.00 out per 1M tokens — GPT-6 Sol's input and output exactly, on the pricing page. Same 1.05M-token window and 128K max output. The cached token is the change: $0.10 is 0.05x input, and OpenAI's caching guide singles out 6.1 Sol as the exception to the usual 0.1x. Above 272K input the whole request bills at $4.00 / $0.20 / $15.00.
The catch is on the model
page: reasoning effort starts at low, and none and minimal
are not supported. GPT-6
Sol accepts none. So the arithmetic cuts both ways. An agent loop that re-sends
100K cached tokens a step pays 2¢ for them on 6 Sol and 1¢ on 6.1 Sol. A workload that ran 6 Sol
with no thinking — classification, formatting, short rewrites — now bills reasoning tokens at the
$10 output rate.
OpenAI's models page now lists Astra, 6.1 Sol and Luna as the places to start. GPT-6 Sol has no shutdown date, so on the board it moves to previous, not off.
Claude Sonnet 5.5: Sonnet 5's price, to the cent
Anthropic's release notes launched Claude Sonnet 5.5 on 28 September at $2 in / $0.20 cache read / $10 out — the same as Sonnet 5 on the pricing page. 1M window, 128K output, 300K on the Batch API with a beta header. One cost change in your favour: per the caching docs it caches from 512 tokens, where Sonnet 5 needed 1,024. Sonnet 5 becomes Legacy, with retirement no sooner than 30 June 2027 on the deprecations page.
The model page lists five breaking changes for code written against Sonnet 5:
- To turn thinking off you send
between_tools, notdisabled— and it only drops the thinking before the first tool call. - Forcing
tool_choicetoanyor a named tool returns a 400. - Thinking blocks are tied to the model and the conversation — and to the account that produced them.
- The earlier
computer_20251124tool is refused on the Claude API and Google Cloud. - The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
And a sixth that fails nothing: text between tool calls now arrives inside thinking blocks, so an app that streams it to users goes quiet between calls until it sets a display option.
On the board that moved four job cards: Sonnet 5.5 onto content writing, agentic coding and RAG; GPT-6.1 Sol onto agentic coding and code review. Computer use keeps Sonnet 5 and Opus 5 on purpose — both 5.5 models refuse the older computer-use tool, so a working loop has to move toolsets before it moves model.
Sonar moved to the Agent API, and search moved with it
Perplexity ended support for Sonar Chat Completions on 27 September
(migration
guide). Synchronous and streaming calls keep working while they are reformulated as Agent API
requests, rolling out by model; asynchronous calls are no longer supported. On the
Agent API,
perplexity/sonar is $0.25 in / $0.0625 cache read / $2.50 out, and
search is a tool fee of $0.0025 per call. The
pricing
page still lists Chat Completions Sonar at $1 / $1 plus $5–$12 per 1,000 requests.
| One answer: 2,000 in, 500 out | Chat Completions | Agent API |
|---|---|---|
| One search | ~0.75¢ | ~0.43¢ |
| Three searches | ~0.75¢ | ~0.93¢ |
Tokens and search fees only, at the lowest old request fee. Cheaper for one search, dearer for three, and output costs 2.5x what it did. Sonar Pro and Sonar Reasoning Pro have no Agent API model — the guide maps them to presets that run other labs' models — so the board now shows both as deprecated. One thing we could not find: which rate a reformulated Chat Completions call bills at. If you have not migrated, check this month's invoice.
Alibaba ships a Jev-class model
Last week we covered Jev, a
model that answers in types rather than text. On 24 September Alibaba released
decision-model-preview
in Model Studio: classification, yes/no and scoring in parallel over up to 65,536 tokens of text,
returning probabilities and confidence scores, billed on input only. The
pricing
page says "Limited-time free", with no end date and no list rate. So it is a model to evaluate,
not one to budget, and it sits in the board's ledger rather than on the card until there is a
number. It is the first entry in this class from a hyperscale cloud.
Ours, and said so first: naderu-laya-150m
On 28 September Naderu released naderu-laya-150m on Hugging Face under Apache-2.0. It is in the same class — a probability per option, not generated text — but narrower than Jev. Jev takes your question with each request; Laya answers one fixed question: which of five buckets a message belongs to (billing, technical support, sales, cancellation, or a question about the model itself). 149M parameters, fine-tuned from ModernBERT-base, and it runs on CPU, through ONNX Runtime or Core ML.
On templates it never saw in training it got 93.6% of 140 held-out messages right. That is 32 test templates, so read it as ±5 points, and every miss sits near the boundary between two buckets.
The fence: it is not on the typed-decisions card. It has no hosted price, and that is the rule the board applies to about twenty open-weight replications in this class — ours gets no exception. It is in the ledger with the others, and in a disclosed band on the board page, outside the comparison.
The rest of the board
Every other row was re-read at source and re-dated; no other lab moved a list price. Three notes from Anthropic's pages: from 24 September, refusals that arrive before any output are billed again in the bio, frontier-LLM and reasoning-extraction categories (release notes); Claude Sonnet 4.5 is deprecated and retires on 30 November; and, on Mistral's side, the models page now lists Z.ai's GLM 5.3 as a third-party model it serves.
This week's correction is ours
We carried Gemini 3.8 Flash and 3.7 Flash at 1M tokens in and 64K out — the rounded figures from Google's latest-model page. The models' own pages, such as Gemini 3.8 Flash, state 1,048,576 in and 65,536 out. Google changed nothing; we did. It is logged as a board correction in the data file, so the history does not show a move that never happened.
The calendar
- 15 October 2026 — the earliest retirement date on Claude Haiku 4.5's row. But Anthropic's deprecations page promises at least 60 days' notice, and none has gone out, so it cannot land then.
- 20 October 2026 — Kimi's built-in web search is expected to be deprecated (pricing page).
- 31 October 2026 — Mistral's hosted GLM 5.2 retires.
- 21 November 2026 — the earliest GPT-5.6 Sol's promotional rate can end.
- 30 November 2026 — Claude Sonnet 4.5 retires.
What we would actually do
- Read the migration guide even when the price didn't move. Sonnet 5.5 refuses a forced tool choice, and thinking has no off switch.
- Check your reasoning-effort floor. If you send effort
noneto GPT-6 Sol, 6.1 Sol will refuse it — stay where you are, or budget the thinking. - Price search per call, not per answer. On Sonar, one search an answer got cheaper. Three did not.
Everything above is on the model board — sixty-seven models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in.