Naderu Weekly · Episode 11
Stop paying for words you throw away
Somewhere in your code there is a call that asks a language model a question, gets back a paragraph, and keeps one word. This week OpenAI and Perplexity both started selling the word on its own. Claude Haiku 5.5 arrived at a tenth of Haiku 4.5's price, and Sonnet 5.5's cache read halved to match GPT-6.1 Sol's.
Watch on YouTube — 9:42. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.
Two labs, one endpoint: /v1/decisions
Both shipped the same shape. You send the evidence — text, JSON or images — plus named questions, each with a type: a predicate (yes or no), a choice from options you supply, or a score on a rubric. You get back a probability per option and no generated text, and you are billed for input only.
- OpenAI — the Decisions
API, in beta from 6 October per the changelog,
runs on
gpt-6-lunaonly, at $0.10 per 1M input tokens, with no output, cache-read or cache-write charge. Text and images; ZDR and HIPAA for eligible customers. - Perplexity — the Decisions
API runs
pplx-decider-v1.1-27bat $0.02 per 1M input tokens on the pricing page, output free, under 262,144 input tokens a request, at 10 requests a second on every plan.
Three weeks ago this was one start-up's idea: TypeSafe's Jev opened the category on 15 September. Alibaba's preview in the same class is still free with no list rate, so it is a model to evaluate, not one to budget.
Two things for whoever writes the client. Both take images only as inline base64 — a hosted image URL is refused. And on Perplexity an image over its tile limit does not fail fast: the request waits about a minute, then returns a 504. Resize before you send, and set the timeout to fit. Perplexity's own pages also disagree on price: the changelog entry announcing the API says $0.04, the pricing page and quickstart say $0.02. The board carries $0.02 and says why in the row.
What a decision costs
| A million decisions, 1,000 input tokens each | Input cost |
|---|---|
| Perplexity Decider v1.1 | $20 |
| meraGPT Decider 1 | $40 |
| TypeSafe Jev | $42 |
| OpenAI Decisions (GPT-6 Luna) | $100 |
| Claude Haiku 4.5 — input alone, before it writes a word | $1,000 |
A five-fold spread for the same shape of answer. Each vendor counts tokens its own way, so read these as orders of magnitude, not to the cent.
If you run the team
Look for one pattern: any call whose reply your code parses into a label. Which team gets this ticket. Is this message abusive. Did this agent step succeed. Price those per decision, not per token, and pilot one queue. The pilot is a week's work: take tickets your people have already routed by hand, send them through, and pick the threshold where a wrong route costs less than a person reading it — everything below that threshold still goes to a person.
What decides it is calibration on your own labelled data: when the model says 90%, it needs to be right about nine times in ten on your traffic. Every speed figure on these pages, OpenAI's "10x faster than the Responses API" included, is the vendor's own measurement. And anything that needs prose, facts it would have to look up, or arithmetic still belongs to a language model.
Claude Haiku 5.5: a tenth of the price, four things to check
Anthropic's release notes launched Claude Haiku 5.5 on 7 October at $0.10 in / $0.01 cache read / $0.50 out per 1M tokens on the model page, against $1 / $0.10 / $5 for Haiku 4.5. A 1M window, 128K output, and per the caching docs it caches from 512 tokens, where Haiku 4.5 needed 4,096. Haiku 4.5 is now Legacy, and on the board it moves to previous.
- The 100K line. That price holds for prompts up to 100,000 input tokens. Above it, the whole request bills $0.50 / $2.50 — five times the short rate. Anthropic's announcement says around 90% of requests to the previous Haiku fall below the line. Check where yours fall.
- Thinking is on by default, at medium effort, so a call that never thought on Haiku 4.5 can open with thinking blocks billed as output (what's new).
- The same text counts as about 30% more tokens.
- Three things now return 400: a non-default
temperature,top_portop_k; an assistant prefill; andthinking.budget_tokens. Priority Tier is not supported on Haiku 5.5 either (migration guide).
Run the arithmetic with the extra tokens counted: a 1,000-token prompt with a 200-token reply cost a fifth of a cent on Haiku 4.5 and about three hundredths of a cent on 5.5 — some 87% less, before any thinking. Anthropic's own average is around 75%. Both are big; neither is your number until you measure it.
Eight days apart, the same cache price
On 7 October Anthropic cut Claude Sonnet 5.5's cache read from $0.20 to $0.10 per 1M tokens — 0.05x input. GPT-6.1 Sol launched on 29 September at exactly $0.10 on OpenAI's pricing page. Two models at $2 in and $10 out, and now the same price for a token they have already seen. An agent loop of 200 steps re-sending 100K cached tokens paid $4 in cache reads on Sonnet 5.5 last week and $2 this week.
One thing on the record: Anthropic's pricing page disagrees with itself. Its model table still prints $0.20 for a Sonnet 5.5 cache hit; the multiplier section on the same page, and the dated release note, say $0.10. The board follows the dated change and says so in the row — so if an invoice shows $0.20, that is where to look.
The rest of the board
Seventy rows now, every one re-read at source this week. One list price went up: meraGPT's Decider 1, from $0.03 to $0.04 per 1M input — in the week Perplexity's Decider came in at half that. OpenAI cut its usage tiers from five to three. And Perplexity's low, medium and high Agent API presets now use fast search, which drops their web search from $2.50 to $1.00 per 1,000 calls.
Ours, and said so first: naderu-laya-150m
One model in this class is ours. naderu-laya-150m, released on Hugging Face on 28 September under Apache-2.0, answers one fixed question — which of five buckets a message belongs to — and runs on your own CPU. It has no hosted price, so it stays off the board's card. If your data cannot leave the building, a model you run yourself is the other way to buy a decision.
The calendar
- 15 October 2026 — the earliest retirement date on Claude Haiku 4.5's row. But Anthropic's deprecations page promises at least 60 days' notice, and none has gone out, so it cannot land then.
- 20 October 2026 — Kimi's built-in
$web_searchis expected to be deprecated (pricing page). - 24 October 2026 — Perplexity's Agent API retires older OpenAI model IDs (changelog).
- 31 October 2026 — Mistral's hosted GLM 5.2 retires (models page).
- 21 November 2026 — the earliest GPT-5.6 Sol's promotional rate can end.
- 30 November 2026 — Claude Sonnet 4.5 retires.
Three moves for Monday
- If you lead the team: audit for decisions. Find the calls that keep one word, pilot one queue on a decisions endpoint, and judge it on your own labelled data.
- If you own the code: swap Haiku with a checklist, not a find-and-replace. Drop the sampling parameters and the prefill, set the effort, read content blocks by type, and recount your tokens.
- Re-price your agent loops. A cached token on Sonnet 5.5 or GPT-6.1 Sol is now 5% of input — make the stable prefix bigger, and keep it byte for byte the same.
Everything above is on the model board — seventy models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in.