Naderu Weekly · Episode 11

Stop paying for words you throw away

8 October 2026 Week of 5 October 10 min episode · 7 min read

Somewhere in your code there is a call that asks a language model a question, gets back a paragraph, and keeps one word. This week OpenAI and Perplexity both started selling the word on its own. Claude Haiku 5.5 arrived at a tenth of Haiku 4.5's price, and Sonnet 5.5's cache read halved to match GPT-6.1 Sol's.

Watch or listen

Watch on YouTube — 9:42. Ira and Nitin are synthetic voices; the research and the arithmetic are ours, and every figure below is linked to where we got it. If you would rather have it arrive than remember to check, subscribe to the weekly report — one email a week, one click to leave.

Two labs, one endpoint: /v1/decisions

Both shipped the same shape. You send the evidence — text, JSON or images — plus named questions, each with a type: a predicate (yes or no), a choice from options you supply, or a score on a rubric. You get back a probability per option and no generated text, and you are billed for input only.

Three weeks ago this was one start-up's idea: TypeSafe's Jev opened the category on 15 September. Alibaba's preview in the same class is still free with no list rate, so it is a model to evaluate, not one to budget.

Two things for whoever writes the client. Both take images only as inline base64 — a hosted image URL is refused. And on Perplexity an image over its tile limit does not fail fast: the request waits about a minute, then returns a 504. Resize before you send, and set the timeout to fit. Perplexity's own pages also disagree on price: the changelog entry announcing the API says $0.04, the pricing page and quickstart say $0.02. The board carries $0.02 and says why in the row.

What a decision costs

A million decisions, 1,000 input tokens eachInput cost
Perplexity Decider v1.1$20
meraGPT Decider 1$40
TypeSafe Jev$42
OpenAI Decisions (GPT-6 Luna)$100
Claude Haiku 4.5 — input alone, before it writes a word$1,000

A five-fold spread for the same shape of answer. Each vendor counts tokens its own way, so read these as orders of magnitude, not to the cent.

If your code turns a model's reply into a label, that call is a decision — and you are paying for the paragraph.

If you run the team

Look for one pattern: any call whose reply your code parses into a label. Which team gets this ticket. Is this message abusive. Did this agent step succeed. Price those per decision, not per token, and pilot one queue. The pilot is a week's work: take tickets your people have already routed by hand, send them through, and pick the threshold where a wrong route costs less than a person reading it — everything below that threshold still goes to a person.

What decides it is calibration on your own labelled data: when the model says 90%, it needs to be right about nine times in ten on your traffic. Every speed figure on these pages, OpenAI's "10x faster than the Responses API" included, is the vendor's own measurement. And anything that needs prose, facts it would have to look up, or arithmetic still belongs to a language model.

Claude Haiku 5.5: a tenth of the price, four things to check

Anthropic's release notes launched Claude Haiku 5.5 on 7 October at $0.10 in / $0.01 cache read / $0.50 out per 1M tokens on the model page, against $1 / $0.10 / $5 for Haiku 4.5. A 1M window, 128K output, and per the caching docs it caches from 512 tokens, where Haiku 4.5 needed 4,096. Haiku 4.5 is now Legacy, and on the board it moves to previous.

Run the arithmetic with the extra tokens counted: a 1,000-token prompt with a 200-token reply cost a fifth of a cent on Haiku 4.5 and about three hundredths of a cent on 5.5 — some 87% less, before any thinking. Anthropic's own average is around 75%. Both are big; neither is your number until you measure it.

Eight days apart, the same cache price

On 7 October Anthropic cut Claude Sonnet 5.5's cache read from $0.20 to $0.10 per 1M tokens — 0.05x input. GPT-6.1 Sol launched on 29 September at exactly $0.10 on OpenAI's pricing page. Two models at $2 in and $10 out, and now the same price for a token they have already seen. An agent loop of 200 steps re-sending 100K cached tokens paid $4 in cache reads on Sonnet 5.5 last week and $2 this week.

One thing on the record: Anthropic's pricing page disagrees with itself. Its model table still prints $0.20 for a Sonnet 5.5 cache hit; the multiplier section on the same page, and the dated release note, say $0.10. The board follows the dated change and says so in the row — so if an invoice shows $0.20, that is where to look.

The rest of the board

Seventy rows now, every one re-read at source this week. One list price went up: meraGPT's Decider 1, from $0.03 to $0.04 per 1M input — in the week Perplexity's Decider came in at half that. OpenAI cut its usage tiers from five to three. And Perplexity's low, medium and high Agent API presets now use fast search, which drops their web search from $2.50 to $1.00 per 1,000 calls.

Ours, and said so first: naderu-laya-150m

One model in this class is ours. naderu-laya-150m, released on Hugging Face on 28 September under Apache-2.0, answers one fixed question — which of five buckets a message belongs to — and runs on your own CPU. It has no hosted price, so it stays off the board's card. If your data cannot leave the building, a model you run yourself is the other way to buy a decision.

The calendar

Three moves for Monday

Everything above is on the model board — seventy models this week — with the page it came from and the day we read it, and the file behind it is open: no key, no sign-up. When we get something wrong we correct it in public and say which pass it was wrong in.