Anthropic Just Cut Its Cheapest Model's Price 90% — and Matched OpenAI's to the Cent


The cheapest frontier-adjacent model just got 90% cheaper

On Tuesday, Anthropic closed out the Claude 5.5 family with Haiku 5.5 — the small model built for high-volume, latency-sensitive work: classification, routing, extraction, summaries, compaction, browser automation, and subagents. The headline is the price. For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. The previous Haiku 4.5 charged $1 and $5. That is a 90% cut, and it is not a coincidence that it lands at exactly the list price of OpenAI’s GPT-6 Luna.

This is the small-model pricing war, fought in public, to the second decimal place. Opus 5.5 shipped September 22, Sonnet 5.5 on September 28, and GPT-6 Luna launched the same afternoon as Opus 5.5. Now Haiku 5.5 arrives with identical short-context pricing to Luna — $0.10 input, $0.50 output — and Anthropic’s own comparison table shows Haiku 5.5 ahead of Luna on every benchmark row where both have a score. Independent evaluators at Artificial Analysis put Haiku 5.5 at 43 on their Intelligence Index at max effort, the top of the small-model group.

The message to builders is deliberate: the small model is no longer the dumb model. It is the cheap model that can do real work — and the two labs are now competing on who can make it cheapest.

What actually shipped

Haiku 5.5 is a substantial upgrade, not just a price cut. The context window jumps from 200K to 1 million tokens. Maximum output goes from 64K to 128K tokens (300K on the Batch API with a beta header). Knowledge cutoff is June 2026. It takes text and images, outputs text, and the API model ID is claude-haiku-5-5 — generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.

The biggest functional change is the effort parameter. Haiku 5.5 is the first Haiku-class model with adjustable effort, from low to max, defaulting to medium, with adaptive thinking on by default. For the first time, the cheap model gets a dial: turn effort down for classification and routing where speed and cost matter, turn it up for harder subagent tasks where you need the extra reasoning. The frontier labs spent 2026 pushing reasoning controls up the stack; now the control has reached the bottom of the lineup.

The benchmark jumps are genuinely startling for a “small” model. At maximum effort, Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 — Haiku 4.5 scored 0.0%. OSWorld 2.1’s offline subset: 72.4% versus 15.7%. SWE-bench Multilingual: 83.7% versus 67.4%. GDPval-AA v2.1: 1620 Elo versus 735. These are not incremental deltas; the small model went from “cannot do terminal tasks” to “does a third of them.” For agentic workloads — the subagent fleets that sit behind Opus and Sonnet doing the bulk of the actual calls — that changes what you can route to the cheap tier.

The pricing cliff and the tokenizer asterisk

Now the fine print, because this is where the bill gets decided. Haiku 5.5’s pricing splits at 100K prompt tokens: above that, rates jump to $0.50 input and $2.50 output per million — a 5x markup at the threshold, applied to the whole request. Anthropic says roughly 90% of Haiku 4.5 requests stayed under 100K tokens, so most traffic lands in the cheap tier. But as the HN crowd flagged immediately, a 5x cliff is awkward for cost planning: one long-context agent loop crossing the line re-prices the entire run.

Then there is the tokenizer. Haiku 5.5 uses the tokenizer introduced with Claude Opus 4.7, which counts the same text as roughly 30% more tokens than Haiku 4.5. Independent token counts on October 8 put Haiku 5.5 at about 1.5 to 1.7 times as many tokens as GPT-6 Luna for identical text — meaning the “identical list price” is not an identical bill. Anthropic’s headline figure — Haiku 5.5 runs about 75% cheaper on average than Haiku 4.5 — already accounts for the tokenizer difference, but the Luna comparison needs the same adjustment: at equal list prices with ~1.6x the tokens, Haiku costs more per unit of actual work. At maximum effort, one review found Haiku 5.5 writes roughly three times as many output tokens per task as Luna.

Luna’s structure is gentler on long prompts: its higher tier starts at 272K input tokens, not 100K, with a smaller markup (2x input, 1.5x output) rather than 5x. For a 150K-token prompt, Luna is cheaper on list price — before tokenization differences. The pricing war is real, but it is being fought with two different scoreboards, and the scoreboard that matters is your bill, not the press release.

One more migration gotcha for existing Haiku users: non-default temperature, top_p, or top_k values now return a 400 error. The migration guide covers this and the tokenizer change — read it before swapping model IDs in production.

Why this is happening now

The timing is not mysterious. This week alone, the industry’s cost pressure has been showing in public: Meta and Microsoft are quietly weaning employees off Claude internally, and the conversation across the industry is about who pays for all this inference. Small models are where the volume is — subagents, compaction, classification, the millions of background calls behind every agentic product — and volume is where a 90% price cut moves the most money.

Anthropic paired the launch with a 50% cut to Sonnet 5.5’s cache-read price, from $0.20 to $0.10 per million tokens, which it says makes Sonnet 5.5 around 20% cheaper on most agentic workloads. The pattern is clear: the labs are competing on the unit economics of agents, not the headline intelligence of flagships. The frontier model gets the keynote; the small model gets the margin.

And the early evidence says the economics are real. Plotly’s benchmark team reported their data analytics score jumped two letter grades at 9x lower cost on Haiku 5.5. When the cheap model gets smarter faster than the expensive one gets cheaper, the default choice for high-volume work flips.

What builders should do now

1. Re-run your routing logic. If you built a two-tier setup — expensive model for hard tasks, cheap model for easy ones — the boundary just moved. Haiku 5.5 at max effort does terminal and OSWorld tasks its predecessor scored zero on. Test your current “hard” tasks on Haiku 5.5 at high effort before assuming they need Sonnet or Opus.

2. Budget against bills, not list prices. Model your per-token costs with the new tokenizer: budget ~30% more tokens than your Haiku 4.5 counts, and compare against Luna on your actual prompts, not the identical $0.10/$0.50 sticker. Include the 100K cliff in your cost model — cap agent-loop context or split long prompts deliberately.

3. Use the effort dial as a cost control, not just a quality knob. Effort low for classification and routing, medium for summaries and extraction, max for subagent work that used to need a bigger model. This is the first time the Haiku line lets you tune cost per request type — build it into your routing layer instead of one-sizing every call.

4. Watch your temperature settings before migrating. The 400 error on non-default temperature/top_p/top_k will break pipelines that set these habitually. Audit your Haiku 4.5 call sites for sampling parameters and the batch output limit before switching the model ID.

5. Revisit cache-heavy architectures. The Sonnet 5.5 cache-read cut to $0.10 and Haiku 5.5’s $0.01 cache reads change the math on prompt caching for agentic loops. If you dismissed caching as marginal savings, re-run the numbers — at these prices, caching is a first-class cost control.

The war is about the bill, not the benchmark

Strip away the press releases and this is a fight over who owns the economics of the agentic era. The flagship models are the billboards; the small models are the business. Anthropic just matched OpenAI’s price to the cent, shipped a small model that does real agent work, and hid a 5x pricing cliff and a 30%-heavier tokenizer in the footnotes. OpenAI answered with a gentler long-context structure and a leaner tokenizer.

Both labs are telling you the same thing: the future is millions of cheap, smart, background calls — and they each want to be the one selling them to you. Your job is to read past the sticker price, count your actual tokens, and route accordingly. The labs will keep cutting; make sure your architecture is the thing that captures the savings.