A small AI task can become a large business expense when it runs millions of times. Anthropic's latest release is aimed at that part of the market: the summaries, lookups, classifications and background steps that keep agents working.

On October 7, the company introduced Claude Haiku 5.5 and said it costs around 75% less to run on average than Haiku 4.5. It also reduced the price of reading cached material with Sonnet 5.5, targeting the repeated context that can make longer agent workflows expensive.

The competition with OpenAI is becoming a question of unit economics. Businesses need to know how much useful work they get for the total bill, including the work that has to be checked or done again.

Why 75% and 90% Are Both in the Announcement

The Haiku 5.5 launch announcement distinguishes the reduction in listed token prices from the average cost of completing work. For shorter prompts, the listed input and output prices are 90% below Haiku 4.5. Longer prompts have a smaller discount, and the newer model counts text differently.

The model documentation specifies the two pricing tiers:

Standard API usageInput per million tokensOutput per million tokens
Haiku 5.5, prompts up to 100,000 tokens$0.10$0.50
Haiku 5.5, prompts over 100,000 tokens$0.50$2.50
Haiku 4.5 reference pricing$1.00$5.00
Published rates checked October 8, 2026. Caching, batch processing and other service charges are separate.

The threshold refers to each prompt's size. It is not a monthly allowance that runs out after 100,000 tokens. Anthropic's documentation also says the same text counts as approximately 30% more tokens with the newer tokenizer than with Haiku 4.5. A token is a unit used to process text, not a fixed number of words shared by every model.

Anthropic announces Haiku 5.5 and its estimated average running-cost reduction.

What the Rates Mean for a Repeated Task

Consider an illustrative workload totaling 100 million input tokens and 10 million output tokens, with every individual prompt within the lower tier. At Haiku 5.5's standard rates, those tokens cost $15: $10 for input and $5 for output. The same token counts at the listed Haiku 4.5 rates cost $150.

That is a price-sheet calculation, not a prediction that a business's bill will fall by exactly $135. The same underlying job may use different token counts after migration. Longer reasoning, retries, tools, caching and review also affect the amount ultimately paid.

The example shows why small-model pricing matters. A routine operation that was too expensive to run across an entire document archive may become practical when its per-call cost falls. But the company still needs to verify that the result is useful.

OpenAI Has a Direct Price Competitor

OpenAI's GPT‑6 Luna documentation lists the same standard short-context rates: $0.10 per million input tokens and $0.50 per million output tokens. Luna's higher-price tier begins above 272,000 input tokens, while Haiku's begins above 100,000.

The matching base rates make a simple claim that one provider is universally cheaper misleading. Prompt length, output size, cache behavior, service tier and task success all change the comparison. A longer prompt can sit in the higher tier for one model and the lower tier for another.

This developer pricing contest is separate from OpenAI's new GPT‑6 ChatGPT interface, which Blockster covered today. A consumer subscription and a metered API serving an application are different purchasing decisions, even when they use related model families.

Sonnet's Cache Cut Targets a Different Expense

For an agent performing several steps, much of the context can repeat: instructions, tool descriptions and earlier information. Caching allows supported repeated material to be reused at a lower read price.

Sonnet 5.5's current documentation lists cache reads at $0.10 per million tokens, alongside regular input at $2 and output at $10. Anthropic's announcement says the cache-read price has been halved and estimates roughly 20% savings on most agentic work.

The size of the saving depends on how much of a workload actually qualifies for cache reads. A workflow generating mostly new output will have a different cost mix from one repeatedly processing a large, stable context. Cache creation also has its own price.

Anthropic's October 7 follow-up gives the new Sonnet cache-read rate and its estimated effect on longer workflows.

Customers Are Testing Where the Smaller Model Fits

The most useful early examples concern narrow assignments. In a customer statement supplied for Anthropic's launch, Rogo described using Haiku to retrieve a revenue figure from a filing while a larger model works on a presentation.

“The short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries.”

Alex Wang, Applied AI at Rogo, in Anthropic's launch announcement

Asana's Aaron Vinh reported faster task completion in the company's evaluation of agent workflows, including bug triage and project setup.

“It's a noticeably snappier experience.”

Aaron Vinh, Staff Software Engineer at Asana, in Anthropic's launch announcement

These are customer evaluations published by the vendor, not universal guarantees. They do, however, illustrate a practical model-selection strategy: use the smaller model where its work is reliable, and reserve more expensive reasoning for tasks that need it.

Lower Costs Make Budget Controls More Important

Cheaper calls can encourage applications to do more. A lower unit price may reduce the cost of one task while a growing number of tasks increases total spending.

That becomes especially relevant when agents purchase services themselves. Blockster's reporting on Sui and Alibaba Cloud's planned computing payments examines how software could buy resources within a user-set budget. The price of the model is one part of that budget; the services it invokes are another.

Our coverage of NVIDIA's work on agent guardrails addresses the accompanying permission problem. A cheaper agent still needs clear limits on what it can access and spend.

Haiku 5.5 gives businesses a reason to retest workflows that were previously too slow or costly. The useful benchmark is the cost of an accepted result, including failures and human review. That is where a lower price becomes a real operating saving.

Reporting by Lidia Yadlos

AIAnthropicOpenAIClaude