Anthropic launches Claude Haiku 5.5, its cheapest and fastest small model yet
IA en un minuto newsroom · Editor: Jon Elgezabal

In 30 seconds
Haiku 5.5 is Anthropic's new small model: on average it comes in around 75% cheaper than Haiku 4.5, answers faster than any other model from the company at standard speed and lets users choose the effort level. Sonnet 5.5 is also getting cheaper, because reading from its cache now costs half as much. This week, Max and Team plans start receiving a monthly balance for the API, up to $500 on Team.
Anthropic has introduced Claude Haiku 5.5, the cheapest, fastest and most capable small model it has ever released.
It is designed for high-volume, cost-sensitive tasks such as summaries or classification. According to Anthropic, it costs around 75% less on average than Haiku 4.5 and is its fastest model to date at standard speed (its Opus models in Fast Mode run more quickly), so it works especially well for live customer support and for operating a web browser. It is also the first Haiku model with an adjustable effort setting to optimize for cost or intelligence.
For requests of up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, compared with $1 and $5 for Haiku 4.5; above that size, $0.50 and $2.50. Anthropic notes that its new tokenizer uses slightly more tokens per task, and says the 75% average already takes this into account. It is available from today on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure.
Anthropic is also halving the price of cache reads (reusing text that has already been processed) for Sonnet 5.5, from $0.20 to $0.10 per million tokens. Because these reads make up a large share of what models consume, Sonnet 5.5 now runs around 20% cheaper on most agentic work.
And this week it will start giving Claude Max and Team subscribers a monthly credit to build tools, apps and agents with its API: $100 or $200 a month on Max, depending on the tier (5x or 20x), and up to $500 on Team, pooled across users.
Why it matters · analysis and opinion
More than a new model, the announcement is a price cut on three fronts at once: the small model, the Sonnet cache and a balance for the API on paid plans. For anyone already automating repetitive work, such as classifying requests, summarizing documents or handling live customer support, it can make uses that used to be too expensive pay off. The effort setting adds a practical choice: not every task needs the same capability, and now it can also be tuned on the cheap model. What is still missing is checking the real quality outside Anthropic's own tests, and keeping in mind that the new tokenizer uses a little more per task. The sensible move is to try it on a sample of your own work before switching models.
Official source: Anthropic · Written with the help of AI: how we make the news


