Anthropic has launched Claude Haiku 5.5, its newest small model, with a 1M token context window and input pricing of $0.10 per million tokens.
In this article
The pricing breakdown
The cost structure splits at 100K prompt tokens. Requests under that limit cost $0.10 for input and $0.50 for output per million tokens. Cache reads are $0.01 per million, while 5-minute cache writes cost $0.125. Beyond the 100K threshold, rates jump to $0.50 for input and $2.50 for output. This represents a 90% reduction compared to Claude Haiku 4.5 for prompts up to 100K tokens, after adjusting for the new tokenizer.
Haiku 4.5 charged $1.00 for input and $5.00 for output. Anthropic notes that roughly 90% of requests for the previous model stayed under 100K tokens. They estimate the average cost for Haiku 5.5 is 75% lower than its predecessor. Batch processing adds another 50% discount.
GPT-6 Luna lists identical rates for short contexts. Its higher tier begins only above 272K input tokens, at $0.20 and $0.75 per million. For a 150K-token prompt, Luna is cheaper on the list price.
Technical specifications
Haiku 5.5 supports text and images, outputs text, and has a knowledge cutoff of June 2026. It allows up to 128K output tokens by default, though beta batch jobs support up to 300K output tokens.
The model features an adjustable effort setting with adaptive thinking enabled by default. The effort parameter defaults to medium. Developers must use default values for temperature, top_p, and top_k, or the API returns a 400 error. The new tokenizer counts roughly 30% more tokens for the same text compared to Haiku 4.5.
Deployment is available via hosted API on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS.
Performance data
Anthropic reports the following benchmark figures:
- OSWorld 2.1 (offline subset): 72.4%, compared to 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5.
- Terminal-Bench 4.0: 39.2%, versus 16.4% for Luna and 0.0% for Haiku 4.5.
- FrontierCode 1.1 (Main): 46.4%, versus 42.4% for Luna.
- Humanity’s Last Exam: 45.9% without tools and 57.4% with tools.
- GDPval-AA v2.1: 1620, versus 1437 for Luna and 735 for Haiku 4.5.
Sonnet 5.5 remains the leader across all metrics, including 70.6% on Terminal-Bench 4.0. Anthropic recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding tasks.
Practical applications
Three specific workloads fit this model best:
- Subagent work running under Opus 5.5 or Sonnet 5.5. For example, a Haiku 5.5 subagent can extract a 10-K revenue line while a larger model builds the presentation deck.
- High-volume document Q&A and summarisation. AlphaSense tested it on a feature handling approximately 8M calls a week.
- Speed-sensitive work such as live customer support and browser use.
What it means
Developers can now run subagents on a budget that is significantly lower than previous iterations, enabling more complex multi-step workflows without breaking the bank. The pricing tiering encourages batching, which reduces costs for high-volume users. The strict limits on temperature and sampling parameters mean developers must configure the model carefully, but the performance gains on coding and terminal benchmarks suggest it is a capable tool for specific, constrained tasks.



