Anthropic Launches Claude Haiku 5.5 with Lower Costs and Faster Response Times

Anthropic Launches Claude Haiku 5.5 with Lower Costs and Faster Response Times

Anthropic has released Claude Haiku 5.5, introducing its fastest and most affordable small model to date. The update cuts operating costs by roughly 75% compared to the previous version while adding an adjustable computation dial for balancing speed and accuracy. The model is now live across Anthropic platforms alongside major cloud providers.

Anthropic Launches Claude Haiku 5.5 with Lower Costs and Faster Response Times

Anthropic designed the architecture for rapid, high volume operations. Common use cases include generating summaries, classifying text, querying databases, and serving as a subagent for coding alongside larger models like Sonnet 5.5 and Opus 5.5. The system introduces 5 separate effort settings labeled Low, Med, High, Xhigh, and Max. Users can select how much compute the model spends on a response, giving developers direct control over runtime costs versus output quality.

Performance gains appear sharpest in automated system control. In the OSWorld 2.1 benchmark, the model scored 72.4%, surpassing the 15.7% mark of Haiku 4.5 and the 48.9% mark of GPT 6 Luna. On the GDPval AA v2.1 test, it reached 1620 points compared to 735 for its predecessor. Aaron Vinh, a staff software engineer at Asana, confirmed the latency improvements during internal testing.

We are very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It is a noticeably snappier experience.

For prompts up to 100,000 tokens, Anthropic charges $0.10 per million input tokens and $0.50 per million output tokens, with cache reads priced at $0.01. Prompts exceeding 100,000 tokens cost $0.50 for input and $2.50 for output per million tokens. Alongside this release, the company cut the price of Sonnet 5.5 cache reads by 50% to $0.10, which lowers operational costs by roughly 20% on agent workflows. Max and Team tier subscribers also receive monthly API credits ranging from $100 to $500 to test new integrations.

About the author

mgtid
mgtid
Owner of Technetbook | 10+ Years of Expertise in Technology | Seasoned Writer, Designer, and Programmer | Specialist in In-Depth Tech Reviews and Industry Insights | Passionate about Driving Innovation and Educating the Tech Community Technetbook

Join the conversation

Newsletter Subscription