Access Anthropic Claude Haiku 5.5—also searched as Haiku 5.5— through EvoLink's unified chat API. Test tool use, structured output, prompt caching before integrating.
Anthropic·Text Generation·Available
From $0.096 / 1M input tokens$0.100 official price-4%
Adaptive thinkingEffort controlTool useStructured outputPrompt cachingChat + Messages API
Production routeThis rate reflects platform-side availability — only confirmed server errors (HTTP 500 / empty response) count as failures. User-side issues (content moderation, invalid params, cancellation) plus rate limits, timeouts and auth errors are excluded. Before real traffic arrives, empty buckets may display as available.Live
Live
Context
1M context · 128K max output
Best For
High-volume classification and extraction, routing, sub-agent tasks
Anthropic’s newest Haiku model for high-volume classification, extraction, routing, and sub-agent work. The prompt length selects one of two rate cards, and five billing dimensions — input, cache write, cache read, output, and the web search tool — are metered separately.
Claude Haiku 5.5
Anthropic’s newest Haiku model
Selected
Model ID
claude-haiku-5-5
Best for
Ticket and intent classification, field extraction into JSON, request routing, and sub-agent steps under a larger model — the high-volume, latency-sensitive work where a Haiku-rate route should finish the task without escalating to Sonnet.
≤ 100K
Input
$0.096 / 1M-4%
6.5 cr / 1M$0.100official price
Cache write
$0.120 / 1M-5%
8.1 cr / 1M$0.125official price
Cache read
$0.011 / 1M
0.7 cr / 1M
Output
$0.475 / 1M-5%
32.3 cr / 1M$0.500official price
> 100K
Input
$0.478 / 1M-4%
32.5 cr / 1M$0.500official price
Cache write
$0.596 / 1M-5%
40.5 cr / 1M$0.625official price
Cache read
$0.052 / 1M
3.5 cr / 1M
Output
$2.375 / 1M-5%
161.5 cr / 1M$2.500official price
All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing. Above 100K prompt tokens the long-context tier applies (×5) to every token role.
Claude Haiku 5.5 pricing
Estimate what one Claude Haiku 5.5 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.
Request calculator
Enter the token mix for one request and the number of successful tool calls.
Estimated request cost
Claude Haiku 5.5
Short-context rate
USD$0.0003
Credits0.0164
Official estimate $0.0003 · save $0.0001 (4%)
Input tokens0.0065 cr
Cache write tokens0 cr
Cache read tokens0.0002 cr
Output tokens0.0097 cr
Minimum charge: 0.01 credits per request. Above 100K prompt tokens the long-context tier applies (×5) to every token role.
Web search depends on the selected route and is unavailable on some routes. Confirm support before sending web_search; the rates below apply only where it is supported.
Web search$0.010/ call·0.68 cr / call
Claude Haiku 5.5 API
Use Claude Haiku 5.5 API through EvoLink’s unified gateway. Run high-volume classification, extraction, routing, and sub-agent calls through EvoLink’s unified API, with a 1M-token context window and up to 128K output tokens. Compare $0.096 input / $0.475 output per million tokens for prompts up to 100K tokens, reuse one API key across models, and check thinking and max_tokens settings before your first request.
Claude Haiku 5.5 is served on EvoLink under the model ID claude-haiku-5-5 through Chat Completions · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window and up to 128K output tokens, plus tool use, structured output, prompt caching.
Claude Haiku 5.5 specs and capabilities
Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.
Context window
1M tokens
Max output
128K tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Thinking mode (on by default)
Tool use
Function calling with multi-step tool sequences
Prompt caching
Cache write + cache read rates
Long-context tier
×5 above 100K prompt tokens
Protocols
Chat Completions · Anthropic Messages
Model ID
claude-haiku-5-5
Where Claude Haiku 5.5 is the right default route
Haiku 5.5 is Anthropic’s fast, low-cost Claude for high-volume and latency-sensitive work. Route to it by the cost of a completed task, and tune effort before touching the prompt.
High-volume classification, routing, and extraction
Label tickets, detect intent, route requests, and pull fields out of documents into JSON. Keep each prompt compact and constrain the answer with structured output or an enum-valued tool, so the reply stays short and the request stays on the base rate card.
Sub-agents under a larger model
Run search, file-reading, and other narrow steps on Haiku 5.5 while an Opus or Sonnet model plans the task. Anthropic describes Haiku 5.5 as substantially better than Haiku 4.5 at following instructions and at running as a sub-agent. Measure completed steps, not single replies.
Low-latency chat and support assistants
Answer routine questions and call a small set of tools with a short system prompt. Start at low effort for chat and short tool tasks, and raise it where the assistant has to follow strict rules across a long conversation.
When another route is the better choice
Multi-file coding, long-horizon agents, and analyst-grade writing belong on Claude Sonnet 5.5, with Claude Opus 5.5 above it. Keep Claude Haiku 4.5 where a workload depends on sampling parameters, a thinking budget, or assistant prefill and has not been re-tested yet.
Two ways to use Claude Haiku 5.5: EvoLink API or Agent
Use the EvoLink API for product backends and batch jobs, or call Claude Haiku 5.5 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.
Option 1
Integrate with the EvoLink API
Best for: product backends, batch jobs, automated pipelines
Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.
1Create an EvoLink API key in the console
2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
3Send one representative request and read the usage field for input, cached, and output tokens
4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini
Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls Claude Haiku 5.5 through EvoLink, and returns the answer with token usage.
1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
2Describe the task, the inputs to include, and the expected output format
3Ask the Agent to call Claude Haiku 5.5 through EvoLink and show the request before sending
4Let the Agent report the answer, token usage, and any error body
Claude Haiku 5.5 API code example and error handling
This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.
curl -X POST https://direct.evolink.ai/v1/chat/completions \
-H "Authorization: Bearer $EVOLINK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"messages": [
{ "role": "system", "content": "Classify each support ticket as billing, bug, or how-to. Reply with one word." },
{ "role": "user", "content": "I was charged twice for my subscription." }
],
"max_tokens": 4096
}'
# Thinking is on by default and counts toward max_tokens, so leave room for
# it. This Chat Completions request returns choices[0].message and usage
# (prompt_tokens, completion_tokens, and cache read / write tokens when
# caching applies).
# For the Anthropic Messages format at /v1/messages, follow the linked
# Claude reference; its request fields and response format are different.
Invalid request or unsupported parameter
Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.
Authentication or balance issue
Check the Authorization bearer token and confirm the available balance in the console.
Context length exceeded
Prompt tokens exceed the Claude Haiku 5.5 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.
Rate limited (429)
Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.
Content or tool call rejected
Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.
Claude Haiku 5.5 API integration checks
Keep one EvoLink client and API key. Most requests carry over unchanged; the four steps below cover the settings that behave differently from Claude Haiku 4.5.
01
Switch the model ID
Point the client at EvoLink, keep the same API key, and select claude-haiku-5-5. Both the OpenAI-compatible Chat Completions endpoint and the Anthropic Messages endpoint accept it.
Low migration
02
Leave room for thinking
Thinking is on by default and counts toward max_tokens, so a cap sized for a one-word answer can be used up before any text is written. Raise max_tokens or lower effort, and read content by block type because a response can begin with a thinking block.
Behavior change
03
Remove Haiku 4.5-only settings
Anthropic rejects thinking budgets, non-default sampling values, and assistant prefill on Haiku 5.5. EvoLink converts a thinking budget to adaptive thinking with an effort level and removes temperature, top_p, and top_k; replace prefill with structured output or a system-prompt instruction.
Breaking change
04
Watch the 100K prompt line
Input, cache writes, and cache reads together decide the rate card. Trim history, retrieve less, or split the job to keep high-volume requests at or under 100K prompt tokens, and budget the higher card for the ones that need more context.
Cost control
Prompt caching for repeated context
Cache stable system prompts, tool schemas, and reference documents to reuse them across requests. Writes and reads are billed separately. Anthropic lowered the minimum cacheable prompt length compared with Haiku 4.5, so shorter prompts can now be cached. Cached tokens still count toward the prompt length that selects the rate card.
1M-token context with a rate card that changes at 100K
The route records a 1M-token context window and up to 128K output tokens. Prompts up to 100K tokens — input, cache writes, and cache reads together — use the base rates; above 100K the whole request, output included, is billed at 5×. Thinking counts toward max_tokens, so leave room for it.
Thinking is on by default — effort is the control
A request that does not set thinking runs adaptive thinking, at medium effort unless you choose a level. Thinking is billed as output and its text is omitted unless you ask for summaries. thinking: disabled is still accepted at effort high or below; Anthropic recommends lowering effort instead.
Haiku 4.5 parameters that no longer apply
Anthropic rejects thinking budgets, non-default sampling values, and assistant prefill on Haiku 5.5. EvoLink converts a budget to adaptive thinking with an effort level and removes temperature, top_p, and top_k for compatibility. Forced tool choice is accepted, but the response then starts with the tool call and no thinking.
Claude Haiku 5.5: two rate cards and five billing dimensions
Claude Haiku 5.5 does not bill as a single blended rate. The prompt length selects the rate card, each dimension is metered separately, and thinking is billed as output.
Input and output tokens, by prompt length
For prompts up to 100K tokens: $0.096 per 1M input tokens and $0.475 per 1M output tokens on EvoLink, compared with Anthropic’s $0.100 and $0.500. Above 100K prompt tokens the whole request moves to the higher rate card: $0.478 input and $2.375 output. Thinking is billed as output even when its text is not returned.
5-minute cache writes and cache reads
Storing a prefix for 5 minutes costs $0.120 per 1M tokens; reusing it costs $0.011 per 1M tokens. Cached tokens still count toward the prompt length that selects the rate card. EvoLink rates are rounded up to a whole billing unit, so the cache-read rate can sit slightly above the official list price; use the displayed prices when estimating savings.
Web search tool
On routes that support web_search, each search costs $0.010 and is not multiplied by the higher rate card. Search-result tokens are billed as ordinary input. Confirm route support before integrating; a configured price does not mean every route supports the tool.
Compare Claude routes after workload testing
EvoLink
First check whether Haiku 5.5 holds quality on your own tasks at the effort level you plan to run. Then compare price, context, caching, and workload fit to choose the production route.
Ticket and intent classification, field extraction into JSON, request routing, and sub-agent steps under a larger model — the high-volume, latency-sensitive work where a Haiku-rate route should finish the task without escalating to Sonnet.
The previous Haiku, billed at one rate for any prompt length; keep it where a workload still needs sampling parameters or a thinking budget.
Anthropic’s newest Sonnet, the step up for multi-file coding, tool-heavy agents, and work that Haiku 5.5 does not finish.
Claude Haiku 5.5
Input / output$0.096 / $0.475
Context1M
CachingRead + write
Best forTicket and intent classification, field extraction into JSON, request routing, and sub-agent steps under a larger model — the high-volume, latency-sensitive work where a Haiku-rate route should finish the task without escalating to Sonnet.
This page covers Claude Haiku 5.5, available on EvoLink under model ID "claude-haiku-5-5". Haiku names a model family; select the exact version in your request rather than sending "haiku" as the model ID. Claude Haiku 4.5 has a separate model page and rates.
What model ID should I use for the Claude Haiku 5.5 API?
Send model "claude-haiku-5-5". The same ID works for both the OpenAI-compatible Chat Completions endpoint and the Anthropic Messages endpoint.
How much does the Claude Haiku 5.5 API cost on EvoLink?
For prompts up to 100K tokens, EvoLink charges $0.096 per 1M input tokens and $0.475 per 1M output tokens, against Anthropic’s $0.100 and $0.500. Above 100K prompt tokens the rates are $0.478 and $2.375.
How does the 100K-token rate card work?
The prompt length is input tokens plus cache-write and cache-read tokens. At 100K or fewer the base rates apply. Above 100K the whole request is billed at 5× — input, output, cache writes, and cache reads — not only the tokens past the line. Web search calls are not multiplied.
Is Claude Haiku 5.5 cheaper than Claude Haiku 4.5?
Anthropic lists Haiku 5.5 at a tenth of the Haiku 4.5 rate for prompts up to 100K tokens and at half of it above that. Haiku 5.5 uses a newer tokenizer that counts roughly 30% more tokens for the same text, and thinking is on by default, so compare the cost of a completed task on your own workload rather than the per-token rate alone.
How is Claude Haiku 5.5 prompt caching billed?
On the base rate card, 5-minute cache writes cost $0.120 per 1M tokens and cache reads cost $0.011 on EvoLink; Anthropic lists $0.125 and $0.010. They are metered separately, and both are billed at 5× above 100K prompt tokens. Use the displayed rates, including rounding, to estimate your workload.
How is the Claude Haiku 5.5 web search tool billed?
On routes that support web_search, each search costs $0.010. Search-result tokens are billed as ordinary input. Confirm route support before integrating; a configured price does not mean every route supports the tool.
Can I turn thinking off on Claude Haiku 5.5?
Yes, at effort high or below: send thinking: disabled. At xhigh or max Anthropic rejects it, so lower the effort level or use adaptive thinking. Anthropic recommends lowering effort instead of turning thinking off, because at low effort the model can skip thinking on simple requests.
Which effort level should I use on Claude Haiku 5.5?
The default is medium. Anthropic suggests low for chat, short tool tasks, and simple high-volume requests, medium for most work including agentic coding, and high for knowledge work and strict instruction following. Compare xhigh and max against Claude Sonnet 5.5 before settling on them.
What changes when I migrate from Claude Haiku 4.5?
Change the model ID, replace thinking budgets with adaptive thinking and an effort level, remove temperature, top_p, and top_k, and replace assistant prefill. Thinking is on by default, so read content by block type and revisit small max_tokens values. The tokenizer counts more tokens for the same text, context grows to 1M tokens, max output grows to 128K, and on the Claude API computer use requires the computer toolset.
Why does a short classification request stop at max_tokens with no text?
Thinking counts toward max_tokens. A cap sized for a one-word answer can be used up by thinking, and the response then ends after a thinking block and before any text. Raise max_tokens to leave room for thinking, or lower the effort level.
What does a refusal from Claude Haiku 5.5 look like?
Safety classifiers can decline a request. The response is a normal success with stop_reason "refusal" and a category, so check stop_reason before reading content. Anthropic does not re-serve a declined Haiku 5.5 request on another model; handle the refusal in your client and try a different prompt.
Which protocols and SDKs work with Claude Haiku 5.5?
Use the OpenAI-compatible Chat Completions endpoint or the Anthropic Messages endpoint with the same EvoLink API key. Claude Code and other Anthropic-native clients work by pointing the base URL at EvoLink.
Do Claude Code subscription usage limits apply to the Claude Haiku 5.5 API on EvoLink?
Claude subscription quotas are separate from EvoLink API billing. Requests through EvoLink consume your EvoLink balance at the displayed rates and remain subject to your account and route limits. Configure an EvoLink API key and the documented endpoint in a compatible client.
What context window and max output does Claude Haiku 5.5 support?
The EvoLink route records a 1,000,000-token context window with up to 128K output tokens per response. Thinking counts toward max_tokens, and prompts above 100K tokens are billed on the higher rate card.
What should a production Claude Haiku 5.5 evaluation measure?
Track first-pass success, valid structured output, retries, model requests per task, output tokens at your chosen effort, cache hit rate, the share of requests above 100K prompt tokens, time to accepted result, and human correction — not the per-token rate alone.