Access Anthropic Claude Sonnet 5—also searched as Sonnet 5 / Claude 5 Sonnet— through EvoLink's unified chat API. Test long-context coding, tool use, prompt caching before integrating.
Anthropic·Text Generation·Available
From $1.800 / 1M input tokens$2.000 official price-10%
Adaptive thinkingAgentic codingTool usePrompt cachingChat + Messages API
Production routeThis rate reflects platform-side availability — only confirmed server errors (HTTP 500 / empty response) count as failures. User-side issues (content moderation, invalid params, cancellation) plus rate limits, timeouts and auth errors are excluded. Before real traffic arrives, empty buckets may display as available.Live
Anthropic's previous Sonnet for coding and agentic work, and a drop-in replacement for Sonnet 4.6. Input, cache write, cache read, output, and the web search tool are metered as separate billing dimensions.
Claude Sonnet 5
Anthropic Sonnet-tier model for coding and agents
Selected
Model ID
claude-sonnet-5
Best for
Everyday production coding, terminal and browser agents, long-horizon tool use, and large-context analysis where balanced cost matters more than an Opus-tier rate.
Input
$1.800 / 1M-10%
122.4 cr / 1M$2.000official price
Cache write
$2.250 / 1M-10%
153 cr / 1M$2.500official price
Cache read
$0.181 / 1M-10%
12.3 cr / 1M$0.200official price
Output
$9.000 / 1M-10%
612 cr / 1M$10.000official price
All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.
Claude Sonnet 5 pricing
Estimate what one Claude Sonnet 5 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.
Request calculator
Enter the token mix for one request and the number of successful tool calls.
Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
Web search$0.010/ call·0.68 cr / call
Claude Sonnet 5 API — Anthropic's previous Sonnet for coding and agents
Claude Sonnet 5 (Claude 5 Sonnet) delivers strong coding and agentic performance across terminals and browsers, with a 1M context window, up to 128K output tokens, and adaptive thinking on by default.
Claude Sonnet 5 is served on EvoLink under the model ID claude-sonnet-5 through Chat Completions · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window, plus long-context coding, tool use, prompt caching.
Claude Sonnet 5 specs and capabilities
Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.
Context window
1M tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Thinking mode (on by default)
Tool use
Function calling with multi-step tool sequences
Prompt caching
Cache write + cache read rates
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Anthropic Messages
Model ID
claude-sonnet-5
What can you build with the Claude Sonnet 5 API?
Agentic Coding Assistant
Sonnet 5 posts its largest gains over Sonnet 4.6 in coding — planning multi-step changes, running terminals, and iterating on results. With up to 128K output and a 1M context window, it handles large codebases and generates comprehensive diffs, test suites, and implementation plans in a single request.
Autonomous Agents
Build agents that plan a sequence of steps, call tools, read the result, and keep going without a human nudge at every turn. Sonnet 5 is built for terminal and browser use and long-horizon runs, delivering reliable tool use across multi-step workflows.
Large-Context Analysis
Read very large documents, codebases, or long agent traces in a single request with the 1M context window, and let adaptive thinking apply deeper reasoning only when a task needs it — keeping cost predictable for research, planning, and technical strategy.
Two ways to use Claude Sonnet 5: EvoLink API or Agent
Use the EvoLink API for product backends and batch jobs, or call Claude Sonnet 5 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.
Option 1
Integrate with the EvoLink API
Best for: product backends, batch jobs, automated pipelines
Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.
1Create an EvoLink API key in the console
2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
3Send one representative request and read the usage field for input, cached, and output tokens
4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini
Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls Claude Sonnet 5 through EvoLink, and returns the answer with token usage.
1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
2Describe the task, the inputs to include, and the expected output format
3Ask the Agent to call Claude Sonnet 5 through EvoLink and show the request before sending
4Let the Agent report the answer, token usage, and any error body
Claude Sonnet 5 API code example and error handling
This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.
curl -X POST https://api.evolink.ai/v1/chat/completions \
-H "Authorization: Bearer $EVOLINK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"messages": [
{ "role": "system", "content": "You are a senior engineer. Produce clean, well-tested code." },
{ "role": "user", "content": "Review this diff and list the blocking issues with file and line:\n<diff>" }
],
"max_tokens": 8192
}'
# Anthropic Messages API is also available at /v1/messages with the same model ID.
# The response includes choices[0].message and a usage object
# (prompt_tokens, completion_tokens, and cache read / write tokens when
# caching applies).
Invalid request or unsupported parameter
Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.
Authentication or balance issue
Check the Authorization bearer token and confirm the available balance in the console.
Context length exceeded
Prompt tokens exceed the Claude Sonnet 5 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.
Rate limited (429)
Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.
Content or tool call rejected
Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.
How to integrate the Claude Sonnet 5 API
Connect through EvoLink, choose your model ID, and start building in minutes.
1
Step 1 — Create your EvoLink API key
Sign up for EvoLink to get a single API key that routes to Anthropic, Bedrock, or Google Cloud.
2
Step 2 — Select the model ID
Use `claude-sonnet-5` to access the Sonnet 5 model through EvoLink's unified API. It is a drop-in replacement for `claude-sonnet-4-6`.
3
Step 3 — Optimize quality and cost
Adaptive thinking is on by default and can be tuned with the effort parameter; use prompt caching for stable system prompts and repeated long context to lower repeat costs.
Prompt caching billed as two separate dimensions
Cache write and cache read are metered separately, and cache hits are billed at 0.1x the base input rate. Stable system prompts, repository instructions, and tool schemas pay the write once and then read cheaply across a long agent run.
1M-token context with a 128K max output
1M is both the default and the maximum context, with no smaller variant. Keep connected code, specifications, and agent state in one working context, and treat the limit as capacity rather than a target: retrieve relevant evidence, cache the stable prefix, and set task-appropriate output budgets.
Adaptive thinking is on by default and replaces manual budgets
Sonnet 5 scales reasoning to task difficulty automatically. Setting `budget_tokens` for manual extended thinking returns a 400 error — use the effort parameter to guide depth, or pass `thinking: {type: "disabled"}` to turn it off. Non-default `temperature`, `top_p`, and `top_k` also return a 400 error, so remove them when migrating.
Recount tokens before shifting production traffic
Sonnet 5 uses a new tokenizer that produces roughly 30% more tokens for the same text. Per-token pricing is unchanged, but billed cost and context capacity for an equivalent request can differ — recount your prompts, revisit `max_tokens` limits sized close to your expected output, and compare success rate and total cost on production samples.
Claude Sonnet 5 API capabilities
Key specs and model features for production use
1M Context Window
Read very large documents or codebases in a single request. 1M is both the default and the maximum context.
128K Max Output
Generate long-form answers, plans, and code without early truncation.
Adaptive Thinking
Reasoning scales to task difficulty automatically, on by default — tune it with the effort parameter instead of manual thinking budgets.
Agentic Coding & Tool Use
Built for terminal and browser agents, reliable function calling, and long-horizon multi-step runs.
Vision + Multilingual Input
Accept text and image inputs with strong multilingual understanding.
Prompt Caching Rates
Cache writes and reads are priced separately; cache hits are billed at 0.1x the base input price.
Compare Claude routes after workload testing
EvoLink
First verify on your own tasks whether Sonnet 5 already clears the quality bar. Then compare price, context, caching, and workload fit to choose the production route.
Everyday production coding, terminal and browser agents, long-horizon tool use, and large-context analysis where balanced cost matters more than an Opus-tier rate.
The previous Sonnet generation, kept for workloads already validated against it and for cost-sensitive traffic.
The Opus-tier flagship for the hardest reasoning, review, and agent tasks where quality outweighs cost.
Claude Sonnet 5
Input / output$1.8 / $9
Context1M
CachingRead + write
Best forEveryday production coding, terminal and browser agents, long-horizon tool use, and large-context analysis where balanced cost matters more than an Opus-tier rate.
What is the context window and max output for Claude Sonnet 5?
Claude Sonnet 5 supports a 1M token context window — both the default and the maximum, with no smaller variant — and up to 128K output tokens in a single request, making it suitable for large codebases, long documents, and comprehensive generation tasks.
Which model ID should I use, and how do I access it through EvoLink?
Use `claude-sonnet-5` through EvoLink's unified, OpenAI-compatible API to access the Sonnet 5 model with a single API key. It is a drop-in replacement for `claude-sonnet-4-6` — update the model ID and, if needed, review token budgets and thinking settings.
How does Claude Sonnet 5 handle extended thinking?
Sonnet 5 uses adaptive thinking, on by default, which scales reasoning to task difficulty. Manual extended thinking (setting `budget_tokens`) is no longer supported and returns a 400 error — use the effort parameter to guide reasoning depth instead. To turn thinking off, pass `thinking: {type: "disabled"}`.
What changes when migrating from Sonnet 4.6 to Sonnet 5?
Sonnet 5 is a drop-in replacement, but note three things: adaptive thinking is on by default, manual extended thinking returns a 400 error, and non-default sampling parameters (`temperature`, `top_p`, `top_k`) return a 400 error — remove them when migrating. Tool definitions and response shapes are unchanged.
Why do my token counts change on Claude Sonnet 5?
Sonnet 5 uses a new tokenizer that produces roughly 30% more tokens for the same text. Per-token pricing is unchanged, but the cost and context capacity of an equivalent request can differ — recount your prompts and revisit `max_tokens` limits sized close to your expected output.
How is Claude Sonnet 5 priced?
EvoLink pricing: $1.800 per 1M input tokens and $9.000 per 1M output tokens, against Anthropic's $2.000 and $10.000. Prompt caching is billed separately — $2.250 per 1M cache write tokens and $0.181 per 1M cache read tokens on EvoLink, versus $2.500 and $0.200 official.
How is the Claude Sonnet 5 web search tool billed?
Web search is charged per search at $0.010. Tokens the search results add to the request are billed normally.
How does Sonnet 5 compare to Opus 4.8?
Sonnet 5 delivers Opus-class coding and agentic performance and is the right default for most production workloads. Opus 4.8 is the flagship for the hardest reasoning and agent tasks where top-tier quality matters most.
Where is Claude Sonnet 5 available?
Claude Sonnet 5 is available via the Anthropic API, AWS Bedrock, Google Cloud (Vertex AI), and Microsoft Foundry (preview). EvoLink can route to the provider you choose.
Does Claude Sonnet 5 support vision and multimodal input?
Yes. Sonnet 5 supports text and image input with strong multilingual capabilities, so you can combine documents, screenshots, and visuals in one request.