GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

Claude Sonnet 5 API

Access Anthropic Claude Sonnet 5—also searched as Sonnet 5 / Claude 5 Sonnet— through EvoLink's unified chat API. Test long-context coding, tool use, prompt caching before integrating.

AnthropicText GenerationAvailable
From $1.800 / 1M input tokens$2.000 official price-10%
API docs
Adaptive thinkingAgentic codingTool usePrompt cachingChat + Messages API
Production routeLive
Context
1M context
Best For
Agentic coding, autonomous agents, large-context analysis
Input
Text + images
Output
Text · JSON (structured output) · tool calls

Choose Claude Sonnet 5

Anthropic's previous Sonnet for coding and agentic work, and a drop-in replacement for Sonnet 4.6. Input, cache write, cache read, output, and the web search tool are metered as separate billing dimensions.

Claude Sonnet 5

Anthropic Sonnet-tier model for coding and agents

Selected
Model ID
claude-sonnet-5
Best for

Everyday production coding, terminal and browser agents, long-horizon tool use, and large-context analysis where balanced cost matters more than an Opus-tier rate.

Input
$1.800 / 1M-10%
122.4 cr / 1M$2.000official price
Cache write
$2.250 / 1M-10%
153 cr / 1M$2.500official price
Cache read
$0.181 / 1M-10%
12.3 cr / 1M$0.200official price
Output
$9.000 / 1M-10%
612 cr / 1M$10.000official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing.

Claude Sonnet 5 pricing

Estimate what one Claude Sonnet 5 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request and the number of successful tool calls.

Estimated request cost

Claude Sonnet 5
USD$0.0046
Credits0.3085

Official estimate $0.0051 · save $0.0006 (10%)

Input tokens0.1224 cr
Cache write tokens0 cr
Cache read tokens0.0025 cr
Output tokens0.1836 cr

Minimum charge: 0.01 credits per request.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 2106 requests

For quick testing

$50
About 11021 requests

For regular development

$100
About 22042 requests

For production evaluation

Server-side tool rates

Only successful server-side calls are billed per call; failed attempts have no tool fee, but tokens still apply.
  • Web search$0.010/ call0.68 cr / call

Claude Sonnet 5 API — Anthropic's previous Sonnet for coding and agents

Claude Sonnet 5 (Claude 5 Sonnet) delivers strong coding and agentic performance across terminals and browsers, with a 1M context window, up to 128K output tokens, and adaptive thinking on by default.

Claude Sonnet 5 is served on EvoLink under the model ID claude-sonnet-5 through Chat Completions · Anthropic Messages, with the same API key and balance you use for every other model. It offers a 1M context window, plus long-context coding, tool use, prompt caching.

Claude Sonnet 5

Claude Sonnet 5 specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
1M tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Thinking mode (on by default)
Tool use
Function calling with multi-step tool sequences
Prompt caching
Cache write + cache read rates
Server-side tools
Web search, billed per successful call
Protocols
Chat Completions · Anthropic Messages
Model ID
claude-sonnet-5

What can you build with the Claude Sonnet 5 API?

Agentic Coding Assistant

Sonnet 5 posts its largest gains over Sonnet 4.6 in coding — planning multi-step changes, running terminals, and iterating on results. With up to 128K output and a 1M context window, it handles large codebases and generates comprehensive diffs, test suites, and implementation plans in a single request.

Autonomous Agents

Build agents that plan a sequence of steps, call tools, read the result, and keep going without a human nudge at every turn. Sonnet 5 is built for terminal and browser use and long-horizon runs, delivering reliable tool use across multi-step workflows.

Large-Context Analysis

Read very large documents, codebases, or long agent traces in a single request with the 1M context window, and let adaptive thinking apply deeper reasoning only when a task needs it — keeping cost predictable for research, planning, and technical strategy.

Two ways to use Claude Sonnet 5: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call Claude Sonnet 5 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls Claude Sonnet 5 through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call Claude Sonnet 5 through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

Claude Sonnet 5 API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "messages": [
      { "role": "system", "content": "You are a senior engineer. Produce clean, well-tested code." },
      { "role": "user", "content": "Review this diff and list the blocking issues with file and line:\n<diff>" }
    ],
    "max_tokens": 8192
  }'

# Anthropic Messages API is also available at /v1/messages with the same model ID.
# The response includes choices[0].message and a usage object
# (prompt_tokens, completion_tokens, and cache read / write tokens when
# caching applies).

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the Claude Sonnet 5 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

How to integrate the Claude Sonnet 5 API

Connect through EvoLink, choose your model ID, and start building in minutes.

1

Step 1 — Create your EvoLink API key

Sign up for EvoLink to get a single API key that routes to Anthropic, Bedrock, or Google Cloud.

2

Step 2 — Select the model ID

Use `claude-sonnet-5` to access the Sonnet 5 model through EvoLink's unified API. It is a drop-in replacement for `claude-sonnet-4-6`.

3

Step 3 — Optimize quality and cost

Adaptive thinking is on by default and can be tuned with the effort parameter; use prompt caching for stable system prompts and repeated long context to lower repeat costs.

Prompt caching billed as two separate dimensions

Cache write and cache read are metered separately, and cache hits are billed at 0.1x the base input rate. Stable system prompts, repository instructions, and tool schemas pay the write once and then read cheaply across a long agent run.

1M-token context with a 128K max output

1M is both the default and the maximum context, with no smaller variant. Keep connected code, specifications, and agent state in one working context, and treat the limit as capacity rather than a target: retrieve relevant evidence, cache the stable prefix, and set task-appropriate output budgets.

Adaptive thinking is on by default and replaces manual budgets

Sonnet 5 scales reasoning to task difficulty automatically. Setting `budget_tokens` for manual extended thinking returns a 400 error — use the effort parameter to guide depth, or pass `thinking: {type: "disabled"}` to turn it off. Non-default `temperature`, `top_p`, and `top_k` also return a 400 error, so remove them when migrating.

Recount tokens before shifting production traffic

Sonnet 5 uses a new tokenizer that produces roughly 30% more tokens for the same text. Per-token pricing is unchanged, but billed cost and context capacity for an equivalent request can differ — recount your prompts, revisit `max_tokens` limits sized close to your expected output, and compare success rate and total cost on production samples.

Claude Sonnet 5 API capabilities

Key specs and model features for production use

1M Context Window

Read very large documents or codebases in a single request. 1M is both the default and the maximum context.

128K Max Output

Generate long-form answers, plans, and code without early truncation.

Adaptive Thinking

Reasoning scales to task difficulty automatically, on by default — tune it with the effort parameter instead of manual thinking budgets.

Agentic Coding & Tool Use

Built for terminal and browser agents, reliable function calling, and long-horizon multi-step runs.

Vision + Multilingual Input

Accept text and image inputs with strong multilingual understanding.

Prompt Caching Rates

Cache writes and reads are priced separately; cache hits are billed at 0.1x the base input price.

Compare Claude routes after workload testing

EvoLink

First verify on your own tasks whether Sonnet 5 already clears the quality bar. Then compare price, context, caching, and workload fit to choose the production route.

Claude Sonnet 5
Input / output$1.8 / $9
Context1M
CachingRead + write
Best forEveryday production coding, terminal and browser agents, long-horizon tool use, and large-context analysis where balanced cost matters more than an Opus-tier rate.
Claude Sonnet 4.6
Input / output$2.7 / $13.5
Context1M
CachingRead + write
Best forThe previous Sonnet generation, kept for workloads already validated against it and for cost-sensitive traffic.
Claude Opus 5
Input / output$4.75 / $23.75
Context1M
CachingRead + write
Best forThe Opus-tier flagship for the hardest reasoning, review, and agent tasks where quality outweighs cost.

Claude Model Family

Same API key and balance — switch tiers without changing your integration.

Compare all Claude models
Claude Sonnet API (Sonnet 5.5)

Claude Sonnet API (Sonnet 5.5)

The newer Sonnet at the same official list price; re-test your workload on it before moving traffic from Sonnet 5.

View model
Claude Opus 5.5

Claude Opus 5.5

Anthropic’s newest Opus for the hardest long-horizon coding, knowledge work, and visual analysis.

View model
Claude Fable 5.1

Claude Fable 5.1

A higher-capability Claude route above Opus for the longest and most demanding agent workloads.

View model

Other text models on EvoLink besides Claude Sonnet 5

GPT-5.6

GPT-5.6

OpenAI's tiered frontier family for routing across capability, latency, and cost.

View model
Grok 4.5

Grok 4.5

xAI's reasoning and tool-use route with a 500K context window and server-side search tools.

View model
Kimi K3

Kimi K3

A long-context reasoning model for repository-scale coding and multi-document work.

View model
Gemini 3.6 Flash

Gemini 3.6 Flash

A fast, low-cost route for high-volume traffic that does not need Opus-tier judgment.

View model

Claude Sonnet 5 API - FAQ

What is the context window and max output for Claude Sonnet 5?

Claude Sonnet 5 supports a 1M token context window — both the default and the maximum, with no smaller variant — and up to 128K output tokens in a single request, making it suitable for large codebases, long documents, and comprehensive generation tasks.

Which model ID should I use, and how do I access it through EvoLink?

Use `claude-sonnet-5` through EvoLink's unified, OpenAI-compatible API to access the Sonnet 5 model with a single API key. It is a drop-in replacement for `claude-sonnet-4-6` — update the model ID and, if needed, review token budgets and thinking settings.

How does Claude Sonnet 5 handle extended thinking?

Sonnet 5 uses adaptive thinking, on by default, which scales reasoning to task difficulty. Manual extended thinking (setting `budget_tokens`) is no longer supported and returns a 400 error — use the effort parameter to guide reasoning depth instead. To turn thinking off, pass `thinking: {type: "disabled"}`.

What changes when migrating from Sonnet 4.6 to Sonnet 5?

Sonnet 5 is a drop-in replacement, but note three things: adaptive thinking is on by default, manual extended thinking returns a 400 error, and non-default sampling parameters (`temperature`, `top_p`, `top_k`) return a 400 error — remove them when migrating. Tool definitions and response shapes are unchanged.

Why do my token counts change on Claude Sonnet 5?

Sonnet 5 uses a new tokenizer that produces roughly 30% more tokens for the same text. Per-token pricing is unchanged, but the cost and context capacity of an equivalent request can differ — recount your prompts and revisit `max_tokens` limits sized close to your expected output.

How is Claude Sonnet 5 priced?

EvoLink pricing: $1.800 per 1M input tokens and $9.000 per 1M output tokens, against Anthropic's $2.000 and $10.000. Prompt caching is billed separately — $2.250 per 1M cache write tokens and $0.181 per 1M cache read tokens on EvoLink, versus $2.500 and $0.200 official.

How is the Claude Sonnet 5 web search tool billed?

Web search is charged per search at $0.010. Tokens the search results add to the request are billed normally.

How does Sonnet 5 compare to Opus 4.8?

Sonnet 5 delivers Opus-class coding and agentic performance and is the right default for most production workloads. Opus 4.8 is the flagship for the hardest reasoning and agent tasks where top-tier quality matters most.

Where is Claude Sonnet 5 available?

Claude Sonnet 5 is available via the Anthropic API, AWS Bedrock, Google Cloud (Vertex AI), and Microsoft Foundry (preview). EvoLink can route to the provider you choose.

Does Claude Sonnet 5 support vision and multimodal input?

Yes. Sonnet 5 supports text and image input with strong multilingual capabilities, so you can combine documents, screenshots, and visuals in one request.