GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

API Updates

Stay informed about the latest changes and improvements across all our APIs.

September 2026

2026-09-30

New Models | Suno V6, Suno V6 Mini and Suno V6 Wild

models

Suno V6, Suno V6 Mini and Suno V6 Wild are now live on EvoLink, and Suno V6 becomes the default recommended version on the Suno page. All three cost the same as the existing Suno music models, billed per generation, with 2 tracks per generation.

Suno V6

  • Model ID: suno-v6-beta
  • Highlights: more natural vocals, richer detail and stronger musical expression
  • Supports: Persona and custom duration (60 seconds or longer recommended)

Suno V6 Mini

  • Model ID: suno-v6-mini-beta
  • Highlights: lightweight and fast, balancing quality and speed
  • Supports: Persona; custom duration is not supported yet
  • Try it: switch to Suno V6 Mini in the model selector

Suno V6 Wild

  • Model ID: suno-v6-wild-beta
  • Highlights: bolder, more distinctive musical expression
  • Supports: Persona and custom duration
  • Try it: switch to Suno V6 Wild in the model selector

Shared by all three models

  • Pricing: same as the existing Suno music models, billed per generation regardless of track length. See pricing
  • Length limits: in custom mode (custom_mode=true), lyrics up to 5000 characters and style up to 1000 characters
  • Persona and duration: available in custom mode only
  • Async tasks: submit, then poll status or use callback_url

All Suno versions

  • The prompt limit in simple mode (custom_mode=false) is raised from 500 to 3000 characters
2026-09-29

New Model | GPT-6.1 Sol

models

GPT-6.1 Sol is now live on EvoLink. OpenAI positions it as near-Astra performance for complex work at a lower cost: complex coding, computer use, and professional work.

GPT-6.1 Sol

  • Model ID: gpt-6.1-sol — written with a dot; GPT-6 Sol stays available under its own ID
  • Built for: complex coding and debugging, computer use, and multi-step professional work
  • Context window: 1,050,000 tokens, up to 128K max output
  • Pricing: uncached input, cache read, cache write and output are metered separately; OpenAI's cache read price is half that of GPT-6 Sol. Current rates are on the GPT-6.1 Sol model page
  • Long context: requests above 272K prompt tokens bill the entire request at long-context rates — 2x for input and both cache roles, 1.5x for output
  • Protocols: OpenAI-compatible Chat Completions and Responses, on the same EvoLink API key

Changes from GPT-6 Sol

  • Reasoning: reasoning_effort (Chat) or reasoning.effort (Responses) accepts low, medium (default), high, and xhigh, plus max on Responses; none is not supported, so switch none to low
  • Tool calling: Chat Completions accepts requests without tools only; send function calling and built-in tools to Responses
  • Parameters: remove temperature, top_p, logprobs, and top_logprobs from requests
2026-09-29

Claude Sonnet 5.5 | Anthropic's Newest Sonnet for Coding and Agents

models

Claude Sonnet 5.5 from Anthropic is now live on EvoLink — the successor to Claude Sonnet 5 at the same list price, for everyday coding, tool-using agents, and high-volume production traffic, with a 1M-token context window and up to 128K max output.

Claude Sonnet 5.5

  • Model ID: claude-sonnet-5-5
  • Context window: 1M tokens
  • Max output: 128K tokens
  • Capabilities: agentic coding in real repositories, reliable tool use, adaptive thinking (on by default), vision input, prompt caching
  • Recommended workloads: everyday coding, tool-using agents, and high-volume production traffic
  • Billing: token-based — Input, Output, Cache Write, Cache Hit; the web search tool is billed per search on routes that support it. See current rates
  • API: OpenAI-compatible Chat Completions and Anthropic Messages on the same EvoLink API key
2026-09-29

GLM, DeepSeek & Kimi K3 | Lower Pricing

pricing

GLM, DeepSeek and Kimi K3 get lower prices on EvoLink, effective September 30, 2026 (UTC+8). The update covers GLM-5.3, GLM-5.3 Flash, GLM-5.3 FlashX, GLM-5.2, DeepSeek V4 Pro, DeepSeek V4 Flash, DeepSeek V4.1 Flash and Kimi K3: input and output are lower on every one of these models, and cache reads are lower or unchanged on most of them. The new rates apply automatically, with no changes to your API calls.

  • Cache reads on GLM-5.3 and GLM-5.3 FlashX rise slightly under the new rates
  • deepseek-v4-flash-vision-exp is now served by DeepSeek V4.1 Flash and billed at DeepSeek V4.1 Flash rates
2026-09-24

New Models | Seedream 5.0 Flash and Seedream 5.0 Flash Layerize

models

Seedream 5.0 Flash and Seedream 5.0 Flash Layerize are now live on EvoLink. Flash is the faster Seedream 5.0 model for image generation and editing; Flash Layerize splits one image into editable layers. Both bill one flat price per output image, 10% below the official price.

Seedream 5.0 Flash

  • Model ID: doubao-seedream-5.0-flash
  • Generation and editing: text-to-image, image-to-image and editing with up to 10 reference images; size accepts auto, an aspect ratio or custom pixels; quality is 1K, 1.5K or 2K. Transparent output preserves existing alpha from one reference image and requires PNG; it does not remove backgrounds
  • Pricing: one flat rate per output image — 1K, 1.5K and 2K cost the same, and reference images are not billed. See pricing

Seedream 5.0 Flash Layerize

  • Model ID: doubao-seedream-5.0-flash-layerize
  • Input: exactly one png or jpeg image; prompt is optional
  • Output: one base image plus up to 16 transparent PNG layers, each with its stacking order and bounding box
  • Pricing: every output image — the base image and each layer — is billed at the same flat rate as Flash. Up to 17 images are reserved when the task starts and the unused part is refunded

Shared by both models

  • Same request as Seedream 5.0 Pro: quality is 1K, 1.5K or 2K (Layerize also takes auto)
  • Prompt mode: prompt_priority accepts standard only — requests with fast are rejected and not billed
  • Async tasks: submit, then poll status or use callback_url
2026-09-23

New Models | GPT-6 Sol and GPT-6 Luna

models

GPT-6 Sol and GPT-6 Luna join GPT-6 Astra on EvoLink. Sol is built for complex coding and agentic workflows; Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks.

GPT-6 Sol

  • Model ID: gpt-6-sol — there is no generic gpt-6 alias
  • Built for: complex coding, multi-file changes, and tool-driven agents
  • Pricing: uncached input, cache read, cache write and output are metered separately. Current rates are on the GPT-6 Sol model page

GPT-6 Luna

  • Model ID: gpt-6-luna
  • Built for: classification, extraction, routing, and subagent calls at high volume
  • Pricing: metered the same four ways. Current rates are on the GPT-6 Luna model page

Shared by both models

  • Context window: 1,050,000 tokens, up to 128K max output
  • Long context: requests above 272K prompt tokens bill the entire request at long-context rates — 2x for input and both cache roles, 1.5x for output
  • Protocols: OpenAI-compatible Chat Completions and Responses, on the same EvoLink API key
  • Reasoning: reasoning_effort (Chat) or reasoning.effort (Responses) accepts none, low, medium (default), high, xhigh, and max — none is new compared with GPT-6 Astra
  • Tool calling: use Responses for built-in tools and function calling; on Chat Completions, function calling works only with reasoning_effort set to none
2026-09-23

GPT-5.6 Sol | Lower Pricing

pricing

GPT-5.6 Sol gets an across-the-board price cut on EvoLink, following OpenAI's new list price for the model, effective September 23, 2026 (UTC+8). Input, cache read, cache write and output are all lower, including long-context rates. The new rates apply automatically — keep calling gpt-5.6-sol, with no changes to your API calls.

  • GPT-5.6 Terra and GPT-5.6 Luna pricing is unchanged
2026-09-22

New Model | Claude Opus 5.5

models

Claude Opus 5.5 — Anthropic's newest Opus model, succeeding Claude Opus 5 at a lower list price — is now live on EvoLink.

Availability

Claude Opus 5.5 can be called on the EvoLink production API with the same EvoLink API key. The Claude Opus 5.5 model page is the source of truth for live pricing and parameters.

Model ID

  • claude-opus-5-5 — 1,000,000-token context window, up to 128K max output

Billing dimensions

Five dimensions are metered separately: input, cache write, cache read, output, and the web search tool. The web search tool is billed per search, and search results additionally count as ordinary input tokens. See the pricing section for current rates.

Coming from Claude Opus 5

Anthropic lists Opus 5.5 below Opus 5 on every token rate, and cache reads drop from a tenth to a twentieth of the input rate. Three request settings behave differently: thinking cannot be turned off — send thinking: {type: "adaptive"} or omit the parameter and control depth with effort; the default effort is medium instead of high; and forced tool choice (any or tool) is rejected, so use auto.

Protocols

OpenAI-compatible Chat Completions and Anthropic Messages, on the same EvoLink API key.

Try Claude Opus 5.5 | View pricing

2026-09-22

Grok 4.7 | xAI's Newest Reasoning Model, at Grok 4.6 Rates

models

Grok 4.7 is live on EvoLink with a 500,000-token context window, available on both Chat Completions and Responses.

Pricing — identical to Grok 4.6, line for line

  • Input: $2.00 per 1M tokens (list price)
  • Cached input: $0.50 per 1M tokens
  • Output: $6.00 per 1M tokens, reasoning tokens included
  • Prompts at or above 200,000 tokens move every token in the request to the long-context tier at 2×
  • The model page shows the rate your account actually pays

What changed from Grok 4.6

  • Request model ID is grok-4.7. The page slug grok-4-7 is not a valid request value
  • reasoning_effort accepts low, medium, high (default) and xhigh
  • On Responses, encrypted reasoning comes back on every response, without listing it in include
  • Knowledge cutoff moves to May 2026
2026-09-22

Grok X Search | Now Billed by Posts and Profiles Fetched

modelspricing

X search on Grok is now billed by what a search returns, not by how many times it runs. xAI changed this on September 21, 2026 at 12:00 PM PT, and it applies to grok-4.5, grok-4.6 and grok-4.7.

Rates — effective September 21, 2026

  • Posts fetched: $5 per 1,000 posts, replacing $5 per 1,000 calls
  • User profiles fetched: $10 per 1,000 profiles, a new billing item
  • Every post a search or thread fetch returns counts, including parent and quoted posts, with no de-duplication
  • Other server-side tools are unchanged: web search and code execution at $5 per 1,000 calls, attachment search at $10 per 1,000, collections search at $2.50 per 1,000

What this changes for your bill

  • Under the old model, one x_search call cost the same whether it returned 1 post or 30. It is now counted per item
  • Searches that pull large threads get noticeably more expensive; cap the volume with search_parameters.max_search_results
  • Token charges are unaffected and are still calculated separately from tool fees
2026-09-22

Seedance 2.5 Update | Draft Mode and Draft-to-Video

models

Seedance 2.5 now supports a two-step draft workflow: preview fast, then generate the final only when it looks right.

New

  • Draft mode: pass draft: true on any of the five Seedance 2.5 generation models (or switch on Draft mode in the Playground) to get a fast 480p preview of scene structure, camera moves and subject motion. Billed the same as a normal 480p video
  • seedance-2.5-draft-to-video: convert a completed draft into a 1080p final video with just source_task_id and an optional output_format. Prompt, assets, duration, ratio, seed and audio settings are reused from the draft

Notes

  • The draft and the final video are billed separately, and the final is billed at the 1080p rate; the content filter setting carries over from the draft and cannot be changed at conversion time
  • A completed draft returns draft_expires_at; convert before that time. Converting an expired draft returns 410 draft_task_expired
2026-09-21

GLM-5.3 FlashX | New Throughput Tier in the 5.3 Family

models

GLM-5.3 FlashX is live. It sits between GLM-5.3 Flash and the GLM-5.3 flagship: the same natively multimodal 1M-token context and always-on reasoning, at roughly a quarter of the flagship price.

  • Model ID glm-5.3-flashx, available over Chat Completions, Responses, and Anthropic Messages
  • Text, image, video, and file input; prompt caching reads billed on their own lower rate
  • Web search is metered per call, separate from token usage
  • Reasoning effort accepts low, high, and max; other values are mapped to the nearest supported level

The model page shows the live per-token rate for your account. See the pricing page for the current numbers.

2026-09-17

New Model | MiniMax H3 Max Reference-to-Video

models

MiniMax H3 Max now takes reference media — images, video and audio — through a new route. Output is billed per second; reference video carries its own, higher rate.

MiniMax H3 Max Reference-to-Video

  • Model ID: minimax-h3-max-reference-to-video
  • Reference media: up to 9 images, 3 videos and 3 audio clips, 12 items in total. At least one reference image or one reference video is required — audio on its own is rejected
  • Output: 768p or 480p MP4, 5–15 seconds, aspect ratio follows the reference material by default. H3 Max has no 2K tier
  • Pricing: three parts — output seconds, reference-video seconds at a separate and higher per-second rate, and reference images past the free allowance. Reference audio is not charged. See pricing
  • Free image allowance: the first 2 reference images are free. MiniMax H3 gives 5, so the same brief can cost more here — check the estimate before moving a workload across
  • Not accepted: image_start / image_end
  • Async tasks: submit, then poll status or use callback_url
2026-09-17

New Model | MiniMax H3 Max Turbo

models

MiniMax H3 Max Turbo — the speed tier of MiniMax H3 Max — is now live on EvoLink as two routes, with 768p / 480p output and per-second billing.

MiniMax H3 Max Turbo

  • Model IDs: minimax-h3-max-turbo-text-to-video, minimax-h3-max-turbo-image-to-video
  • Output: 768p or 480p MP4, 5–15 seconds. Aspect ratio 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 — text-to-video needs an explicit ratio (defaults to 16:9)
  • Pricing: per output second, 5% below the official Turbo rate at the same resolution; minimum charge equals the cheapest valid request (480p × 5s). See pricing
  • Image-to-video: image_start (first frame) is required and image_end (last frame) is optional. Sending a last frame on its own is not supported. This route does not take an aspect ratio — the output follows the input image proportions
  • Default resolution: requests that omit quality generate at 768p. Send "quality": "480p" for the cheapest tier
  • Not on Turbo: 2K output, 4-second clips, and every form of reference-driven generation — no reference images, video or audio. Use MiniMax H3 or MiniMax H3 Max when a brief needs reference media
  • Async tasks: submit, then poll status or use callback_url
2026-09-17

Seedance 2.5 | 1080p Back to Regular Pricing

pricing

The limited-time 28% discount on Seedance 2.5 1080p output ended on September 17, 2026 (UTC+8), as scheduled. 1080p is back to its regular rate on all five models. No changes to your API calls.

  • 480p and 720p rates are unchanged
  • Applies to seedance-2.5-text-to-video, seedance-2.5-image-to-video, seedance-2.5-reference-to-video, seedance-2.5-video-edit, and seedance-2.5-video-extend

The Playground estimate and the pricing page already show the regular rate. See the pricing page for the current per-second rate.

2026-09-16

MiniMax H3 Max | 5% Price Cut Across Both Tiers

modelspricing

MiniMax H3 Max gets a 5% price cut across both output tiers, covering text-to-video and image-to-video.

Pricing — 5% off, effective September 16, 2026

  • 768p: $0.076 per output second, was $0.080
  • 480p: $0.0475 per output second, was $0.050
  • Minimum charge: $0.2375 per task, was $0.250 (480p × 5 seconds)
  • Both routes share the same rates: minimax-h3-max-text-to-video and minimax-h3-max-image-to-video

Choose your output tier

  • Requests that omit quality use 768p. For the lowest rate, send "quality": "480p" explicitly
  • Both tiers support 5–15 seconds of output, billed per output second
2026-09-10

New Model | DeepSeek V4.1 Flash

models

DeepSeek V4.1 Flash adds native image understanding for coding, document and screenshot workflows. See the model page for current pricing before rollout.

DeepSeek V4.1 Flash

  • Model ID: deepseek-v4.1-flash on EvoLink; the page URL is /deepseek-v4-1-flash
  • Context window: 1,000,000 tokens, with a maximum output of 384,000 tokens; choose an output budget for the task
  • Billing: input, cache hit and output have separate rate entries. Check the model pricing section and final account charges; per-token rates alone do not determine task cost
  • Image input: PNG, JPEG, WebP and GIF. Measure representative image usage before scaling
  • Protocols: Chat Completions, Messages and Responses take the same model ID; image fields differ: image_url, an image block and input_image
  • Thinking and caching: Thinking is enabled by default and prefix caching is automatic; record the thinking setting you send, check the protocol documentation, and read cached tokens from usage
  • Legacy routes: on EvoLink, deepseek-v4-flash and deepseek-v4-pro are not affected and continue to serve V4 Flash and V4 Pro; deepseek-v4-flash-vision-exp now redirects to V4.1 Flash. See the migration guide
2026-09-09

New Model | GPT Image 2.5 Sunburst & Flare

models

GPT Image 2.5 Sunburst and GPT Image 2.5 Flare are now available on EvoLink — the two GPT Image 2.5 models OpenAI released on September 8, 2026, at the same token rates as GPT Image 2.

GPT Image 2.5 Sunburst

  • Model ID: gpt-image-2.5-sunburst (no alias; use the exact ID)
  • Positioning: OpenAI's pick for workflows where editing precision matters most, with tighter control across multi-turn edits and longer generation times
  • Input: text prompt, up to 16 reference images and an optional inpainting mask; output: image
  • Quality: five tiers low / medium / high / xhigh / max (default medium); size auto, 15 aspect ratios or explicit WxH; resolution 1K / 2K / 4K
  • Billing: token-based — image output, image input, cached input and text input tokens are settled on the upstream usage object; rates in the pricing section
  • Endpoint: POST /v1/images/generations, async task with polling or callback_url

GPT Image 2.5 Flare

  • Model ID: gpt-image-2.5-flare (no alias; use the exact ID)
  • Positioning: OpenAI's fast default for high-quality everyday generation — higher quality than GPT Image 2 at up to 50% lower latency, by OpenAI's figure
  • Input: text prompt, up to 16 reference images and an optional inpainting mask; output: image
  • Quality: five tiers low / medium / high / xhigh / max (default medium); size auto, 15 aspect ratios or explicit WxH; resolution 1K / 2K / 4K
  • Billing: token-based — image output, image input, cached input and text input tokens are settled on the upstream usage object; rates in the pricing section
  • Endpoint: POST /v1/images/generations, async task with polling or callback_url

Same token rates do not mean the same bill: each quality tier spends output tokens differently, and GPT Image 2.5 uses a lower tile base per tier than GPT Image 2 (its high costs about what GPT Image 2 medium did). Estimate a request in the Playground or the price calculator before you standardise on a tier.

2026-09-07

Discount Extended | Seedance 2.0 Mini & Fast

pricing

The limited-time price drop on Seedance 2.0 Mini and Seedance 2.0 Fast, originally set to end today, is extended through October 6, 2026 (UTC+8). Rates and discount levels are unchanged, and the discount still applies automatically — no changes to your API calls. Seedance 2.0 Standard remains outside the scope of this promotion.

  • Seedance 2.0 Mini: 60% off
  • Seedance 2.0 Fast: 25% off

The reduction continues to apply across every resolution and every generation mode (text-to-video, image-to-video, reference-to-video). See each model's pricing page for current rates.

The promotion now ends October 6, 2026 (UTC+8), after which both models return to their original pricing.

2026-09-04

New Model | GPT-6 Astra

models

GPT-6 Astra — OpenAI's most capable model, built for the hardest end-to-end reasoning, coding, computer-use, and research work — is now live on EvoLink at 10% below OpenAI list pricing.

GPT-6 Astra

  • Model ID: gpt-6-astra — there is no generic gpt-6 alias
  • Context window: 1,050,000 tokens, up to 128K max output
  • Billing: uncached input, cache read, cache write and output are metered separately, each 10% below OpenAI list. Current rates are on the GPT-6 Astra model page
  • Long context: requests above 272K prompt tokens bill the entire request at long-context rates — 2x for input and both cache roles, 1.5x for output
  • Protocols: OpenAI-compatible Chat Completions and Responses, on the same EvoLink API key; OpenAI requires Responses for tool calling
  • Reasoning: set reasoning_effort (Chat) or reasoning.effort (Responses) to low, medium, high, xhigh, or max — none and minimal are rejected, and temperature / top_p are not accepted
  • Also supported: prompt caching and structured outputs on both surfaces
2026-09-03

New Model | MiniMax H3 Max

models

MiniMax H3 Max — the fast generation variant in MiniMax's Hailuo family — is now live on EvoLink as two routes, with 768p / 480p output and per-second billing.

MiniMax H3 Max

  • Model IDs: minimax-h3-max-text-to-video, minimax-h3-max-image-to-video
  • Output: 768p or 480p MP4, 5–15 seconds. Aspect ratio 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 — text-to-video needs an explicit ratio (defaults to 16:9)
  • Pricing: $0.080 per output second at 768p, $0.050 at 480p, same rate on both routes; minimum charge equals the cheapest valid request (480p × 5s, $0.25 per task)
  • Image-to-video: send a first frame, a last frame, or both. This route does not support choosing an aspect ratio — the output always follows the input image proportions
  • Default resolution: requests that omit quality generate at 768p. Send "quality": "480p" for the cheapest tier
  • Not on H3 Max: 2K output, 4-second clips, and reference-driven generation. Use MiniMax H3 when you need 2K, or reference images, video and audio
  • Async tasks: submit, then poll status or use callback_url
2026-09-02

New Model | Gemini 3.8 Flash

models

Gemini 3.8 Flash is now available on EvoLink, priced the same as Gemini 3.7 Flash.

It is Google's most capable Flash model to date, with clear gains over 3.7 in software engineering, computer use, and multi-step agentic work while keeping Flash-tier throughput and cost.

Gemini 3.8 Flash

  • Model ID: gemini-3.8-flash
  • Context: 1,048,576 tokens; max output: 65,536 tokens
  • Input: text, image, audio, video, PDF; output: text
  • Capabilities: thinking, system instructions, structured output, context caching, function calling, code execution, Google Search / Maps grounding
  • Pricing: input $0.675/M tokens, output $3.375/M tokens, cache hit $0.0676/M tokens (promotional rate, through December 31, 2026)
  • Endpoints: both /v1/chat/completions (OpenAI-style) and the native Gemini API

Thinking levels are unchanged from 3.7

Gemini 3.8 Flash does not support the minimal level, same as 3.7. Available levels are low / medium / high, defaulting to medium.

When migrating from 3.7, if your code sets reasoning_effort: "none" or "minimal" (thinkingLevel: "MINIMAL" on the native API), EvoLink automatically downgrades it to low — your requests will not fail. For explicit control, set low directly.

Beyond the model ID, nothing else needs to change: 3.8 matches 3.7 field for field on context window, max output, modalities, and official pricing. If you are coming from 3.6 Flash or earlier, Google's migration checklist still requires removing custom temperature, top_p, top_k, and candidate_count, replacing the numeric thinking_budget with the thinking_level string, and removing prefilled model turns.

2026-09-02

New Model | Claude Fable 5.1

models

Claude Fable 5.1 — Anthropic's most capable model, the Fable tier above Opus — is now live on EvoLink at 5% below official token pricing.

Availability

Claude Fable 5.1 can be called on the EvoLink production API with the same EvoLink API key. The Claude Fable 5.1 model page is the source of truth for live pricing and parameters.

Model ID

  • claude-fable-5-1 — 1,000,000-token context window, up to 128K max output

Billing dimensions

Five dimensions are metered separately: input, cache write, cache read, output, and the web search tool. Token rates run 5% below Anthropic's official pricing; the web search tool is billed per search at the official rate with no discount, and search results additionally count as ordinary input tokens. See the pricing section for current rates.

Coming from Claude Fable 5

Anthropic's list rates for input, output and cache write are unchanged. The difference is cache read, which Anthropic cut to a quarter of the Fable 5 rate — cache-heavy agent runs and long coding sessions get noticeably cheaper. On EvoLink the other three rates sit about 5.6% above Fable 5, because Fable 5 runs at 10% off list while 5.1 runs at 5% off.

Protocols

OpenAI-compatible Chat Completions and Anthropic Messages, on the same EvoLink API key. Thinking is always on: send thinking: {type: "adaptive"} or omit the parameter, and control depth with effort instead of sampling parameters.

Try Claude Fable 5.1 | View pricing

2026-09-01

Pricing Update | Claude Sonnet 5

pricing

Anthropic has made its Claude Sonnet 5 price reduction permanent. EvoLink's launch offer becomes the long-term rate at the same time: what you pay per token does not change, and it will no longer rise when the promotion expires.

What changed

  • Official list price is now permanently lower: input $3.00 → $2.00/1M tokens, output $15.00 → $10.00/1M tokens, cache write $3.75 → $2.50/1M tokens, cache read $0.30 → $0.20/1M tokens.
  • Your EvoLink rate is unchanged. The limited-time launch offer that was scheduled to end on August 31 is now the standard rate, so no expiry increase is coming.
  • The discount on the page reads differently. It was measured against the old $3.00 list price; now that the official price is permanently $2.00, the same EvoLink rate shows as 10% below official instead. The number you are billed is identical.

What did not change

The model ID stays claude-sonnet-5, with the same 1M-token context window and 128K max output, on both OpenAI-compatible Chat Completions and Anthropic Messages. No changes to your API calls are required. The web search tool is still billed per search at the official rate.

Current rates are always shown on the Claude Sonnet 5 model page.

Claude Sonnet 5 | View pricing