GPT Image 2.5 Flare & Sunburst are live on EvoLinkTry GPT Image 2.5

Grok 4.6 API

Access xAI Grok 4.6—also searched as grok-4-6— through EvoLink's unified chat API. Test reasoning, tool use, server-side tools before integrating.

xAIText GenerationAvailable
from $1.700 / 1M input tokens$2.000 official price-15%
API docs
Long-context reasoningConfigurable reasoningServer-side toolsChat + Responses
Production routeLive
Context
500K context · 450K max output
Best For
Difficult reasoning, research with server-side tools, long-context analysis
Input
Text + images
Output
Text · JSON (structured output) · tool calls

Choose Grok 4.6

An evaluation route for difficult reasoning, long-context analysis, research, and workflows that combine text generation with xAI server-side tools.

Grok 4.6

xAI reasoning and tool-use model

Selected
Model ID
grok-4.6
Best for

Research with web and X search, code-assisted analysis, document and collection retrieval, complex reasoning, and provider-diverse production routing.

< 200K
Input
$1.700 / 1M-15%
115.6 cr / 1M$2.000official price
Cache read
$0.425 / 1M-15%
28.9 cr / 1M$0.500official price
Output
$5.100 / 1M-15%
346.8 cr / 1M$6.000official price
≥ 200K
Input
$3.400 / 1M-15%
231.2 cr / 1M$4.000official price
Cache read
$0.850 / 1M-15%
57.8 cr / 1M$1.000official price
Output
$10.200 / 1M-15%
693.6 cr / 1M$12.000official price

All rates are per 1M tokens, shown in USD and credits, and reflect your account's current pricing. Above 200K prompt tokens the long-context tier applies (×2) to every token role.

Grok 4.6 pricing

Estimate what one Grok 4.6 request costs before you integrate. The calculator uses your account's current rates, with official pricing as a reference.

Request calculator

Enter the token mix for one request and the number of successful tool calls.

Estimated request cost

Grok 4.6
Short-context rate
USD$0.0034
Credits0.2255

Official estimate $0.0039 · save $0.0006 (15%)

Input tokens0.1156 cr
Cache read tokens0.0058 cr
Output tokens0.1041 cr

Minimum charge: 0.01 credits per request. At 200K prompt tokens every token role is billed at ×2. Tool calls are not multiplied.

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 3015 requests

For quick testing

$50
About 15077 requests

For regular development

$100
About 30155 requests

For production evaluation

Server-side tool rates

Only successful server-side calls are billed; X search is metered by the posts and user profiles it fetches, the other tools per call. Failed attempts have no tool fee, but tokens still apply.
  • Web search$0.0050/ call0.34 cr / call
  • X search (posts fetched)$0.0050/ post0.34 cr / post
  • Code execution$0.0050/ call0.34 cr / call
  • Attachment search$0.010/ call0.68 cr / call
  • Collections search$0.0025/ call0.17 cr / call
  • X user profiles$0.010/ profile0.68 cr / profile

What is the Grok 4.6 API?

Grok 4.6 is xAI’s reasoning model for coding agents, connected research, long-document analysis, and tool-driven workflows. The page URL uses grok-4-6, while API requests use the model ID grok-4.6. Grok 4.6 is live on the EvoLink production route; review pricing boundaries, supported workflows, and model configuration in the Pricing and API sections below.

Grok 4.6 is served on EvoLink under the model ID grok-4.6 through Chat Completions · Responses, with the same API key and balance you use for every other model. It offers a 500K context window and up to 450K output tokens, plus reasoning, tool use, server-side tools.

Grok 4.6

Grok 4.6 specs and capabilities

Numbers come from the EvoLink route configuration; capabilities are what the API exposes today.

Context window
500K tokens
Max output
450K tokens
Input
Text + images
Output
Text · JSON (structured output) · tool calls
Reasoning
Configurable reasoning effort
Tool use
Function calling with multi-step tool sequences
Prompt caching
Automatic cache reads at a lower rate
Server-side tools
Web search, X search, Code execution, Attachment search, Collections search; X search billed per post and user profile fetched, other tools per successful call
Long-context tier
×2 above 200K prompt tokens
Protocols
Chat Completions · Responses
Model ID
grok-4.6

grok-4-6 is the page URL. Requests on Chat Completions and Responses must use the exact model ID shown here.

What is the Grok 4.6 API best suited for?

Grok 4.6 combines a 500K-token context window, configurable reasoning, image input, and tools for work that must retain evidence, call external systems, and return verifiable results. These use cases explain where it may add value and what to confirm before integration.

Repository-scale coding and code review

Keep relevant source files, issues, test results, and previous changes in one task for cross-file debugging, implementation planning, and review. Before rollout, use fixed repository tasks to check test pass rate, incomplete steps, structured results, and required human edits instead of judging coding ability from a single demo.

Long documents and multi-source analysis

The 500K-token context can hold reports, contracts, knowledge-base passages, conversation history, and retrieved sources together. More context does not automatically improve an answer, so verify that key evidence is preserved and use the long-context and cached-input rates in Pricing to calculate cost per acceptable result.

Connected research and tool workflows

Grok 4.6 can combine web search, X search, code execution, attachment search, and collections search into sourced research. When using EvoLink, check tool availability, token usage, and per-call charges in the existing API and Pricing sections, then test citations and result quality on your actual research flow.

Structured output and agent orchestration

Text and image inputs can feed JSON Schema, function calling, and multi-step agent flows for extraction, review, and downstream automation. Before production use, test schema validity, function arguments, streaming completion events, and safe recovery after a tool failure.

What changes when you access Grok 4.6 through different platforms?

Even with the same model, platforms can differ in context configuration, tool support, usage reporting, and billing. EvoLink puts model configuration and usage behind one API so later model changes remain straightforward.

Use the correct API model ID

grok-4-6 is the page URL and a common search form; the API request model ID is grok-4.6. Existing Chat Completions or Responses applications can use the correct ID in model configuration, with exact fields available in the API section.

Follow the active API route configuration

xAI documents a 500K-token model window, but platforms may expose different context settings, tools, and rate limits. When using EvoLink, rely on the model configuration, available features, and actual usage shown for the current route.

Choose Chat or Responses by workflow

Start with Chat Completions for standard chat, streaming, and client-side functions. Choose Responses when you need research agents, server-side search, or code execution. This preserves a familiar OpenAI-style integration without adding complexity the workload does not need.

Keep model choice open in one gateway

Use one EvoLink account, balance, and API pattern for Grok, GPT, Claude, and Kimi. Keeping model selection in configuration lets teams route by task quality, cost, and availability without rebuilding application code for every provider.

How can you control Grok 4.6 API cost more accurately?

The existing Pricing section shows current token and tool rates. In practice, caching, context management, reasoning settings, and model routing help reduce unnecessary usage and connect spend to completed business tasks.

Make repeated context cache-friendly

Stable system prompts, tool schemas, and shared context are easier to reuse through caching. Configure the supported cache or conversation identifier for the selected protocol and inspect cached tokens in usage to confirm that repeated requests receive the expected benefit.

Send only the context the task needs

A 500K-token window is useful for large repositories and documents, but it does not need to be filled on every request. Selecting only relevant files, messages, and retrieved passages reduces input cost and helps the model focus on the evidence that matters.

Tune reasoning, output, and tool calls by task

Start simple work with lower reasoning effort and shorter output, then increase the reasoning budget for harder analysis. Research and agent workflows should also track tool-call count to avoid repeated searches, executions, or unproductive loops.

Compare models by total task cost

Evaluate tokens, cached input, server tools, and required retries within the same completed task. EvoLink centralizes model and usage information so teams can compare the total cost of completing equivalent work across Grok 4.6 and other routes.

Two ways to use Grok 4.6: EvoLink API or Agent

Use the EvoLink API for product backends and batch jobs, or call Grok 4.6 from Codex, Claude, or Gemini for coding and analysis workflows. Both paths share the same EvoLink API key, balance, model ID, and request history.

Option 1

Integrate with the EvoLink API

Best for: product backends, batch jobs, automated pipelines

Send OpenAI-compatible Chat Completions (or Anthropic Messages) requests to EvoLink and control the model ID, system prompt, output budget, tools, and structured output.

  1. 1Create an EvoLink API key in the console
  2. 2Point your OpenAI or Anthropic SDK at the EvoLink base URL and select the model ID shown above
  3. 3Send one representative request and read the usage field for input, cached, and output tokens
  4. 4Set max_tokens and retries per task; keep tool-call IDs and results across turns
Option 2

Call it with an Agent

Best for: coding, review, and analysis tasks in Codex, Claude, and Gemini

Give the Agent the task, the inputs to include, and the acceptance criteria. It assembles the request, calls Grok 4.6 through EvoLink, and returns the answer with token usage.

  1. 1Set EVOLINK_API_KEY in your local environment; never put it in code or a prompt
  2. 2Describe the task, the inputs to include, and the expected output format
  3. 3Ask the Agent to call Grok 4.6 through EvoLink and show the request before sending
  4. 4Let the Agent report the answer, token usage, and any error body

Grok 4.6 API code example and error handling

This example shows the shortest runnable request: an OpenAI-compatible Chat Completions call with a system prompt, a user message, and an output budget. Open the API tab for the complete parameter and response reference.

View complete API docs
cURL
curl -X POST https://api.evolink.ai/v1/chat/completions \
  -H "Authorization: Bearer $EVOLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.6",
    "messages": [
      { "role": "system", "content": "You are a research assistant. Cite sources." },
      { "role": "user", "content": "Summarize the three most important changes in <topic> this month." }
    ],
    "max_tokens": 2048,
    "temperature": 0.2
  }'

# Server-side tools (web search, X search, code execution, attachment /
# collections search)
# run on /v1/responses with the same model ID. X search is billed per post
# and user profile fetched; the other tools per successful call.
# The response includes choices[0].message and a usage object (prompt,
# cached, and completion tokens).

Invalid request or unsupported parameter

Check the model ID, the messages array, and parameter ranges against the API reference; remove fields this route does not support.

Authentication or balance issue

Check the Authorization bearer token and confirm the available balance in the console.

Context length exceeded

Prompt tokens exceed the Grok 4.6 context window. Trim or retrieve only the relevant evidence and reuse cached prefixes.

Rate limited (429)

Back off and retry with jitter; batch or queue requests instead of sending parallel bursts.

Content or tool call rejected

Review sensitive content, malformed tool-call arguments, and JSON schema mismatches before retrying.

How should you choose Chat Completions, Responses, and production traffic for Grok 4.6?

The two protocols serve different workflows. This is a selection summary; use the existing API section and EvoLink documentation for exact request fields, and the existing Pricing section for token and tool charges.

01

Standard chat and client functions: Chat Completions

If you already use OpenAI-compatible chat, streaming, or client-side function calling, start with Chat Completions. Confirm message format, streaming completion, function arguments, and usage against what the current client expects.

Chat and client functions
02

Research agents and server tools: Responses

Evaluate Responses when you need web or X search, code execution, attachment or collections search, and longer multi-step agent flows. Confirm tool availability, call charges, citations, failure states, and retry boundaries before integration.

Server tools
03

Large-context work: watch cache and total cost together

Long documents, repositories, and conversations should not automatically put everything into one request. Compare representative context sizes, cache hits and misses, output length, and retries, then use the final charge to calculate successful-task cost.

Context cost
04

Production traffic: start small and keep a fallback

Send a small, observable set of tasks to Grok 4.6 first and retain a proven GPT, Claude, or Kimi route. EvoLink’s unified API centralizes model choice, usage, and balance, making it easier to change routes after rate limits, timeouts, schema errors, or tool failures.

Gradual rollout

Cached input is billed at its own lower rate

Stable prefixes — system prompts, tool schemas, shared context — are cached automatically on repeat requests and billed as cache reads instead of fresh input. Keep the prefix identical across turns and check cached tokens in usage to confirm the saving.

500K-token context with a long-context tier from 200K prompt tokens

Fill the window only when a task needs it. Once a prompt reaches 200K tokens the long-context tier applies to every token role, so trim retrieved passages and history before scaling up, and watch the tier badge in the estimator when planning budgets.

Compare Grok 4.6 with GPT-5.6 and Claude Opus 5

EvoLink

Compare input/output rates, context, caching, and workload fit. Benchmark the same requests before choosing a route; prices below follow your account’s current rates.

Grok 4.6
Input / output$1.7 / $5.1
Context500K
CachingAutomatic cache reads
Best forResearch with web and X search, code-assisted analysis, document and collection retrieval, complex reasoning, and provider-diverse production routing.
GPT-5.6
Input / output$0.181 / $1.081
Context1.05M
CachingCache reads / writes
Best forCoding, production agents, complex reasoning
Claude Opus 5
Input / output$4.75 / $23.75
Context1M
CachingCache reads / writes
Best forRepository-scale coding, long-horizon agents, high-stakes review

What should you confirm before using Grok 4.6 in production?

Start with a small set of real workloads to determine whether the model meets your quality, latency, cost, and reliability needs before deciding which traffic should move.

Integration and usage data are clear

Confirm that the application uses the correct grok-4.6 model ID and intended protocol, and that required responses, usage, and cache information are returned. Clear usage data supports cost analysis and gives Chat and Responses workflows a consistent observation method.

Outputs meet real business requirements

Test real code changes, long-document analysis, research, or structured extraction. Beyond answer quality, check whether tests pass, citations are reliable, function arguments are correct, and JSON Schema output can be consumed directly by downstream systems.

Latency and error handling meet expectations

Observe response time under representative traffic and prepare retry behavior for rate limits, timeouts, invalid structured output, and tool failures. Configurable model selection in EvoLink makes it easier to switch to a verified alternative when one route is temporarily unavailable.

Cost and model choice remain controllable

Use Pricing, usage, and final charges to calculate total cost for the same class of task, then decide whether Grok 4.6 belongs on the default route, difficult tasks, or fallback traffic. A unified gateway keeps quality and budget decisions separate from integration work.

Begin with a small, observable, reversible set of Grok 4.6 tasks. Expand only after quality, latency, and cost meet expectations. Keeping multiple model options behind the EvoLink unified API makes later scaling, switching, and cost optimization easier.

Grok Model Family

Same API key and balance — switch tiers without changing your integration.

Grok 4.6

Grok 4.6

Current model

xAI reasoning and tool-use model

Grok 4.7

Grok 4.7

xAI reasoning and tool-use model

View model
Grok 4.5

Grok 4.5

xAI reasoning and tool-use model

View model

Other text models on EvoLink besides Grok 4.6

GPT-5.6

GPT-5.6

OpenAI’s tiered frontier family for comparing capability, latency, and cost-routing flexibility.

View model
Claude Opus 5

Claude Opus 5

Anthropic’s premium route for long-running agents, tool use, and complex review.

View model
Kimi K3

Kimi K3

Moonshot’s long-context reasoning route with a separate cached-input rate.

View model
DeepSeek V4 Flash

DeepSeek V4 Flash

DeepSeek’s cost-sensitive open-model route for high-volume coding, reasoning, and agent workloads.

View model

Grok 4.6 guides and related reading

Grok 4.6 vs Grok 4.5

Grok 4.6 vs Grok 4.5

Compare benchmarks, cached-input cost, reasoning controls, and a safe production upgrade plan.

Read guide
Grok 4.6 vs Kimi K3

Grok 4.6 vs Kimi K3

Choose a route by coding workload, context size, open-weight requirements, cost, and fallback.

Read guide

Grok 4.6 API FAQ

Is the Grok 4.6 API available through EvoLink now?

Yes. Grok 4.6 is live on the EvoLink production route. Send requests with the model ID grok-4.6 on Chat Completions or Responses, and check the model ID, pricing, context, and supported workflows on this page before scaling up.

Is grok-4-6 the API model ID?

No. grok-4-6 is the page slug and search label. The actual request model ID is grok-4.6, shown in the model ID section on this page.

What is the Grok 4.6 context window and long-context pricing boundary?

xAI documents a 500,000-token context window. Requests enter the long-context price tier at the prompt-token threshold shown in Pricing, so test representative context sizes and cached input instead of assuming every request should use the full window.

What inputs and outputs does Grok 4.6 support?

Upstream documentation lists text and image input with text output. Image input means the model can analyze visual information; it does not make Grok 4.6 an image-generation model. Confirm the effective EvoLink route limits during validation.

Should I use Chat Completions or the Responses API?

Use Chat Completions for familiar chat, streaming, and client function flows. Evaluate Responses for longer-running agent workflows and server-side tools. The current API section and EvoLink docs remain the source for exact request fields and verified route support.

How should reasoning effort be selected?

Start with the lowest setting that meets the task quality requirement, then compare medium and high on the same tasks. Complex work may benefit from a larger reasoning budget, while simple work can prioritize response time and cost.

How do cached input and long agent loops affect Grok 4.6 pricing?

Cached input can reduce repeated-context cost, while cache misses, long outputs, retries, and repeated tool steps can make an agent loop more expensive than the headline input rate suggests. Use the live Pricing section and final request charges when calculating cost per accepted task.

What is the difference between server tools and client functions?

Client functions run in your application. Web and X search, code execution, attachment search, and collections search are xAI server tools. They differ in execution location, failure handling, and billing; check the current support and rates in the API and Pricing sections.

How should teams compare Grok 4.6 with GPT, Claude, or Kimi?

Run the same real tasks with consistent context, tools, and reasoning settings. Compare result quality, response time, token mix, tool performance, and total cost to decide which EvoLink route fits each traffic class.

What fallback should teams keep during rollout?

Keep a model that already handles the same workload reliably and leave route selection configurable. If rate limits, timeouts, or invalid output occur, EvoLink can route to GPT, Claude, Kimi, or another suitable alternative.