Gemini 3.5 Flash Lite API
Choose Gemini 3.5 Flash Lite
Google’s cheapest, fastest 3.5-class model for low-latency classification, routing, extraction, and lightweight agent subagents, with a full 1M-token context at $0.30 / $2.50 per 1M tokens.
Gemini 3.5 Flash Lite
Google fastest, most cost-effective 3.5-class model
gemini-3.5-flash-liteHigh-throughput classification and extraction, low-latency real-time actions, focused agent subagents, and cost-sensitive workloads where speed and price matter more than maximum reasoning.
Gemini 3.5 Flash Lite pricing
Estimate a request with the interactive pricing calculator. All user groups use the official Gemini 3.5 Flash Lite rate.
Token calculator
Enter the token mix for one request.Estimated request cost
Gemini 3.5 Flash LiteMinimum charge: 0.01 credits per request.
EvoLink vs Google direct
Same token mix, default group price.Budget guide
Approximate requests using the current token mix.For quick testing
For regular development
For production evaluation
Model pricing
| Model | Context | Input tokens | Cache read tokens | Output tokens |
|---|---|---|---|---|
Gemini 3.5 Flash Litegemini-3.5-flash-lite | All context sizes | $0.271 / 1M-10% 18.4 cr / 1M$0.300official price | $0.028 / 1M-7% 1.9 cr / 1M$0.030official price | $2.250 / 1M-10% 153 cr / 1M$2.500official price |
Gemini 3.5 Flash Lite
All context sizesUSD and credits are shown per 1M tokens. Live backend pricing takes priority over these frozen fallback rates.
Gemini 3.5 Flash Lite API for fast, low-cost, high-throughput workloads
Call Google’s cheapest, fastest 3.5-class model through EvoLink’s unified API. Gemini 3.5 Flash Lite is built for low-latency classification, routing, extraction, and lightweight agent subagents — with a 1,048,576-token context window, prompt caching, structured output, and tool use — at $0.30 / $2.50 per million input / output tokens.
Where Gemini 3.5 Flash Lite earns a place in a production model stack
Google positions Gemini 3.5 Flash Lite as its fastest, most cost-effective 3.5-class model. Its strongest fit is high-volume, latency-sensitive work where low price and speed matter more than maximum reasoning — and where you can escalate the few hard requests to a bigger model.
High-throughput classification and extraction
Gemini 3.5 Flash Lite is built for high-throughput classification, routing, and JSON extraction. At $0.30 / $2.50 per 1M tokens it lets you process large request volumes cheaply where a premium model would be wasted.
Low-latency, cost-sensitive scale
Google reports about 350 output tokens per second, making it a strong fit for real-time UI actions, autocomplete, moderation, and other latency-critical paths. Measure cost and speed per request, not headline benchmarks.
Lightweight and focused agent subagents
It suits subagents that execute focused tasks inside larger multi-agent systems. It defaults to minimal thinking for speed; for multi-step tool use, validate that agents do not terminate early before you scale.
When to escalate to a bigger model
Deep repository refactoring, long autonomous coding, and hard multi-step reasoning belong on Gemini 3.6 Flash or a Pro-tier model. Use Flash Lite for the high-volume majority and route only the difficult minority upward.
What early Gemini 3.5 Flash Lite reactions suggest—and what still needs proof
Gemini 3.5 Flash Lite launched on 2026-07-21 as a replacement for Gemini 2.5 Flash. Treat launch-day reactions as hypotheses to verify on your own tasks, tools, budgets, and acceptance criteria.
Price-to-performance is the headline
At $0.30 input / $2.50 output per 1M tokens and roughly 350 tokens/second, its appeal is throughput per dollar. On simple, high-volume work this can beat pricier models on total cost with acceptable quality.
It replaces Gemini 2.5 Flash and lighter 3-Flash workloads
Google frames it as a drop-in for Gemini 2.5 Flash and less complex Gemini 3 Flash tasks. If you run those today, re-benchmark your prompts before migrating so quality and formatting stay within tolerance.
Minimal thinking is the default—validate multi-step agents
Flash Lite defaults to minimal thinking for speed and cost. Google notes that minimal thinking can cause premature tool termination on multi-step tasks, so test agent completion before high-throughput rollout.
It is a closed, API-only Google model
Gemini 3.5 Flash Lite has no open weights and cannot be self-hosted. Access is through the Gemini API and platforms such as Google AI Studio, plus unified gateways like EvoLink—route by cost and latency, not local deployment.
Why Gemini 3.5 Flash Lite can handle these workloads
Gemini 3.5 Flash Lite pairs a low price and high speed with a full 1M-token context and prompt caching. It is not a maximum-reasoning model; the value is doing simple work fast and cheap at scale.
A 1M-token context on the cheapest tier
The 1,048,576-token window lets even the cheapest 3.5-class model take long documents, transcripts, or logs in one pass. Retrieval and context compaction still reduce cost, since irrelevant input still consumes tokens.
Speed and throughput over deep reasoning
With about 350 output tokens/second and minimal thinking by default, Flash Lite optimizes for fast, high-volume responses. Reserve heavier reasoning configurations and models for the requests that actually need them.
Prompt caching lowers cost further
Stable system prompts, instructions, and tool schemas create cache hits that read at the dedicated lower cache rate. Keep prefix ordering consistent so repeated high-volume calls stay cheap.
What to verify before routing production traffic to Gemini 3.5 Flash Lite
A suitable workload can still fail because the integration uses the wrong identifier, sends unsupported parameters, or assumes deep reasoning that minimal-thinking defaults do not provide. Verify the request contract before evaluating quality.
Use the exact model ID gemini-3.5-flash-lite
Send model "gemini-3.5-flash-lite" (with a dot after 3.5) on the EvoLink API route. The dashed gemini-3-5-flash-lite is only the page URL, not the API model parameter.
Use OpenAI Chat Completions or the Gemini native API
EvoLink exposes Gemini 3.5 Flash Lite through OpenAI-compatible /v1/chat/completions and the Gemini native generateContent endpoint. Keep the same EvoLink API key for either protocol.
Mind the breaking parameter changes
Custom temperature, top-K, and top-P are ignored; custom frequency and presence penalties return an error; and a request whose last turn has the model role is rejected. Update older Gemini request builders accordingly.
Validate multi-step agents under minimal thinking
Because Flash Lite defaults to minimal thinking, confirm that multi-step tool-calling agents complete their steps rather than stopping early. Reserve harder reasoning for a larger model instead of over-tuning this one.
Compare cost per accepted task, not token price alone
Gemini 3.5 Flash Lite wins when its low price and speed clear your quality bar on high-volume work. Evaluate identical task sets, and escalate the requests it cannot handle rather than paying premium rates for all traffic.
If Gemini 3.5 Flash Lite clears your quality bar on simple, high-volume tasks, its low rate and speed make it the cheapest production default. Route only the requests it fails to Gemini 3.6 Flash or a Pro-tier model.
Compare leading long-context models after workload testing
EvoLinkFirst verify whether Gemini 3.5 Flash Lite reduces retries and review effort on your tasks. Then compare price, context, caching, and workload fit to choose the production route.
| Model | Gemini 3.5 Flash Lite | Gemini 3.6 Flash | Gemini 3.1 Pro |
|---|---|---|---|
| Input / output | $0.271 / $2.25 | $1.5 / $7.5 | $1.68 / $10.08 |
| Context | 1M | 1M | 1M |
| Caching | Automatic cache reads | Context cache | Context cache |
| Best for | High-throughput classification and extraction, low-latency real-time actions, focused agent subagents, and cost-sensitive workloads where speed and price matter more than maximum reasoning. | Larger Gemini to escalate to when a task outgrows Flash-Lite’s minimal-thinking speed. | Larger Gemini to escalate to when a task outgrows Flash-Lite’s minimal-thinking speed. |
Other Gemini models on EvoLink

Gemini 3.6 Flash
The workhorse to escalate to when a task needs stronger coding or reasoning than the Lite tier provides.
View model
Gemini 3.5 Flash
The mid Flash tier between Flash-Lite and the Pro line, balancing cost and capability.
View model
Gemini 3.1 Pro
The Pro-tier step up for the deepest reasoning tasks the Flash line is not built to handle.
View model
Gemini 3.1 Flash Lite
The lightweight 3.1 route for high-volume, cost-sensitive, and low-latency workloads.
View modelOther text models

GPT-5.6
OpenAI’s tiered frontier family (Sol/Terra/Luna) for capability, latency, and cost-routing flexibility.
View model
Claude Opus 4.8
Anthropic’s premium baseline for long-running coding agents and judgment-heavy review.
View model
DeepSeek V4
Cost-sensitive open baseline for high-volume coding, reasoning, and agent workloads.
View model
Kimi K3
Moonshot’s long-context reasoning route for repository-scale coding and multi-document work.
View modelRelated reading for production teams

Gemini 3.6 Flash release date & rollout
What launched, the model ID, official pricing, and where the model is available.
Read guideGemini 3.5 Flash Lite API FAQ
Is the Gemini 3.5 Flash Lite API available through EvoLink?
Yes. Gemini 3.5 Flash Lite is available as a production model with live backend pricing and official fallback rates.
What model ID should I use for the Gemini 3.5 Flash Lite API?
Send model "gemini-3.5-flash-lite" (with a dot after 3.5). The dashed gemini-3-5-flash-lite is only the page URL, not the API model parameter.
Which protocols and SDKs work with Gemini 3.5 Flash Lite?
Use the OpenAI-compatible Chat Completions endpoint or the Gemini native generateContent endpoint with the same EvoLink API key. There is no Anthropic Messages route for this model.
Does the Gemini 3.5 Flash Lite API support 1M context or only 256K?
The EvoLink route records 1,048,576 tokens. Smaller limits usually come from a particular Gemini product surface or client configuration, not the API model itself.
How fast and how cheap is Gemini 3.5 Flash Lite?
It is Google’s fastest, most cost-effective 3.5-class model — $0.30 input / $2.50 output per 1M tokens at roughly 350 output tokens per second. That throughput per dollar is its main advantage.
What is Gemini 3.5 Flash Lite best for?
High-throughput classification, routing, JSON extraction, low-latency UI actions, and focused agent subagents. For deep coding or hard reasoning, escalate to Gemini 3.6 Flash or a Pro-tier model.
How is Gemini 3.5 Flash Lite different from 3.5 Flash and 3.6 Flash?
Flash Lite is the cheapest, fastest tier and defaults to minimal thinking. 3.5 Flash and 3.6 Flash are more capable, higher-priced workhorses for coding and agents.
Does Gemini 3.5 Flash Lite replace Gemini 2.5 Flash?
Google positions it as a replacement for Gemini 2.5 Flash and less complex Gemini 3 Flash workloads. Re-benchmark your prompts before migrating.
What breaking changes should I handle when migrating?
Custom temperature, top-K, and top-P are ignored; custom frequency and presence penalties return an error; and a request whose last turn has the model role is rejected.
Is Gemini 3.5 Flash Lite a good default for real-time or high-volume requests?
Yes — that is what it is built for. It defaults to minimal thinking for speed, so validate that multi-step agents complete their steps before a high-throughput rollout.
How should I compare Gemini 3.5 Flash Lite with GPT, Claude, GLM, or DeepSeek?
Compare it against other cheap, fast tiers on cost per 1K requests and latency, not against premium reasoning models. Escalate the requests it cannot handle to a larger model.
What should a production Gemini 3.5 Flash Lite evaluation measure?
Track cost per 1K requests, tokens per second, first-pass success, accepted deliverables, cache hits, and how often requests must escalate to a larger model.