Kimi K3 is now availableExplore Kimi K3

Gemini 3.5 Flash Lite API

Google-Text generation-from $0.271 / 1M input tokens$0.300 official price-Available
1M context65K default outputPrompt cachingChat + Gemini API
API docs
Production routeLive
Provider
Google
Model
Flash Lite
Context window
1,048,576 tokens
Protocols
Chat + Gemini

Choose Gemini 3.5 Flash Lite

Google’s cheapest, fastest 3.5-class model for low-latency classification, routing, extraction, and lightweight agent subagents, with a full 1M-token context at $0.30 / $2.50 per 1M tokens.

Gemini 3.5 Flash Lite

Google fastest, most cost-effective 3.5-class model

Selected
From $0.271 / 1M input tokens$0.300 official pricegemini-3.5-flash-lite
Best for

High-throughput classification and extraction, low-latency real-time actions, focused agent subagents, and cost-sensitive workloads where speed and price matter more than maximum reasoning.

Input
$0.271 / 1M-10%
18.4 cr / 1M$0.300official price
Cache read
$0.028 / 1M-7%
1.9 cr / 1M$0.030official price
Output
$2.250 / 1M-10%
153 cr / 1M$2.500official price

Gemini 3.5 Flash Lite pricing

Estimate a request with the interactive pricing calculator. All user groups use the official Gemini 3.5 Flash Lite rate.

Flash Lite

Token calculator

Enter the token mix for one request.

Estimated request cost

Gemini 3.5 Flash Lite
Official rate
USD$0.0010
Credits0.0647
Input tokens0.0184 cr
Cache read tokens0.0004 cr
Output tokens0.0459 cr

Minimum charge: 0.01 credits per request.

EvoLink vs Google direct

Same token mix, default group price.
-10%
EvoLink estimate$0.0010
Google official$0.0011
You save$0.0002

Budget guide

Approximate requests using the current token mix.
Add credits
$10
About 10510 requests

For quick testing

$50
About 52550 requests

For regular development

$100
About 105100 requests

For production evaluation

Model pricing

Gemini 3.5 Flash Lite

All context sizes
Input tokens
$0.271 / 1M-10%
18.4 cr / 1M$0.300official price
Cache read tokens
$0.028 / 1M-7%
1.9 cr / 1M$0.030official price
Output tokens
$2.250 / 1M-10%
153 cr / 1M$2.500official price

USD and credits are shown per 1M tokens. Live backend pricing takes priority over these frozen fallback rates.

Gemini 3.5 Flash Lite API for fast, low-cost, high-throughput workloads

Call Google’s cheapest, fastest 3.5-class model through EvoLink’s unified API. Gemini 3.5 Flash Lite is built for low-latency classification, routing, extraction, and lightweight agent subagents — with a 1,048,576-token context window, prompt caching, structured output, and tool use — at $0.30 / $2.50 per million input / output tokens.

Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite use cases

Where Gemini 3.5 Flash Lite earns a place in a production model stack

Google positions Gemini 3.5 Flash Lite as its fastest, most cost-effective 3.5-class model. Its strongest fit is high-volume, latency-sensitive work where low price and speed matter more than maximum reasoning — and where you can escalate the few hard requests to a bigger model.

High-throughput classification and extraction

Gemini 3.5 Flash Lite is built for high-throughput classification, routing, and JSON extraction. At $0.30 / $2.50 per 1M tokens it lets you process large request volumes cheaply where a premium model would be wasted.

Low-latency, cost-sensitive scale

Google reports about 350 output tokens per second, making it a strong fit for real-time UI actions, autocomplete, moderation, and other latency-critical paths. Measure cost and speed per request, not headline benchmarks.

Lightweight and focused agent subagents

It suits subagents that execute focused tasks inside larger multi-agent systems. It defaults to minimal thinking for speed; for multi-step tool use, validate that agents do not terminate early before you scale.

When to escalate to a bigger model

Deep repository refactoring, long autonomous coding, and hard multi-step reasoning belong on Gemini 3.6 Flash or a Pro-tier model. Use Flash Lite for the high-volume majority and route only the difficult minority upward.

Launch-day signals

What early Gemini 3.5 Flash Lite reactions suggest—and what still needs proof

Gemini 3.5 Flash Lite launched on 2026-07-21 as a replacement for Gemini 2.5 Flash. Treat launch-day reactions as hypotheses to verify on your own tasks, tools, budgets, and acceptance criteria.

Price-to-performance is the headline

At $0.30 input / $2.50 output per 1M tokens and roughly 350 tokens/second, its appeal is throughput per dollar. On simple, high-volume work this can beat pricier models on total cost with acceptable quality.

It replaces Gemini 2.5 Flash and lighter 3-Flash workloads

Google frames it as a drop-in for Gemini 2.5 Flash and less complex Gemini 3 Flash tasks. If you run those today, re-benchmark your prompts before migrating so quality and formatting stay within tolerance.

Minimal thinking is the default—validate multi-step agents

Flash Lite defaults to minimal thinking for speed and cost. Google notes that minimal thinking can cause premature tool termination on multi-step tasks, so test agent completion before high-throughput rollout.

It is a closed, API-only Google model

Gemini 3.5 Flash Lite has no open weights and cannot be self-hosted. Access is through the Gemini API and platforms such as Google AI Studio, plus unified gateways like EvoLink—route by cost and latency, not local deployment.

Core capabilities

Why Gemini 3.5 Flash Lite can handle these workloads

Gemini 3.5 Flash Lite pairs a low price and high speed with a full 1M-token context and prompt caching. It is not a maximum-reasoning model; the value is doing simple work fast and cheap at scale.

A 1M-token context on the cheapest tier

The 1,048,576-token window lets even the cheapest 3.5-class model take long documents, transcripts, or logs in one pass. Retrieval and context compaction still reduce cost, since irrelevant input still consumes tokens.

Speed and throughput over deep reasoning

With about 350 output tokens/second and minimal thinking by default, Flash Lite optimizes for fast, high-volume responses. Reserve heavier reasoning configurations and models for the requests that actually need them.

Prompt caching lowers cost further

Stable system prompts, instructions, and tool schemas create cache hits that read at the dedicated lower cache rate. Keep prefix ordering consistent so repeated high-volume calls stay cheap.

Gemini 3.5 Flash Lite API production checks

What to verify before routing production traffic to Gemini 3.5 Flash Lite

A suitable workload can still fail because the integration uses the wrong identifier, sends unsupported parameters, or assumes deep reasoning that minimal-thinking defaults do not provide. Verify the request contract before evaluating quality.

01

Use the exact model ID gemini-3.5-flash-lite

Send model "gemini-3.5-flash-lite" (with a dot after 3.5) on the EvoLink API route. The dashed gemini-3-5-flash-lite is only the page URL, not the API model parameter.

Model ID
02

Use OpenAI Chat Completions or the Gemini native API

EvoLink exposes Gemini 3.5 Flash Lite through OpenAI-compatible /v1/chat/completions and the Gemini native generateContent endpoint. Keep the same EvoLink API key for either protocol.

Protocols
03

Mind the breaking parameter changes

Custom temperature, top-K, and top-P are ignored; custom frequency and presence penalties return an error; and a request whose last turn has the model role is rejected. Update older Gemini request builders accordingly.

Migration
04

Validate multi-step agents under minimal thinking

Because Flash Lite defaults to minimal thinking, confirm that multi-step tool-calling agents complete their steps rather than stopping early. Reserve harder reasoning for a larger model instead of over-tuning this one.

Agents
Production economics

Compare cost per accepted task, not token price alone

Gemini 3.5 Flash Lite wins when its low price and speed clear your quality bar on high-volume work. Evaluate identical task sets, and escalate the requests it cannot handle rather than paying premium rates for all traffic.

First-pass successAccepted deliverable rateRetry countOutput tokensCache-hit rateTokens per secondCost per 1K requestsEscalation rate to a larger model

If Gemini 3.5 Flash Lite clears your quality bar on simple, high-volume tasks, its low rate and speed make it the cheapest production default. Route only the requests it fails to Gemini 3.6 Flash or a Pro-tier model.

Compare leading long-context models after workload testing

EvoLink

First verify whether Gemini 3.5 Flash Lite reduces retries and review effort on your tasks. Then compare price, context, caching, and workload fit to choose the production route.

Gemini 3.5 Flash Lite
Input / output$0.271 / $2.25
Context1M
CachingAutomatic cache reads
Best forHigh-throughput classification and extraction, low-latency real-time actions, focused agent subagents, and cost-sensitive workloads where speed and price matter more than maximum reasoning.
Gemini 3.6 Flash
Input / output$1.5 / $7.5
Context1M
CachingContext cache
Best forLarger Gemini to escalate to when a task outgrows Flash-Lite’s minimal-thinking speed.
Gemini 3.1 Pro
Input / output$1.68 / $10.08
Context1M
CachingContext cache
Best forLarger Gemini to escalate to when a task outgrows Flash-Lite’s minimal-thinking speed.

Other Gemini models on EvoLink

Gemini 3.6 Flash

Gemini 3.6 Flash

The workhorse to escalate to when a task needs stronger coding or reasoning than the Lite tier provides.

View model
Gemini 3.5 Flash

Gemini 3.5 Flash

The mid Flash tier between Flash-Lite and the Pro line, balancing cost and capability.

View model
Gemini 3.1 Pro

Gemini 3.1 Pro

The Pro-tier step up for the deepest reasoning tasks the Flash line is not built to handle.

View model
Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite

The lightweight 3.1 route for high-volume, cost-sensitive, and low-latency workloads.

View model

Other text models

GPT-5.6

GPT-5.6

OpenAI’s tiered frontier family (Sol/Terra/Luna) for capability, latency, and cost-routing flexibility.

View model
Claude Opus 4.8

Claude Opus 4.8

Anthropic’s premium baseline for long-running coding agents and judgment-heavy review.

View model
DeepSeek V4

DeepSeek V4

Cost-sensitive open baseline for high-volume coding, reasoning, and agent workloads.

View model
Kimi K3

Kimi K3

Moonshot’s long-context reasoning route for repository-scale coding and multi-document work.

View model

Related reading for production teams

Gemini 3.6 Flash release date & rollout

Gemini 3.6 Flash release date & rollout

What launched, the model ID, official pricing, and where the model is available.

Read guide

Gemini 3.5 Flash Lite API FAQ

Is the Gemini 3.5 Flash Lite API available through EvoLink?

Yes. Gemini 3.5 Flash Lite is available as a production model with live backend pricing and official fallback rates.

What model ID should I use for the Gemini 3.5 Flash Lite API?

Send model "gemini-3.5-flash-lite" (with a dot after 3.5). The dashed gemini-3-5-flash-lite is only the page URL, not the API model parameter.

Which protocols and SDKs work with Gemini 3.5 Flash Lite?

Use the OpenAI-compatible Chat Completions endpoint or the Gemini native generateContent endpoint with the same EvoLink API key. There is no Anthropic Messages route for this model.

Does the Gemini 3.5 Flash Lite API support 1M context or only 256K?

The EvoLink route records 1,048,576 tokens. Smaller limits usually come from a particular Gemini product surface or client configuration, not the API model itself.

How fast and how cheap is Gemini 3.5 Flash Lite?

It is Google’s fastest, most cost-effective 3.5-class model — $0.30 input / $2.50 output per 1M tokens at roughly 350 output tokens per second. That throughput per dollar is its main advantage.

What is Gemini 3.5 Flash Lite best for?

High-throughput classification, routing, JSON extraction, low-latency UI actions, and focused agent subagents. For deep coding or hard reasoning, escalate to Gemini 3.6 Flash or a Pro-tier model.

How is Gemini 3.5 Flash Lite different from 3.5 Flash and 3.6 Flash?

Flash Lite is the cheapest, fastest tier and defaults to minimal thinking. 3.5 Flash and 3.6 Flash are more capable, higher-priced workhorses for coding and agents.

Does Gemini 3.5 Flash Lite replace Gemini 2.5 Flash?

Google positions it as a replacement for Gemini 2.5 Flash and less complex Gemini 3 Flash workloads. Re-benchmark your prompts before migrating.

What breaking changes should I handle when migrating?

Custom temperature, top-K, and top-P are ignored; custom frequency and presence penalties return an error; and a request whose last turn has the model role is rejected.

Is Gemini 3.5 Flash Lite a good default for real-time or high-volume requests?

Yes — that is what it is built for. It defaults to minimal thinking for speed, so validate that multi-step agents complete their steps before a high-throughput rollout.

How should I compare Gemini 3.5 Flash Lite with GPT, Claude, GLM, or DeepSeek?

Compare it against other cheap, fast tiers on cost per 1K requests and latency, not against premium reasoning models. Escalate the requests it cannot handle to a larger model.

What should a production Gemini 3.5 Flash Lite evaluation measure?

Track cost per 1K requests, tokens per second, first-pass success, accepted deliverables, cache hits, and how often requests must escalate to a larger model.