Gemini 3.5 Flash Lite
Google · released Jul 21, 2026
- Type
- Closed
- Input
- $0.30
- Output
- $2.50
- Cached
- None held
List price · per 1M tokens · Google AI at 1M context · machine-readable source ↗
Our take
Written Sep 2, 2026Gemini 3.5 Flash Lite is a hosted-only multimodal model from Google that accepts text, images, files, audio and video, and can handle up to one million tokens in a single request. Released in July 2026, it is positioned as the lighter sibling to the full Flash line, with measured strengths in coding and a very large context window for its price class.
Pick this for long-context document and media analysis where you need broad input support, or for budget-conscious coding tasks where its measured coding scores lead its other skills. Use it when proprietary hosting is acceptable and you can navigate tier selection carefully. Skip it if you need agentic or autonomous task performance, predictable throughput, or if you want to self-host.
The case for it
- One-million-token request limit, extremely large for its price class.
- Coding is the standout skill: LiveBench Coding 76.07%, well above its overall score, and Arena Coding 1499.86 leads its own profile by 43 points.
- Lowest tier is half the price of the next one up on the same model.
The case against it
- Agentic performance is weak across every measured dimension, with negative scores on task outcome, recovery and steerability.
- LiveBench Data Analysis 53.25% and Agentic Coding 45.25% are particular soft spots, the latter nearly 19 points below its overall score.
- Throughput varies wildly from 17 to 108 tokens per second depending on tier, with no way to predict which you will get.
How good is it?
A closed text model for everyday questions and drafting prose, though it struggles with multi-step agent work and changing course.
- getting answers to everyday questionsArena Text (overall) · 42nd of 168
- drafts, rewrites and editingArena Creative Writing · 39th of 168
- multi-step work it carries out for youArena Agent · 52nd of 55
- calling tools to carry out requestsArena Agent · Tool use · 49th of 55
- changing course when you give new instructionsArena Agent · Steerability · 55th of 55
- getting back on track after a step failsArena Agent · Recovery · 51st of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)42nd of 168 · 1456
CodingWriting and fixing code on its own
Arena Coding46th of 168 · 1503
AgenticPlanning, calling tools, staying on task
Arena Agent52nd of 55 · −0.153
WritingDrafting and rewriting prose
Arena Creative Writing39th of 168 · 1435
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to rent it
Prices checked between 59 min and 1 hour ago — each listing carries its own date.
Google AI, direct
The lab is the cheapest at this context. The strip above and this offer are the same one, compared at 1M of context. 2 cheaper rows below are outside that comparison: a non-standard pricing tier.
- per 1M tokens
- $0.30 in / $2.50 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| Google AI Studioflex tierThrough OpenRouter | $0.15 / $1.25checked 1 hour ago | 1M66K max reply | 23 tok/s | No | Yes55 days | Unknown |
| Google Vertex AIflex tierglobalThrough OpenRouter | $0.15 / $1.25checked 1 hour ago | 1M66K max reply | 20 tok/s | No | No | Confirmed |
| Google Vertex AIglobalThrough OpenRouter | $0.30 / $2.50checked 1 hour ago | 1M66K max reply | 66 tok/s | No | No | Confirmed |
| Google AIDirect | $0.30 / $2.50checked 59 min ago | 1M66K max reply | not measured | Unknown | Unknown | Unknown |
| Google AI StudioThrough OpenRouter | $0.30 / $2.50checked 1 hour ago | 1M66K max reply | 87 tok/s | No | Yes55 days | Unknown |
| OpenRouterOpenRouter's own listing | $0.30 / $2.50checked 1 hour ago | 1M | not measured | Unknown | Unknown | Unknown |
| Google Vertex AIusThrough OpenRouter | $0.33 / $2.75checked 1 hour ago | 1M66K max reply | 117 tok/s | No | No | Confirmed |
| Google Vertex AIeuThrough OpenRouter | $0.33 / $2.75checked 1 hour ago | 1M66K max reply | 70 tok/s | No | No | Confirmed |
| Google AI Studiopriority tierThrough OpenRouter | $0.54 / $4.50checked 1 hour ago | 1M66K max reply | 63 tok/s | No | Yes55 days | Unknown |
| Google Vertex AIpriority tierglobalThrough OpenRouter | $0.54 / $4.50checked 1 hour ago | 1M66K max reply | 27 tok/s | No | No | Confirmed |
Across the 10 listings we hold: 8 say they do not train on prompts, 0 say they do and 2 do not say. 5 appear in the zero-retention registry we check; the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| Google AI StudioflexThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIflex · globalThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIglobalThrough OpenRouter | ✓ | ✓ | ✓ |
| Google AIDirect | |||
| Google AI StudioThrough OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Google Vertex AIusThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIeuThrough OpenRouter | ✓ | ✓ | ✓ |
| Google AI StudiopriorityThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIpriority · globalThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 9 of 10 listings say yes, 1 publishes no parameter list. JSON output: 9 of 10 listings say yes, 1 publishes no parameter list. Strict schema: 9 of 10 listings say yes, 1 publishes no parameter list.
Models people weigh against Gemini 3.5 Flash Lite
When we formed this view
Recent changes
Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- 1 of 10 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 10 listings do not say whether they train on prompts.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Text, images, audio, video and documents in, text out
- Catalogue slug
- google-gemini-3-5-flash-lite