Gemini 3.1 Flash Lite
Google · released May 7, 2026
- Type
- Closed
- Input
- $0.25
- Output
- $1.50
- Cached
- None held
List price · per 1M tokens · Google AI at 1M context · machine-readable source ↗
Our take
Written Sep 2, 2026Gemini 3.1 Flash Lite is Google's lightweight multimodal model that can handle up to one million tokens in a single request. It accepts text, images, files, audio and video, and is positioned as the fast, cheap option in the Gemini 3.1 family.
Pick this for long-document or video analysis where a million tokens of context matters, or for budget-conscious multimodal workloads on the cheapest tier. Choose the higher-throughput tier when speed matters more than cost. Skip it if you need verified quality scores — chat, reasoning, coding and multimodal performance are all unmeasured in our data — or if you need guaranteed speed on the cheapest rate.
The case for it
- One-million-token request limit among the largest we track on any model.
- Broad multimodal input: text, images, files, audio and video in a single model.
- Strong price differentiation within Google AI Studio, with a 3.6× spread on input price and throughput varying nearly 3× across tiers.
- Fastest tier at 122 tokens per second, competitive with mid-market speeds.
The case against it
- No benchmark scores at all — chat, reasoning, coding and multimodal performance are unverified.
- Throughput on the cheapest tier is modest at 76 tokens per second, half the fastest tier on the same provider.
- Vertex AI is consistently slower than AI Studio at comparable prices.
How good is it?
We hold no score for this model.
So there is no figure here for everyday use, coding, agent work or writing. That is a gap in our data, not a low score.
Where to rent it
Prices checked 2 hours ago — each listing carries its own date.
Google AI, direct
The lab is the cheapest at this context. The strip above and this offer are the same one, compared at 1M of context. 2 cheaper rows below are outside that comparison: a non-standard pricing tier.
- per 1M tokens
- $0.25 in / $1.50 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| Google AI Studioflex tierThrough OpenRouter | $0.13 / $0.75checked 2 hours ago | 1M66K max reply | 56 tok/s | No | Yes55 days | Unknown |
| Google Vertex AIflex tierglobalThrough OpenRouter | $0.13 / $0.75checked 2 hours ago | 1M66K max reply | 2 tok/s | No | No | Confirmed |
| Google AIDirect | $0.25 / $1.50checked 2 hours ago | 1M66K max reply | not measured | Unknown | Unknown | Unknown |
| Google Vertex AIglobalThrough OpenRouter | $0.25 / $1.50checked 2 hours ago | 1M66K max reply | 64 tok/s | No | No | Confirmed |
| OpenRouterOpenRouter's own listing | $0.25 / $1.50checked 2 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| DeepInfraDirect | $0.25 / $1.50checked 2 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Google AI StudioThrough OpenRouter | $0.25 / $1.50checked 2 hours ago | 1M66K max reply | 76 tok/s | No | Yes55 days | Unknown |
| Google Vertex AIusThrough OpenRouter | $0.28 / $1.65checked 2 hours ago | 1M66K max reply | 109 tok/s | No | No | Confirmed |
| Google Vertex AIeuThrough OpenRouter | $0.28 / $1.65checked 2 hours ago | 1M66K max reply | 106 tok/s | No | No | Confirmed |
| Google AI Studiopriority tierThrough OpenRouter | $0.45 / $2.70checked 2 hours ago | 1M66K max reply | 55 tok/s | No | Yes55 days | Unknown |
| Google Vertex AIpriority tierglobalThrough OpenRouter | $0.45 / $2.70checked 2 hours ago | 1M66K max reply | 95 tok/s | No | No | Confirmed |
Across the 11 listings we hold: 8 say they do not train on prompts, 0 say they do and 3 do not say. 5 appear in the zero-retention registry we check; the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| Google AI StudioflexThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIflex · globalThrough OpenRouter | ✓ | ✓ | ✓ |
| Google AIDirect | |||
| Google Vertex AIglobalThrough OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| DeepInfraDirect | |||
| Google AI StudioThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIusThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIeuThrough OpenRouter | ✓ | ✓ | ✓ |
| Google AI StudiopriorityThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIpriority · globalThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 9 of 11 listings say yes, 2 publish no parameter list. JSON output: 9 of 11 listings say yes, 2 publish no parameter list. Strict schema: 9 of 11 listings say yes, 2 publish no parameter list.
Models people weigh against Gemini 3.1 Flash Lite
When we formed this view
Recent changes
What moved
input $0.25 → $0.125, output $1.5 → $0.75 per 1M tokensWhat moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- No independent board has scored it, so we hold no quality figures at all.
- 2 of 11 listings publish no parameter list, so what their API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 3 of 11 listings do not say whether they train on prompts.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Text, images, audio, video and documents in, text out
- Catalogue slug
- google-gemini-3-1-flash-lite