Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints.

Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.

google/gemini-3.1-flash-lite
Organization
Google
Providers
22
Context
1,048,576
Output limit
65,536
Knowledge
2025-01
Release
2026-05-07
Updated
2026-05-07
Weights
Closed
Input
text, image, video, audio, pdf
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

34 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
302.AI
gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Abacus
gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Cortecs
gemini-3.1-flash-lite
1,048,57665,535$0.272 / $1.631YesYesYes
View
DevPass (LLM Gateway)
gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Eden AI
vertex/gemini-3.1-flash-lite@eu
1,048,57665,536$0.25 / $1.50YesYesYes
View
Eden AI
vertex/gemini-3.1-flash-lite@us
1,048,57665,536$0.25 / $1.50YesYesYes
View
Eden AI
vertex/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Eden AI
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Impossibl
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Kenari
gemini-3-1-flash-lite
1,048,57665,536$0.00 / $0.00YesYesYes
View
Kilo Gateway
google/gemini-3.1-flash-lite
1,048,57665,536$0.125 / $0.75YesYesYes
View
LLM Gateway
google-ai-studio/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
LLM Gateway
google-vertex/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Merge Gateway
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
NanoGPT
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
NEAR AI Cloud
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Ofox
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Googlevia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.125 / $0.75YesYesYes
View
Googlevia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.275 / $1.65YesYesYes
View
Googlevia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.25 / $1.50YesYesYes
View
Googlevia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.275 / $1.65YesYesYes
View
Googlevia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.45 / $2.70YesYesYes
View
Google AI Studiovia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.125 / $0.75YesYesYes
View
Google AI Studiovia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.45 / $2.70YesYesYes
View
Google AI Studiovia OpenRouter
google/gemini-3.1-flash-lite
unknown
1,048,57665,536$0.25 / $1.50YesYesYes
View
OrcaRouter
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Pioneer
gemini-3.1-flash-lite
1,000,00065,000$0.25 / $1.50YesYesYes
View
Requesty
gemini-3.1-flash-lite
1,048,57665,535$0.25 / $1.50YesYesYes
View
Requesty
gemini-3.1-flash-lite@eu
1,048,57665,535$0.275 / $1.65YesYesYes
View
SAP AI Core
gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Tempr Gateway
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
Vercel AI Gateway
google/gemini-3.1-flash-lite
1,000,00065,000$0.25 / $1.50YesYesYes
View
Vertex
gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View
ZenMux
google/gemini-3.1-flash-lite
1,048,57665,536$0.25 / $1.50YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers34Across 22 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 8 lowest-priced configurations
InputOutput
Googlegoogle-vertex/global/flex
$0.125$0.75
Google AI Studiogoogle-ai-studio/flex
$0.125$0.75
Googlegoogle-vertex/global
$0.25$1.50
Google AI Studiogoogle-ai-studio
$0.25$1.50
Googlegoogle-vertex/us
$0.275$1.65
Googlegoogle-vertex/eu
$0.275$1.65
Googlegoogle-vertex/global/priority
$0.45$2.70
Google AI Studiogoogle-ai-studio/priority
$0.45$2.70
google/gemini-3.1-flash-liteObserved
Input
$0.25
Output
$1.50
Cache read
$0.025
Cache write
$0.083333
Additional fees and pricing conditions
audio
$0.0000005 / audio token
image
$0.00000025 / image
web search
$0.014 / search
input audio cache
$0.00000005 / audio token
internal reasoning
$0.0000015 / reasoning token
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
Googlegoogle-vertex/global/flex$0.125$0.75$0.0125$0.041667
Googlegoogle-vertex/us$0.275$1.65$0.0275$0.083333
Googlegoogle-vertex/global$0.25$1.50$0.025$0.083333
Googlegoogle-vertex/eu$0.275$1.65$0.0275$0.083333
Googlegoogle-vertex/global/priority$0.45$2.70$0.045$0.15
Google AI Studiogoogle-ai-studio/flex$0.125$0.75$0.0125$0.041667
Google AI Studiogoogle-ai-studio/priority$0.45$2.70$0.045$0.15
Google AI Studiogoogle-ai-studio$0.25$1.50$0.025$0.083333

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Googlegoogle-vertex/global/flex
16,022–134,475.41 ms
Googlegoogle-vertex/us
1,195–2,553.62 ms
Googlegoogle-vertex/global
786–4,066.62 ms
Googlegoogle-vertex/eu
479–3,510.11 ms
Googlegoogle-vertex/global/priority
1,245–2,162.04 ms
Google AI Studiogoogle-ai-studio/flex
413–2,514.71 ms
All latency percentiles 8 endpoints
EndpointP50P75P90P99
Googlegoogle-vertex/global/flex16,022 ms16,961.25 ms20,423.3 ms134,475.41 ms
Googlegoogle-vertex/us1,195 ms1,406 ms2,244.2 ms2,553.62 ms
Googlegoogle-vertex/global786 ms1,301 ms2,057.8 ms4,066.62 ms
Googlegoogle-vertex/eu479 ms615 ms1,448.1 ms3,510.11 ms
Googlegoogle-vertex/global/priority1,245 ms1,358 ms1,534.4 ms2,162.04 ms
Google AI Studiogoogle-ai-studio/flex413 ms580.25 ms716.9 ms2,514.71 ms
Google AI Studiogoogle-ai-studio/priority507.5 ms677.25 ms1,109.5 ms1,659.93 ms
Google AI Studiogoogle-ai-studio605.5 ms974 ms1,296 ms3,317.9 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Googlegoogle-vertex/global/flex
5–35.23 t/s
Googlegoogle-vertex/us
264–317.32 t/s
Googlegoogle-vertex/global
92–301 t/s
Googlegoogle-vertex/eu
129–325.82 t/s
Googlegoogle-vertex/global/priority
97–265.02 t/s
Google AI Studiogoogle-ai-studio/flex
90–263 t/s
All throughput percentiles 8 endpoints
EndpointP50P75P90P99
Googlegoogle-vertex/global/flex5 t/s7 t/s12 t/s35.23 t/s
Googlegoogle-vertex/us264 t/s278.5 t/s302.2 t/s317.32 t/s
Googlegoogle-vertex/global92 t/s140 t/s191 t/s301 t/s
Googlegoogle-vertex/eu129 t/s152 t/s177 t/s325.82 t/s
Googlegoogle-vertex/global/priority97 t/s117 t/s157 t/s265.02 t/s
Google AI Studiogoogle-ai-studio/flex90 t/s103 t/s128 t/s263 t/s
Google AI Studiogoogle-ai-studio/priority183 t/s313.75 t/s451.8 t/s479.89 t/s
Google AI Studiogoogle-ai-studio149 t/s195 t/s245 t/s322.09 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 8 endpoints
Endpoint5 minutes30 minutes24 hours
Googlegoogle-vertex/global/flex
Googlegoogle-vertex/us
Googlegoogle-vertex/global
Googlegoogle-vertex/eu
Googlegoogle-vertex/global/priority
Google AI Studiogoogle-ai-studio/flex
Google AI Studiogoogle-ai-studio/priority
Google AI Studiogoogle-ai-studio

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
tau bench verified airlineGoogle: Gemini 3.1 Flash Lite
72%
gpqa diamondGoogle: Gemini 3.1 Flash Lite
81.594%
EvaluationConfigurationSourceResultsDetails
tau bench verified airlineGoogle: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite-20260507OpenRouter72%
Details
gpqa diamondGoogle: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite-20260507OpenRouter81.594%
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentgoogle/gemini-3.1-flash-lite-202605072–9 turns
$0.006242
Details
Hermes Agentgoogle/gemini-3.1-flash-lite-2026050710–49 turns
$0.03274
Details
Hermes Agentgoogle/gemini-3.1-flash-lite-202605071 turn
$0.001033
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage6,272,384,505,725 reported tokens · 89 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 89 days
Date (UTC)Reported tokensReported routes
Oct 3, 202659,272,111,1081
Oct 2, 202669,290,882,4341
Oct 1, 202673,675,351,1121
Sep 30, 202680,410,172,9721
Sep 29, 202685,443,010,1991
Sep 28, 202681,560,155,5901
Sep 27, 202653,404,851,6271
Sep 26, 202657,976,742,3711
Sep 25, 202683,867,637,4021
Sep 24, 202686,358,795,3501
Sep 23, 202686,680,884,7911
Sep 22, 202681,580,602,4121
Sep 21, 202678,438,111,8271
Sep 20, 202652,598,058,3641
Sep 19, 202650,812,222,8361
Sep 18, 202677,764,017,3681
Sep 17, 202681,004,414,4471
Sep 16, 202678,634,508,8891
Sep 15, 202677,176,920,6881
Sep 14, 202676,713,260,6721
Sep 13, 202649,484,885,6921
Sep 12, 202651,910,169,3971
Sep 11, 202665,368,249,4641
Sep 10, 202673,499,374,4831
Sep 9, 202671,718,093,9561
Sep 8, 202689,362,465,4851
Sep 7, 202660,480,836,1581
Sep 6, 202649,007,674,3041
Sep 5, 202649,941,160,6471
Sep 4, 202665,042,924,4141
Sep 3, 202670,335,558,6001
Sep 2, 202674,224,251,6141
Sep 1, 202670,537,972,3961
Aug 31, 202665,556,094,6821
Aug 30, 202646,221,382,1581
Aug 29, 202649,653,009,9811
Aug 28, 202664,129,595,2811
Aug 27, 202675,768,610,3801
Aug 26, 202673,359,400,8271
Aug 25, 202676,152,424,7521
Aug 24, 202674,344,855,2471
Aug 23, 202644,670,107,2651
Aug 22, 202646,544,665,1611
Aug 21, 202668,583,110,7331
Aug 20, 202665,780,568,5901
Aug 19, 202673,778,771,6611
Aug 18, 202672,235,243,2281
Aug 17, 202664,099,739,3741
Aug 16, 202639,769,813,1651
Aug 15, 202643,402,618,8521
Aug 14, 202667,625,814,7251
Aug 13, 202668,402,228,1221
Aug 12, 202666,420,565,7501
Aug 11, 202669,476,642,1141
Aug 10, 202662,925,370,3451
Aug 9, 202640,631,900,2141
Aug 8, 202651,235,471,1781
Aug 7, 202666,056,387,6111
Aug 6, 202670,424,964,5651
Aug 5, 202675,563,429,9091
Aug 4, 202680,015,317,6451
Aug 3, 202681,369,729,1051
Aug 2, 202652,970,948,7971
Aug 1, 202649,802,148,9061
Jul 31, 202686,594,183,7511
Jul 30, 202692,116,123,8761
Jul 29, 202692,108,209,5901
Jul 28, 202691,235,661,9031
Jul 27, 202692,005,253,3931
Jul 26, 202660,006,540,2681
Jul 25, 202658,725,629,9131
Jul 24, 202685,766,440,7851
Jul 23, 202698,665,692,3081
Jul 22, 202695,513,594,4131
Jul 21, 202696,121,417,4611
Jul 20, 202690,811,739,5661
Jul 19, 202660,989,199,1321
Jul 18, 202658,942,276,0871
Jul 17, 202683,273,200,2081
Jul 16, 202686,120,832,1891
Jul 15, 202685,560,811,4781
Jul 14, 202683,116,619,7881
Jul 13, 202688,388,996,4571
Jul 12, 202657,415,462,5051
Jul 11, 202657,725,305,2241
Jul 10, 202674,256,294,3791
Jul 9, 202675,410,323,0341
Jul 8, 202682,259,391,3951
Jul 7, 202678,708,047,2301

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
google/gemini-3.1-flash-liteObserved
Context
1,048,576
Maximum output
65,536
Tokenizer
Gemini
Input
text, image, video, file, audio
Output
text
Moderation
No
Added to OpenRouter
2026-05-07
Base URLhttps://openrouter.ai/api/v1
Model IDgoogle/gemini-3.1-flash-lite
Version IDgoogle/gemini-3.1-flash-lite-20260507

Reasoning

Required
No
Enabled by default
Yes
Default effort
minimal
Supported efforts
high, medium, low, minimal

Supported parameters

  • include_reasoning
  • max_tokens
  • reasoning
  • reasoning_effort
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_p

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Gemini 3.1 Flash Lite cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Gemini 3.1 Flash Lite?

There are 34 cataloged offers across 22 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 1,048,576 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · 2a8394b713e4
  • · 29104d169e56