Gemma 4 31B IT

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license.

google/gemma-4-31b-it
Organization
Google
Family
gemma
Providers
37
Context
262,144
Output limit
32,768
Knowledge
—
Release
2026-04-02
Updated
2026-04-02
Weights
Open
Input
text, image
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

58 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
Abacus
google/gemma-4-31b-it
262,144131,072$0.14 / $0.40YesYesYes
View
ai&
google/gemma-4-31b-it
262,14432,768$0.20 / $0.50YesYesYes
View
Amazon Bedrock
google.gemma-4-31b
262,14432,768$0.14 / $0.40YesYesYes
View
Bothub
gemma-4-31b-it:free
262,14432,768$0.00 / $0.00YesYesYes
View
Chutes
google/gemma-4-31B-turbo-TEE
131,07265,536$0.12 / $0.37YesYesYes
View
CoreWeave
google/gemma-4-31B-it
262,144262,144$0.10 / $0.34YesYesYes
View
Cortecs
gemma-4-31b-it
262,000262,000$0.223 / $0.39YesYesYes
View
CrofAI
gemma-4-31b-it
262,144262,144$0.10 / $0.30YesYesYes
View
Crusoe
google/gemma-4-31b-it
262,14432,768$0.14 / $0.40YesYesYes
View
Deep Infra
google/gemma-4-31B-it
262,14432,768$0.15 / $0.40YesYesYes
View
DevPass (LLM Gateway)
gemma-4-31b-it
262,14432,768$0.10 / $0.25YesYesYes
View
FastRouter
google/gemma-4-31b-it
262,14432,768$0.13 / $0.38YesYesYes
View
Friendli
google/gemma-4-31B-it
262,14432,768$0.14 / $0.40YesYesYes
View
Hugging Face
google/gemma-4-31B-it
262,14432,768$0.14 / $0.40YesYesYes
View
InferX
gemma-4-31B-it-fp8
262,14432,768$0.00 / $0.00YesYesYes
View
Infomaniak
google/gemma-4-31B-it
100,00032,768$0.25 / $0.50YesYesYes
View
Kenari
gemma-4-31b-it
262,14432,768$0.00 / $0.00YesYesYes
View
Kilo Gateway
google/gemma-4-31b-it
262,14416,384$0.09 / $0.34YesYesYes
View
Lilac
google/gemma-4-31b-it
262,100262,100$0.11 / $0.35YesYesYes
View
LLM Gateway
scx-ai/gemma-4-31b-it
131,0728,192$0.30 / $0.91NoYesNo
View
LLM Gateway
deepinfra/gemma-4-31b-it
262,14432,768$0.15 / $0.40YesYesNo
View
LLM Gateway
consensusprotocol/gemma-4-31b-it
262,14432,768$0.10 / $0.25YesYesNo
View
LLM Gateway
novita/gemma-4-31b-it
262,14432,768$0.14 / $0.40NoYesNo
View
LLM Gateway
runware/gemma-4-31b-it
262,14465,536$0.102 / $0.297YesYesNo
View
LLM Gateway
cerebras/gemma-4-31b-it
131,07232,768$0.99 / $1.49YesYesNo
View
LLMTR
gemma-4
131,072131,072$2.00 / $5.00YesYesYes
View
Merge Gateway
google/gemma-4-31b-it
262,14465,536$0.14 / $0.40YesNoNo
View
NanoGPT
google/gemma-4-31b-it:thinking
262,144131,072$0.10 / $0.35YesYesYes
View
NanoGPT
google/gemma-4-31b-it
262,144131,072$0.10 / $0.45YesYesYes
View
NanoGPT
TEE/gemma-4-31b-it
262,144262,144$0.15 / $0.46YesYesNo
View
Neuralwatt
gemma-4-31b
262,12816,384$0.144 / $0.42YesYesYes
View
Chutesvia OpenRouter
google/gemma-4-31b-it
fp4
131,07265,536$0.12 / $0.37YesYesYes
View
CoreWeavevia OpenRouter
google/gemma-4-31b-it
fp4
262,144235,929$0.10 / $0.34YesYesYes
View
Crusoevia OpenRouter
google/gemma-4-31b-it
bf16
262,144262,141$0.14 / $0.40YesYesYes
View
DeepInfravia OpenRouter
google/gemma-4-31b-it
fp4
262,14416,384$0.09 / $0.34YesNoYes
View
DeepInfravia OpenRouter
google/gemma-4-31b-it
fp8
262,14416,384$0.15 / $0.40YesYesYes
View
DekaLLMvia OpenRouter
google/gemma-4-31b-it
unknown
262,144235,929$0.10 / $0.33YesYesYes
View
Friendlivia OpenRouter
google/gemma-4-31b-it
unknown
262,1448,192$0.14 / $0.40YesYesYes
View
Google AI Studiovia OpenRouter
google/gemma-4-31b-it:free
unknown
262,14432,768$0.00 / $0.00YesYesNo
View
Io Netvia OpenRouter
google/gemma-4-31b-it
unknown
262,14416,384$0.38 / $1.15YesYesNo
View
ModelRunvia OpenRouter
google/gemma-4-31b-it
fp4
262,144235,929$0.75 / $1.00YesYesYes
View
Novitavia OpenRouter
google/gemma-4-31b-it
bf16
262,144131,072$0.14 / $0.40YesYesYes
View
Parasailvia OpenRouter
google/gemma-4-31b-it
fp8
262,144235,929$0.15 / $0.40YesYesYes
View
SambaNovavia OpenRouter
google/gemma-4-31b-it
unknown
131,072117,964$0.38 / $1.15YesYesNo
View
SiliconFlowvia OpenRouter
google/gemma-4-31b-it
fp8
262,144235,929$0.75 / $1.00YesYesYes
View
Venicevia OpenRouter
google/gemma-4-31b-it
fp4
256,0008,192$0.12 / $0.36YesYesYes
View
Opper
gemma-4-31b-it
256,0008,192$0.46488 / $2.44062YesYesYes
View
OrcaRouter
google/gemma-4-31b-it
262,14432,768$0.13 / $0.38YesYesYes
View
Pioneer
google/gemma-4-31B-it
32,76832,768$0.50 / $0.50YesYesYes
View
QVAC
gemma4-31b
262,14432,768$0.00 / $0.00YesYesYes
View
Regolo AI
gemma4-31b
100,000100,000$0.46 / $2.42YesYesYes
View
Requesty
gemma-4-31b-it
262,1448,192$0.00 / $0.00YesYesNo
View
SiliconFlow
google/gemma-4-31B-it
262,144262,144$0.13 / $0.40NoYesYes
View
Tempr Gateway
google/gemma-4-31b-it
262,14432,768— / —YesYesYes
View
Tinfoil
gemma4-31b
262,14432,768$0.40 / $1.00YesYesYes
View
UnoRouter
gemma-4-31b-it:free
262,14432,768$0.00 / $0.00YesYesYes
View
Venice AI
google-gemma-4-31b-it
256,0008,192$0.12 / $0.36YesYesYes
View
Vercel AI Gateway
google/gemma-4-31b-it
262,144131,072$0.14 / $0.40NoYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers58Across 37 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 10 lowest-priced configurations
InputOutput
Google AI Studiogoogle-ai-studio
$0.00$0.00
DeepInfradeepinfra/turbo
$0.09$0.34
CoreWeavecoreweave/fp4
$0.10$0.34
DekaLLMdekallm
$0.10$0.33
Chuteschutes/fp4
$0.12$0.37
Venicevenice/fp4
$0.12$0.36
Crusoecrusoe/bf16
$0.14$0.40
Friendlifriendli
$0.14$0.40
Novitanovita/bf16
$0.14$0.40
DeepInfradeepinfra/fp8
$0.15$0.40
google/gemma-4-31b-itObserved
Input
$0.09
Output
$0.34
Cache read
$0.05
google/gemma-4-31b-it:freeObserved
Input
$0.00
Output
$0.00
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
Chuteschutes/fp4$0.12$0.37$0.012—
CoreWeavecoreweave/fp4$0.10$0.34$0.10—
Crusoecrusoe/bf16$0.14$0.40$0.14—
DeepInfradeepinfra/turbo$0.09$0.34$0.05—
DeepInfradeepinfra/fp8$0.15$0.40——
DekaLLMdekallm$0.10$0.33$0.05—
Friendlifriendli$0.14$0.40——
Io Netio-net$0.38$1.15$0.19—
ModelRunmodelrun/fp4$0.75$1.00$0.20—
Novitanovita/bf16$0.14$0.40——
Parasailparasail/fp8$0.15$0.40$0.06—
SambaNovasambanova$0.38$1.15——
SiliconFlowsiliconflow/fp8$0.75$1.00$0.25—
Venicevenice/fp4$0.12$0.36$0.09—
Google AI Studiogoogle-ai-studio$0.00$0.00——

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Chuteschutes/fp4
2,558–22,870.27 ms
CoreWeavecoreweave/fp4
391–4,878.62 ms
Crusoecrusoe/bf16
487–8,009 ms
DeepInfradeepinfra/turbo
574.5–5,321.94 ms
DeepInfradeepinfra/fp8
1,762–11,978.26 ms
DekaLLMdekallm
1,107–10,305.72 ms
All latency percentiles 15 endpoints
EndpointP50P75P90P99
Chuteschutes/fp42,558 ms4,912.5 ms8,547.6 ms22,870.27 ms
CoreWeavecoreweave/fp4391 ms642.25 ms1,266.9 ms4,878.62 ms
Crusoecrusoe/bf16487 ms712 ms1,201.8 ms8,009 ms
DeepInfradeepinfra/turbo574.5 ms1,024 ms1,744 ms5,321.94 ms
DeepInfradeepinfra/fp81,762 ms3,125.25 ms5,426 ms11,978.26 ms
DekaLLMdekallm1,107 ms2,156 ms4,018.4 ms10,305.72 ms
Friendlifriendli246 ms809 ms1,670.9 ms4,043.81 ms
Io Netio-net501 ms1,491.5 ms3,283.2 ms6,886.86 ms
ModelRunmodelrun/fp4189 ms268 ms362 ms1,012.09 ms
Novitanovita/bf161,266 ms13,044.5 ms32,622.4 ms73,495.85 ms
Parasailparasail/fp81,320 ms2,242.25 ms3,679.8 ms9,264.13 ms
SambaNovasambanova3,033 ms7,578.75 ms9,804 ms35,853.55 ms
SiliconFlowsiliconflow/fp81,237 ms1,627.5 ms3,455.2 ms5,537.78 ms
Venicevenice/fp4482 ms751 ms1,161 ms2,963.63 ms
Google AI Studiogoogle-ai-studio1,215 ms1,715 ms4,051.6 ms38,563.67 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Chuteschutes/fp4
22–85 t/s
CoreWeavecoreweave/fp4
43–113 t/s
Crusoecrusoe/bf16
34–105.73 t/s
DeepInfradeepinfra/turbo
42–71 t/s
DeepInfradeepinfra/fp8
18–33 t/s
DekaLLMdekallm
24–95 t/s
All throughput percentiles 15 endpoints
EndpointP50P75P90P99
Chuteschutes/fp422 t/s35 t/s50 t/s85 t/s
CoreWeavecoreweave/fp443 t/s59 t/s79 t/s113 t/s
Crusoecrusoe/bf1634 t/s51 t/s73 t/s105.73 t/s
DeepInfradeepinfra/turbo42 t/s51 t/s58 t/s71 t/s
DeepInfradeepinfra/fp818 t/s21 t/s25 t/s33 t/s
DekaLLMdekallm24 t/s36 t/s54 t/s95 t/s
Friendlifriendli118 t/s157 t/s208 t/s295 t/s
Io Netio-net38 t/s44 t/s45 t/s46.74 t/s
ModelRunmodelrun/fp4230 t/s261 t/s324 t/s384 t/s
Novitanovita/bf1628 t/s33 t/s36 t/s39 t/s
Parasailparasail/fp823 t/s29 t/s36 t/s50 t/s
SambaNovasambanova79 t/s114 t/s143.5 t/s165.85 t/s
SiliconFlowsiliconflow/fp841 t/s48 t/s58 t/s64.45 t/s
Venicevenice/fp443 t/s52 t/s59 t/s68 t/s
Google AI Studiogoogle-ai-studio29 t/s33 t/s35 t/s36 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 15 endpoints
Endpoint5 minutes30 minutes24 hours
Chuteschutes/fp4
CoreWeavecoreweave/fp4
Crusoecrusoe/bf16
DeepInfradeepinfra/turbo
DeepInfradeepinfra/fp8
DekaLLMdekallm
Friendlifriendli
Io Netio-net
ModelRunmodelrun/fp4
Novitanovita/bf16
Parasailparasail/fp8
SambaNovasambanova
SiliconFlowsiliconflow/fp8
Venicevenice/fp4
Google AI Studiogoogle-ai-studio

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
google/gemma-4-31b-itObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 50
Coding43.4
Agentic4.2
Intelligence14.7
google/gemma-4-31b-it:freeObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 50
Coding43.4
Agentic4.2
Intelligence14.7

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
gpqa diamondGoogle: Gemma 4 31B
81.794%
tau bench verified airlineGoogle: Gemma 4 31B
76.682%
EvaluationConfigurationSourceResultsDetails
Composite indicesGemma 4 31B (Reasoning)google/gemma-4-31b-it-20260402Artificial AnalysisIntelligence 14.7 · Coding 43.4 · Agentic 4.2
Details
gpqa diamondGoogle: Gemma 4 31Bgoogle/gemma-4-31b-it-20260402OpenRouter81.794%
Details
tau bench verified airlineGoogle: Gemma 4 31Bgoogle/gemma-4-31b-it-20260402OpenRouter76.682%
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentgoogle/gemma-4-31b-it-202604022–9 turns
$0.007523
Details
Hermes Agentgoogle/gemma-4-31b-it-202604021 turn
$0.001703
Details
Hermes Agentgoogle/gemma-4-31b-it-2026040210–49 turns
$0.092514
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage4,905,684,489,070 reported tokens · 89 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 89 days
Date (UTC)Reported tokensReported routes
Oct 3, 202658,887,262,9571
Oct 2, 202663,662,661,5681
Oct 1, 202671,271,670,4531
Sep 30, 202668,051,662,7731
Sep 29, 202670,119,450,2761
Sep 28, 202667,326,788,6341
Sep 27, 202660,886,166,0751
Sep 26, 202658,028,444,9801
Sep 25, 202660,412,029,6091
Sep 24, 202657,208,674,1071
Sep 23, 202658,400,449,7001
Sep 22, 202662,833,338,2561
Sep 21, 202660,695,896,7671
Sep 20, 202650,602,610,9191
Sep 19, 202644,196,526,9181
Sep 18, 202654,634,084,1321
Sep 17, 202657,260,692,8651
Sep 16, 202655,457,705,0001
Sep 15, 202671,488,436,8381
Sep 14, 202654,386,645,6751
Sep 13, 202643,196,690,1921
Sep 12, 202643,656,873,8611
Sep 11, 202651,996,290,5891
Sep 10, 202658,128,636,4161
Sep 9, 202655,589,992,1391
Sep 8, 202659,357,498,1951
Sep 7, 202647,785,651,4721
Sep 6, 202645,994,206,4141
Sep 5, 202642,123,096,1971
Sep 4, 202648,998,207,2631
Sep 3, 202656,986,582,4131
Sep 2, 202658,893,691,2361
Sep 1, 202650,027,952,2461
Aug 31, 202645,667,405,0911
Aug 30, 202641,563,662,9681
Aug 29, 202638,983,336,4771
Aug 28, 202647,157,490,8101
Aug 27, 202652,032,833,9151
Aug 26, 202649,059,527,2361
Aug 25, 202650,338,770,7241
Aug 24, 202649,790,990,9361
Aug 23, 202643,115,910,9141
Aug 22, 202642,028,494,0161
Aug 21, 202654,266,275,1281
Aug 20, 202684,264,076,3211
Aug 19, 202658,139,714,2231
Aug 18, 202654,075,051,3351
Aug 17, 202651,195,386,2241
Aug 16, 202645,677,203,8681
Aug 15, 202649,407,478,7431
Aug 14, 202654,554,005,5961
Aug 13, 202688,309,998,8701
Aug 12, 202690,514,544,0091
Aug 11, 202686,793,779,4221
Aug 10, 202660,349,064,0561
Aug 9, 202662,601,543,5551
Aug 8, 202656,666,442,2771
Aug 7, 202662,261,223,1761
Aug 6, 202659,946,744,1901
Aug 5, 202658,317,628,4751
Aug 4, 202655,964,966,1461
Aug 3, 202658,380,676,1151
Aug 2, 202658,206,755,8801
Aug 1, 202659,832,944,8811
Jul 31, 202667,704,518,3051
Jul 30, 202660,257,543,5481
Jul 29, 202655,898,837,5711
Jul 28, 202657,046,179,5991
Jul 27, 202654,833,836,7351
Jul 26, 202647,150,857,0851
Jul 25, 202646,328,267,2091
Jul 24, 202654,956,549,0111
Jul 23, 202656,373,530,3671
Jul 22, 202659,735,936,8861
Jul 21, 202655,210,350,9971
Jul 20, 202650,749,767,1261
Jul 19, 202645,525,733,0111
Jul 18, 202647,768,892,4221
Jul 17, 202650,878,261,3921
Jul 16, 202648,291,474,7131
Jul 15, 202646,111,510,7011
Jul 14, 202645,717,012,2371
Jul 13, 202648,915,978,4361
Jul 12, 202640,527,579,5401
Jul 11, 202640,620,552,0251
Jul 10, 202650,126,368,5941
Jul 9, 202644,530,904,2151
Jul 8, 202645,658,252,7541
Jul 7, 202646,755,301,8791

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
google/gemma-4-31b-itObserved
Context
262,144
Maximum output
16,384
Tokenizer
Gemma
Input
image, text, video
Output
text
Moderation
No
Added to OpenRouter
2026-04-02
Base URLhttps://openrouter.ai/api/v1
Model IDgoogle/gemma-4-31b-it
Version IDgoogle/gemma-4-31b-it-20260402
Weights IDgoogle/gemma-4-31B-it

Reasoning

Required
No
Enabled by default
No
Default effort
—
Supported efforts
—

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Default parameters

top k
64
top p
0.95
temperature
1
google/gemma-4-31b-it:freeObserved
Context
262,144
Maximum output
32,768
Tokenizer
Gemma
Input
image, text, video
Output
text
Moderation
No
Added to OpenRouter
2026-04-02
Base URLhttps://openrouter.ai/api/v1
Model IDgoogle/gemma-4-31b-it:free
Version IDgoogle/gemma-4-31b-it-20260402
Weights IDgoogle/gemma-4-31B-it

Reasoning

Required
No
Enabled by default
No
Default effort
—
Supported efforts
—

Supported parameters

  • include_reasoning
  • max_tokens
  • reasoning
  • response_format
  • seed
  • temperature
  • tool_choice
  • tools
  • top_p

Default parameters

top k
64
top p
0.95
temperature
1

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Gemma 4 31B IT cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Gemma 4 31B IT?

There are 58 cataloged offers across 37 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 262,144 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Resources

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · fa9973480682
  • · 543f0b05f2db