Gemma 4 26B A4B IT

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video (up to 60s at 1fps). Features a 256K token context window, native function calling, configurable thinking/reasoning mode, and structured output support. Released under Apache 2.0.

google/gemma-4-26b-a4b-it
Organization
Google
Family
gemma
Providers
25
Context
262,144
Output limit
32,768
Knowledge
—
Release
2026-04-02
Updated
2026-04-02
Weights
Open
Input
text, image
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

39 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
AKI.IO
gemma4-26b
256,00032,768$0.10 / $0.50YesYesYes
View
Amazon Bedrock
google.gemma-4-26b-a4b
262,14432,768$0.13 / $0.40YesYesYes
View
Charm Hyper
gemma-4-26b-a4b-it
256,00025,600$0.098 / $0.334YesYesYes
View
Cloudflare Workers AI
@cf/google/gemma-4-26b-a4b-it
256,00016,384$0.10 / $0.30YesYesYes
View
CoreWeave
google/gemma-4-26B-A4B-it
262,144262,144$0.10 / $0.30YesYesYes
View
Cortecs
gemma-4-26b-a4b-it
262,00081,920$0.111 / $0.557YesYesYes
View
Deep Infra
google/gemma-4-26B-A4B-it
262,14432,768$0.07 / $0.34YesYesYes
View
DevPass (LLM Gateway)
gemma-4-26b-a4b-it
262,14432,768$0.07 / $0.34YesYesYes
View
EmpirioLabs AI
gemma-4-26b-a4b
262,14432,768$0.05 / $0.29YesYesYes
View
evroc
google/gemma-4-26B-A4B-it
262,14432,768$0.144 / $0.575YesYesYes
View
GreenPT
gemma4
262,14432,768$0.57 / $1.71YesYesYes
View
Hugging Face
google/gemma-4-26B-A4B-it
262,14432,768$0.13 / $0.40YesYesYes
View
Kilo Gateway
google/gemma-4-26b-a4b-it
262,144235,929$0.042 / $0.22YesYesYes
View
LLM Gateway
novita/gemma-4-26b-a4b-it
262,14432,768$0.13 / $0.40NoYesNo
View
LLM Gateway
deepinfra/gemma-4-26b-a4b-it
262,14432,768$0.07 / $0.34YesYesNo
View
Merge Gateway
google/gemma-4-26b-a4b-it
262,14465,536$0.13 / $0.40YesNoNo
View
NaN
gemma4
262,14432,768$0.00 / $0.00YesYesYes
View
NanoGPT
google/gemma-4-26b-a4b-it
262,144131,072$0.12 / $0.38YesYesYes
View
NanoGPT
google/gemma-4-26b-a4b-it:thinking
262,144131,072$0.13 / $0.40YesYesYes
View
Cloudflarevia OpenRouter
google/gemma-4-26b-a4b-it
unknown
256,000230,400$0.10 / $0.30YesYesNo
View
CoreWeavevia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,144235,929$0.10 / $0.30YesNoYes
View
Darkbloomvia OpenRouter
google/gemma-4-26b-a4b-it
unknown
131,07232,768$0.042 / $0.22YesYesYes
View
DeepInfravia OpenRouter
google/gemma-4-26b-a4b-it
fp8
262,14416,384$0.07 / $0.34YesYesYes
View
DekaLLMvia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,144235,929$0.06 / $0.33YesYesNo
View
Googlevia OpenRouter
google/gemma-4-26b-a4b-it
unknown
262,144235,929$0.15 / $0.60YesYesYes
View
Google AI Studiovia OpenRouter
google/gemma-4-26b-a4b-it:free
unknown
262,14432,768$0.00 / $0.00YesYesNo
View
Io Netvia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,142235,927$0.15 / $0.50YesYesNo
View
NextBitvia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,144235,929$0.0675 / $0.225YesYesYes
View
Novitavia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,144131,072$0.13 / $0.40YesYesNo
View
Parasailvia OpenRouter
google/gemma-4-26b-a4b-it
bf16
262,144235,929$0.13 / $0.40YesNoYes
View
SiliconFlowvia OpenRouter
google/gemma-4-26b-a4b-it
fp8
262,144235,929$0.14 / $0.40YesNoYes
View
Venicevia OpenRouter
google/gemma-4-26b-a4b-it
bf16
256,0008,192$0.13 / $0.40YesYesYes
View
OrcaRouter
google/gemma-4-26b-a4b-it
262,14432,768$0.06 / $0.33YesYesYes
View
Requesty
gemma-4-26b-a4b-it
262,144262,144$0.07 / $0.34YesYesYes
View
Scaleway
gemma-4-26b-a4b-it
256,00016,384$0.25 / $0.50YesYesYes
View
SiliconFlow
google/gemma-4-26B-A4B-it
262,144262,144$0.12 / $0.40NoYesYes
View
Tempr Gateway
google/gemma-4-26b-a4b-it
262,14432,768— / —YesYesYes
View
Venice AI
google-gemma-4-26b-a4b-it
256,0008,192$0.13 / $0.40YesYesYes
View
Vercel AI Gateway
google/gemma-4-26b-a4b-it
262,144131,072$0.15 / $0.60YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers39Across 25 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 10 lowest-priced configurations
InputOutput
Google AI Studiogoogle-ai-studio
$0.00$0.00
Darkbloomdarkbloom
$0.042$0.22
DekaLLMdekallm/bf16
$0.06$0.33
NextBitnextbit/bf16
$0.0675$0.225
DeepInfradeepinfra/fp8
$0.07$0.34
Cloudflarecloudflare
$0.10$0.30
CoreWeavecoreweave/bf16
$0.10$0.30
Novitanovita/bf16
$0.13$0.40
Parasailparasail/bf16
$0.13$0.40
Venicevenice/bf16
$0.13$0.40
google/gemma-4-26b-a4b-itObserved
Input
$0.0675
Output
$0.225
Cache read
$0.0375
google/gemma-4-26b-a4b-it:freeObserved
Input
$0.00
Output
$0.00
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
Cloudflarecloudflare$0.10$0.30$0.05—
CoreWeavecoreweave/bf16$0.10$0.30$0.05—
Darkbloomdarkbloom$0.042$0.22$0.021—
DeepInfradeepinfra/fp8$0.07$0.34——
DekaLLMdekallm/bf16$0.06$0.33$0.04—
Googlegoogle-vertex/global$0.15$0.60——
Io Netio-net/bf16$0.15$0.50$0.15—
NextBitnextbit/bf16$0.0675$0.225$0.0375—
Novitanovita/bf16$0.13$0.40——
Parasailparasail/bf16$0.13$0.40$0.05—
SiliconFlowsiliconflow/fp8$0.14$0.40$0.05—
Venicevenice/bf16$0.13$0.40$0.05—
Google AI Studiogoogle-ai-studio$0.00$0.00——

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Cloudflarecloudflare
686–12,142.18 ms
CoreWeavecoreweave/bf16
147–1,031.36 ms
Darkbloomdarkbloom
851–6,969.36 ms
DeepInfradeepinfra/fp8
818–5,250.36 ms
DekaLLMdekallm/bf16
757–4,521.25 ms
Googlegoogle-vertex/global
958–16,617.72 ms
All latency percentiles 13 endpoints
EndpointP50P75P90P99
Cloudflarecloudflare686 ms1,156 ms3,541.8 ms12,142.18 ms
CoreWeavecoreweave/bf16147 ms244 ms336 ms1,031.36 ms
Darkbloomdarkbloom851 ms1,443 ms2,560.7 ms6,969.36 ms
DeepInfradeepinfra/fp8818 ms1,530 ms2,215 ms5,250.36 ms
DekaLLMdekallm/bf16757 ms1,191.25 ms1,869.8 ms4,521.25 ms
Googlegoogle-vertex/global958 ms2,380 ms3,697 ms16,617.72 ms
Io Netio-net/bf16456 ms1,568 ms4,363.2 ms9,027 ms
NextBitnextbit/bf16473 ms1,029.25 ms1,774.8 ms3,298.18 ms
Novitanovita/bf16865 ms1,223.25 ms1,635.9 ms3,484.26 ms
Parasailparasail/bf16761 ms1,090 ms1,611.3 ms3,868.72 ms
SiliconFlowsiliconflow/fp81,530 ms2,256.25 ms3,734 ms11,211.6 ms
Venicevenice/bf161,133 ms1,516 ms2,073.4 ms4,994.82 ms
Google AI Studiogoogle-ai-studio1,024.5 ms1,365.75 ms2,003.8 ms10,844.8 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
Cloudflarecloudflare
40–169 t/s
CoreWeavecoreweave/bf16
92–150 t/s
Darkbloomdarkbloom
59–112 t/s
DeepInfradeepinfra/fp8
30–58 t/s
DekaLLMdekallm/bf16
45–139 t/s
Googlegoogle-vertex/global
41–134 t/s
All throughput percentiles 13 endpoints
EndpointP50P75P90P99
Cloudflarecloudflare40 t/s49 t/s111 t/s169 t/s
CoreWeavecoreweave/bf1692 t/s106 t/s119 t/s150 t/s
Darkbloomdarkbloom59 t/s73 t/s89 t/s112 t/s
DeepInfradeepinfra/fp830 t/s39 t/s46 t/s58 t/s
DekaLLMdekallm/bf1645 t/s62 t/s84 t/s139 t/s
Googlegoogle-vertex/global41 t/s63 t/s83 t/s134 t/s
Io Netio-net/bf1661 t/s82 t/s98 t/s122.75 t/s
NextBitnextbit/bf1645 t/s79 t/s98 t/s124 t/s
Novitanovita/bf1648 t/s157 t/s183 t/s221 t/s
Parasailparasail/bf1659 t/s86 t/s115 t/s167.68 t/s
SiliconFlowsiliconflow/fp837 t/s61 t/s85 t/s122 t/s
Venicevenice/bf1643 t/s65 t/s87 t/s155.93 t/s
Google AI Studiogoogle-ai-studio40 t/s44 t/s45.5 t/s47 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 13 endpoints
Endpoint5 minutes30 minutes24 hours
Cloudflarecloudflare
CoreWeavecoreweave/bf16
Darkbloomdarkbloom
DeepInfradeepinfra/fp8
DekaLLMdekallm/bf16
Googlegoogle-vertex/global
Io Netio-net/bf16
NextBitnextbit/bf16
Novitanovita/bf16
Parasailparasail/bf16
SiliconFlowsiliconflow/fp8
Venicevenice/bf16
Google AI Studiogoogle-ai-studio

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
google/gemma-4-26b-a4b-itObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 40
Coding39.3
google/gemma-4-26b-a4b-it:freeObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 40
Coding39.3

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
tau bench verified airlineGoogle: Gemma 4 26B A4B
67.917%
gpqa diamondGoogle: Gemma 4 26B A4B
73.064%
EvaluationConfigurationSourceResultsDetails
tau bench verified airlineGoogle: Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-it-20260403OpenRouter67.917%
Details
gpqa diamondGoogle: Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-it-20260403OpenRouter73.064%
Details
Composite indicesGemma 4 26B A4B (Reasoning)google/gemma-4-26b-a4b-it-20260403Artificial AnalysisIntelligence — · Coding 39.3 · Agentic —
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentgoogle/gemma-4-26b-a4b-it-202604032–9 turns
$0.003008
Details
Hermes Agentgoogle/gemma-4-26b-a4b-it-2026040310–49 turns
$0.030981
Details
Hermes Agentgoogle/gemma-4-26b-a4b-it-202604031 turn
$0.00042
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage3,788,973,106,376 reported tokens · 78 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 78 days
Date (UTC)Reported tokensReported routes
Sep 28, 202650,403,161,2531
Sep 27, 202642,454,663,3221
Sep 25, 202650,357,319,0741
Sep 24, 202651,962,449,2381
Sep 22, 202656,336,004,8421
Sep 21, 202649,772,338,9781
Sep 20, 202638,328,806,5421
Sep 19, 202636,742,569,1981
Sep 17, 202646,047,919,5661
Sep 15, 202642,888,051,7131
Sep 13, 202635,902,165,9311
Sep 11, 202644,827,162,5471
Sep 10, 202655,244,402,1341
Sep 9, 202645,212,397,3581
Sep 8, 202644,620,791,7821
Sep 7, 202649,154,356,7261
Sep 6, 202650,951,812,1081
Sep 5, 202648,050,107,0901
Sep 4, 202652,385,661,5551
Sep 3, 202657,475,159,9441
Sep 2, 202653,587,770,2001
Sep 1, 202648,026,919,0341
Aug 31, 202643,651,687,0711
Aug 30, 202634,210,091,0251
Aug 29, 202635,637,610,9491
Aug 28, 202639,779,623,9901
Aug 27, 202644,704,075,0601
Aug 26, 202653,331,209,5031
Aug 25, 202648,527,627,6561
Aug 24, 202647,361,988,4211
Aug 23, 202648,022,912,8481
Aug 22, 202645,469,111,2801
Aug 21, 202647,153,821,9171
Aug 20, 202649,208,552,2641
Aug 19, 202652,441,138,5131
Aug 18, 202658,684,560,2841
Aug 17, 202651,353,420,8991
Aug 16, 202650,853,260,5831
Aug 15, 202647,489,779,3031
Aug 14, 202649,563,625,3181
Aug 13, 202652,357,034,7131
Aug 12, 202649,905,803,3501
Aug 11, 202661,384,692,0781
Aug 10, 202660,430,213,0761
Aug 9, 202657,228,554,4961
Aug 8, 202659,781,700,6301
Aug 7, 202650,092,326,6881
Aug 6, 202646,048,533,7231
Aug 5, 202645,340,977,6871
Aug 4, 202642,479,379,1551
Aug 3, 202644,674,393,6851
Aug 2, 202646,116,286,6281
Aug 1, 202642,866,305,1111
Jul 31, 202653,614,101,2561
Jul 30, 202655,892,929,8891
Jul 29, 202656,677,753,4971
Jul 28, 202651,368,108,2551
Jul 27, 202650,445,279,1221
Jul 26, 202648,287,635,9451
Jul 25, 202640,145,749,4111
Jul 24, 202646,166,667,4011
Jul 23, 202649,204,455,9201
Jul 22, 202647,758,512,3721
Jul 21, 202651,000,706,2051
Jul 20, 202648,628,954,3931
Jul 19, 202659,934,969,6831
Jul 18, 202653,069,270,8511
Jul 17, 202655,728,681,9181
Jul 16, 202649,304,797,3921
Jul 15, 202644,354,653,2861
Jul 14, 202639,375,609,5261
Jul 13, 202642,473,152,3651
Jul 12, 202639,759,972,4651
Jul 11, 202642,463,248,5971
Jul 10, 202649,891,331,0721
Jul 9, 202649,749,362,1251
Jul 8, 202649,238,072,0091
Jul 7, 202657,556,843,3821

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
google/gemma-4-26b-a4b-itObserved
Context
262,144
Maximum output
235,929
Tokenizer
Gemma
Input
image, text, video
Output
text
Moderation
No
Added to OpenRouter
2026-04-03
Base URLhttps://openrouter.ai/api/v1
Model IDgoogle/gemma-4-26b-a4b-it
Version IDgoogle/gemma-4-26b-a4b-it-20260403
Weights IDgoogle/gemma-4-26B-A4B-it

Reasoning

Required
No
Enabled by default
No
Default effort
—
Supported efforts
—

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Default parameters

top k
64
top p
0.95
temperature
1
google/gemma-4-26b-a4b-it:freeObserved
Context
262,144
Maximum output
32,768
Tokenizer
Gemma
Input
image, text, video
Output
text
Moderation
No
Added to OpenRouter
2026-04-03
Base URLhttps://openrouter.ai/api/v1
Model IDgoogle/gemma-4-26b-a4b-it:free
Version IDgoogle/gemma-4-26b-a4b-it-20260403
Weights IDgoogle/gemma-4-26B-A4B-it

Reasoning

Required
No
Enabled by default
No
Default effort
—
Supported efforts
—

Supported parameters

  • include_reasoning
  • max_tokens
  • reasoning
  • response_format
  • seed
  • temperature
  • tool_choice
  • tools
  • top_p

Default parameters

top k
64
top p
0.95
temperature
1

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Gemma 4 26B A4B IT cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Gemma 4 26B A4B IT?

There are 39 cataloged offers across 25 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 262,144 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Resources

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · b5ad30ea35c8
  • · e61c27faf2b5