GPT OSS 20B

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

openai/gpt-oss-20b
Organization
OpenAI
Family
gpt-oss
Providers
19
Context
131,072
Output limit
32,768
Knowledge
—
Release
2025-08-05
Updated
2025-08-05
Weights
Open
Input
text
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

42 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
Amazon Bedrock
openai.gpt-oss-20b-1:0
131,072128,000$0.07 / $0.30YesYesYes
View
Amazon Bedrock
openai.gpt-oss-20b
131,072131,072$0.07 / $0.30YesYesYes
View
Amazon Bedrock
us-gov.openai.gpt-oss-20b-1:0
128,00016,384$0.084 / $0.36YesYesYes
View
Cloudflare Workers AI
@cf/openai/gpt-oss-20b
128,00016,384$0.20 / $0.30YesYesYes
View
CoreWeave
openai/gpt-oss-20b
131,072131,072$0.03 / $0.13YesYesYes
View
Cortecs
gpt-oss-20b
131,000131,000$0.045 / $0.167YesYesYes
View
Eden AI
deepinfra/openai/gpt-oss-20b
131,07232,768$0.03 / $0.14YesYesYes
View
Eden AI
cloudflare/@cf/openai/gpt-oss-20b
128,00032,768$0.20 / $0.30YesYesNo
View
Eden AI
groq/openai/gpt-oss-20b
131,07232,768$0.075 / $0.30YesYesYes
View
Eden AI
databricks/databricks-gpt-oss-20b@eu
131,07232,768$0.07 / $0.30002YesYesNo
View
Eden AI
flexai/gpt-oss-20b
131,07232,768$0.02 / $0.10YesYesYes
View
Eden AI
databricks/databricks-gpt-oss-20b
131,07232,768$0.07 / $0.30002YesYesNo
View
Eden AI
ovhcloud/gpt-oss-20b
131,07232,768$0.05 / $0.18YesNoYes
View
Eden AI
greenference/gpt-oss-20b
131,07232,768$0.009 / $0.045YesNoYes
View
Hugging Face
openai/gpt-oss-20b
131,07232,768$0.10 / $0.50YesYesYes
View
Impossibl
groq/gpt-oss-20b
131,07232,768$0.075 / $0.30YesYesYes
View
Impossibl
fireworks/gpt-oss-20b
131,07232,768$0.07 / $0.30YesYesYes
View
Kilo Gateway
openai/gpt-oss-20b
131,07232,768$0.018 / $0.09YesYesYes
View
LLM Gateway
groq/gpt-oss-20b
131,07232,766$0.10 / $0.50YesYesNo
View
LLM Gateway
consensusprotocol/gpt-oss-20b
128,00032,768$0.04 / $0.19YesYesNo
View
Merge Gateway
openai/gpt-oss-20b
128,00032,000$0.04 / $0.15YesNoNo
View
NanoGPT
openai/gpt-oss-20b
128,00016,384$0.20 / $0.30YesYesNo
View
OCI Generative AI
openai.gpt-oss-20b
128,00016,384$0.07 / $0.30YesYesYes
View
AkashMLvia OpenRouter
openai/gpt-oss-20b
fp4
131,072117,964$0.02 / $0.10YesYesYes
View
Amazon Bedrockvia OpenRouter
openai/gpt-oss-20b
unknown
131,072117,964$0.07 / $0.15YesYesNo
View
Amazon Bedrockvia OpenRouter
openai/gpt-oss-20b
unknown
131,072117,964$0.07 / $0.15YesYesNo
View
CoreWeavevia OpenRouter
openai/gpt-oss-20b
fp4
131,072117,964$0.03 / $0.13YesYesYes
View
Darkbloomvia OpenRouter
openai/gpt-oss-20b
fp8
131,07232,768$0.018 / $0.09YesYesYes
View
DeepInfravia OpenRouter
openai/gpt-oss-20b
bf16
131,072117,964$0.03 / $0.14YesYesYes
View
DekaLLMvia OpenRouter
openai/gpt-oss-20b
bf16
131,072117,964$0.029 / $0.14YesYesYes
View
Googlevia OpenRouter
openai/gpt-oss-20b
unknown
131,07232,768$0.07 / $0.25YesNoYes
View
Groqvia OpenRouter
openai/gpt-oss-20b
unknown
131,07265,536$0.075 / $0.30YesYesYes
View
Novitavia OpenRouter
openai/gpt-oss-20b
fp4
131,07232,768$0.04 / $0.15YesNoYes
View
Parasailvia OpenRouter
openai/gpt-oss-20b
fp4
131,072117,964$0.03 / $0.15YesYesYes
View
PrimeIntellectvia OpenRouter
openai/gpt-oss-20b
unknown
131,07232,768$0.07 / $0.30YesYesYes
View
SiliconFlowvia OpenRouter
openai/gpt-oss-20b
fp8
131,0728,192$0.04 / $0.18YesNoYes
View
Opper
gpt-oss-20b
128,0008,192$0.11622 / $0.488124YesYesYes
View
Pioneer
openai/gpt-oss-20b
131,0728,192$0.07 / $0.30YesYesYes
View
QVAC
gpt-oss-20b
131,07232,768$0.00 / $0.00YesYesYes
View
STACKIT
openai/gpt-oss-20b
131,0728,192$0.18 / $0.29YesYesNo
View
Tempr Gateway
groq/openai/gpt-oss-20b
131,07265,536$0.075 / $0.30YesYesYes
View
Vercel AI Gateway
openai/gpt-oss-20b
131,0728,192$0.03 / $0.14YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers42Across 19 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 10 lowest-priced configurations
InputOutput
Darkbloomdarkbloom/fp8
$0.018$0.09
AkashMLakashml/fp4
$0.02$0.10
DekaLLMdekallm/bf16
$0.029$0.14
CoreWeavecoreweave/fp4
$0.03$0.13
DeepInfradeepinfra/bf16
$0.03$0.14
Parasailparasail/fp4
$0.03$0.15
Novitanovita/fp4
$0.04$0.15
SiliconFlowsiliconflow/fp8
$0.04$0.18
Amazon Bedrockamazon-bedrock
$0.07$0.15
Amazon Bedrockamazon-bedrock/eu-west-1
$0.07$0.15
openai/gpt-oss-20bObserved
Input
$0.018
Output
$0.09
Cache read
$0.009
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
AkashMLakashml/fp4$0.02$0.10——
Amazon Bedrockamazon-bedrock$0.07$0.15——
Amazon Bedrockamazon-bedrock/eu-west-1$0.07$0.15——
CoreWeavecoreweave/fp4$0.03$0.13$0.03—
Darkbloomdarkbloom/fp8$0.018$0.09$0.009—
DeepInfradeepinfra/bf16$0.03$0.14——
DekaLLMdekallm/bf16$0.029$0.14$0.029—
Googlegoogle-vertex/us-central1$0.07$0.25——
Groqgroq$0.075$0.30$0.0375—
Novitanovita/fp4$0.04$0.15——
Parasailparasail/fp4$0.03$0.15$0.02—
PrimeIntellectprimeintellect$0.07$0.30——
SiliconFlowsiliconflow/fp8$0.04$0.18——

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AkashMLakashml/fp4
654–15,007.86 ms
Amazon Bedrockamazon-bedrock
477–2,259.8 ms
Amazon Bedrockamazon-bedrock/eu-west-1
689–1,482.94 ms
CoreWeavecoreweave/fp4
2,696–8,134.09 ms
Darkbloomdarkbloom/fp8
3,738.5–25,852.35 ms
DeepInfradeepinfra/bf16
229–1,777.09 ms
All latency percentiles 13 endpoints
EndpointP50P75P90P99
AkashMLakashml/fp4654 ms1,227.25 ms3,575.7 ms15,007.86 ms
Amazon Bedrockamazon-bedrock477 ms618 ms786 ms2,259.8 ms
Amazon Bedrockamazon-bedrock/eu-west-1689 ms896 ms959 ms1,482.94 ms
CoreWeavecoreweave/fp42,696 ms3,792.25 ms5,032.8 ms8,134.09 ms
Darkbloomdarkbloom/fp83,738.5 ms6,412 ms10,592 ms25,852.35 ms
DeepInfradeepinfra/bf16229 ms341 ms642 ms1,777.09 ms
DekaLLMdekallm/bf16693 ms1,334 ms2,918 ms11,730.4 ms
Googlegoogle-vertex/us-central1474 ms1,094.25 ms1,480.2 ms2,135.05 ms
Groqgroq717 ms1,244 ms1,836.9 ms4,003.09 ms
Novitanovita/fp4633 ms756.75 ms962.4 ms2,406.51 ms
Parasailparasail/fp41,047 ms1,288 ms1,593 ms3,147.25 ms
PrimeIntellectprimeintellect479 ms694 ms1,005 ms2,897 ms
SiliconFlowsiliconflow/fp81,023.5 ms1,227.25 ms1,644.5 ms3,294.95 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AkashMLakashml/fp4
67–214.9 t/s
Amazon Bedrockamazon-bedrock
248–904.26 t/s
Amazon Bedrockamazon-bedrock/eu-west-1
147.5–254.8 t/s
CoreWeavecoreweave/fp4
246–299.09 t/s
Darkbloomdarkbloom/fp8
20–78.09 t/s
DeepInfradeepinfra/bf16
97–140 t/s
All throughput percentiles 13 endpoints
EndpointP50P75P90P99
AkashMLakashml/fp467 t/s91 t/s122 t/s214.9 t/s
Amazon Bedrockamazon-bedrock248 t/s419 t/s561 t/s904.26 t/s
Amazon Bedrockamazon-bedrock/eu-west-1147.5 t/s181 t/s190.5 t/s254.8 t/s
CoreWeavecoreweave/fp4246 t/s262 t/s276 t/s299.09 t/s
Darkbloomdarkbloom/fp820 t/s33 t/s48 t/s78.09 t/s
DeepInfradeepinfra/bf1697 t/s112 t/s122 t/s140 t/s
DekaLLMdekallm/bf1633 t/s40 t/s47.7 t/s70 t/s
Googlegoogle-vertex/us-central1226 t/s295 t/s368 t/s453.08 t/s
Groqgroq405 t/s601 t/s742 t/s902.09 t/s
Novitanovita/fp4157 t/s200 t/s238 t/s271 t/s
Parasailparasail/fp460 t/s85 t/s103 t/s124 t/s
PrimeIntellectprimeintellect180 t/s228 t/s251.4 t/s309.7 t/s
SiliconFlowsiliconflow/fp887 t/s124 t/s151.1 t/s171.81 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 13 endpoints
Endpoint5 minutes30 minutes24 hours
AkashMLakashml/fp4
Amazon Bedrockamazon-bedrock
Amazon Bedrockamazon-bedrock/eu-west-1
CoreWeavecoreweave/fp4
Darkbloomdarkbloom/fp8
DeepInfradeepinfra/bf16
DekaLLMdekallm/bf16
Googlegoogle-vertex/us-central1
Groqgroq
Novitanovita/fp4
Parasailparasail/fp4
PrimeIntellectprimeintellect
SiliconFlowsiliconflow/fp8

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
openai/gpt-oss-20bObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 30
Coding20.7
Agentic1.2
Intelligence9
Design ArenaCategoryEloWin rateRank
modelsdataviz94039.7%115
modelswebsite85927.9%129

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
tau bench verified airlineOpenAI: gpt-oss-20b
53.391%
gpqa diamondOpenAI: gpt-oss-20b
64.399%
EvaluationConfigurationSourceResultsDetails
models · websiteGPT OSS 20Bopenai/gpt-oss-20bDesign Arena859 Elo
Details
tau bench verified airlineOpenAI: gpt-oss-20bopenai/gpt-oss-20bOpenRouter53.391%
Details
models · datavizGPT OSS 20Bopenai/gpt-oss-20bDesign Arena940 Elo
Details
gpqa diamondOpenAI: gpt-oss-20bopenai/gpt-oss-20bOpenRouter64.399%
Details
Composite indicesgpt-oss-20b (high)openai/gpt-oss-20bArtificial AnalysisIntelligence 9 · Coding 20.7 · Agentic 1.2
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentopenai/gpt-oss-20b1 turn
$0.000181
Details
Hermes Agentopenai/gpt-oss-20b2–9 turns
$0.000992
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage1,478,733,052,827 reported tokens · 40 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 40 days
Date (UTC)Reported tokensReported routes
Sep 9, 202636,297,363,6411
Sep 8, 202644,771,168,6191
Sep 7, 202645,685,685,1971
Sep 6, 202643,863,747,3691
Sep 5, 202639,545,040,8351
Sep 4, 202646,673,792,7871
Sep 3, 202662,875,649,3051
Sep 2, 202667,630,538,3661
Sep 1, 202670,098,750,5281
Aug 31, 202683,153,659,7461
Aug 30, 202670,046,628,4891
Aug 29, 202667,189,102,8061
Aug 28, 202664,909,103,3731
Aug 27, 202663,771,547,0201
Aug 26, 202655,479,524,7011
Aug 25, 202629,465,089,3371
Aug 24, 202627,153,001,3541
Aug 22, 202618,516,279,7831
Aug 16, 202619,790,140,3011
Aug 12, 202623,832,008,0611
Aug 11, 202624,180,929,2181
Aug 10, 202619,935,560,3961
Aug 9, 202615,830,224,1841
Aug 8, 202617,078,278,5901
Aug 7, 202619,339,262,0341
Aug 5, 202623,722,184,5651
Aug 4, 202625,741,792,7251
Aug 3, 202627,018,485,8491
Aug 2, 202626,190,729,1041
Aug 1, 202620,007,120,2781
Jul 26, 202617,906,818,3221
Jul 16, 202630,274,340,7191
Jul 15, 202632,139,371,3301
Jul 14, 202631,265,428,3091
Jul 13, 202624,416,482,2741
Jul 11, 202623,073,011,8421
Jul 10, 202635,143,442,7471
Jul 9, 202625,477,604,2711
Jul 8, 202633,587,477,0921
Jul 7, 202625,656,687,3601

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
openai/gpt-oss-20bObserved
Context
131,072
Maximum output
32,768
Tokenizer
GPT
Input
text
Output
text
Moderation
No
Knowledge cutoff
2024-06-30
Added to OpenRouter
2025-08-05
Base URLhttps://openrouter.ai/api/v1
Model IDopenai/gpt-oss-20b
Version IDopenai/gpt-oss-20b
Weights IDopenai/gpt-oss-20b

Reasoning

Required
Yes
Enabled by default
Not reported
Default effort
medium
Supported efforts
high, medium, low

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does GPT OSS 20B cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer GPT OSS 20B?

There are 42 cataloged offers across 19 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 131,072 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Resources

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · 31b4f2cdfeb5
  • · c69dfbe20065