GPT OSS 20B
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.
- Organization
- OpenAI
- Family
- gpt-oss
- Providers
- 19
- Context
- 131,072
- Output limit
- 32,768
- Knowledge
- —
- Release
- 2025-08-05
- Updated
- 2025-08-05
- Weights
- Open
- Input
- text
- Output
- text
- Capabilities
- Tools, Reasoning, Structured, Temperature
Providers
Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.
| Provider | Model ID | Context | Output | Input / output · 1M | Reasoning | Tools | Structured | Details |
|---|---|---|---|---|---|---|---|---|
| Amazon Bedrock | openai.gpt-oss-20b-1:0 | 131,072 | 128,000 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| Amazon Bedrock | openai.gpt-oss-20b | 131,072 | 131,072 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| Amazon Bedrock | us-gov.openai.gpt-oss-20b-1:0 | 128,000 | 16,384 | $0.084 / $0.36 | Yes | Yes | Yes | View |
| Cloudflare Workers AI | @cf/openai/gpt-oss-20b | 128,000 | 16,384 | $0.20 / $0.30 | Yes | Yes | Yes | View |
| CoreWeave | openai/gpt-oss-20b | 131,072 | 131,072 | $0.03 / $0.13 | Yes | Yes | Yes | View |
| Cortecs | gpt-oss-20b | 131,000 | 131,000 | $0.045 / $0.167 | Yes | Yes | Yes | View |
| Eden AI | deepinfra/openai/gpt-oss-20b | 131,072 | 32,768 | $0.03 / $0.14 | Yes | Yes | Yes | View |
| Eden AI | cloudflare/@cf/openai/gpt-oss-20b | 128,000 | 32,768 | $0.20 / $0.30 | Yes | Yes | No | View |
| Eden AI | groq/openai/gpt-oss-20b | 131,072 | 32,768 | $0.075 / $0.30 | Yes | Yes | Yes | View |
| Eden AI | databricks/databricks-gpt-oss-20b@eu | 131,072 | 32,768 | $0.07 / $0.30002 | Yes | Yes | No | View |
| Eden AI | flexai/gpt-oss-20b | 131,072 | 32,768 | $0.02 / $0.10 | Yes | Yes | Yes | View |
| Eden AI | databricks/databricks-gpt-oss-20b | 131,072 | 32,768 | $0.07 / $0.30002 | Yes | Yes | No | View |
| Eden AI | ovhcloud/gpt-oss-20b | 131,072 | 32,768 | $0.05 / $0.18 | Yes | No | Yes | View |
| Eden AI | greenference/gpt-oss-20b | 131,072 | 32,768 | $0.009 / $0.045 | Yes | No | Yes | View |
| Hugging Face | openai/gpt-oss-20b | 131,072 | 32,768 | $0.10 / $0.50 | Yes | Yes | Yes | View |
| Impossibl | groq/gpt-oss-20b | 131,072 | 32,768 | $0.075 / $0.30 | Yes | Yes | Yes | View |
| Impossibl | fireworks/gpt-oss-20b | 131,072 | 32,768 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| Kilo Gateway | openai/gpt-oss-20b | 131,072 | 32,768 | $0.018 / $0.09 | Yes | Yes | Yes | View |
| LLM Gateway | groq/gpt-oss-20b | 131,072 | 32,766 | $0.10 / $0.50 | Yes | Yes | No | View |
| LLM Gateway | consensusprotocol/gpt-oss-20b | 128,000 | 32,768 | $0.04 / $0.19 | Yes | Yes | No | View |
| Merge Gateway | openai/gpt-oss-20b | 128,000 | 32,000 | $0.04 / $0.15 | Yes | No | No | View |
| NanoGPT | openai/gpt-oss-20b | 128,000 | 16,384 | $0.20 / $0.30 | Yes | Yes | No | View |
| OCI Generative AI | openai.gpt-oss-20b | 128,000 | 16,384 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| AkashMLvia OpenRouter | openai/gpt-oss-20b fp4 | 131,072 | 117,964 | $0.02 / $0.10 | Yes | Yes | Yes | View |
| Amazon Bedrockvia OpenRouter | openai/gpt-oss-20b unknown | 131,072 | 117,964 | $0.07 / $0.15 | Yes | Yes | No | View |
| Amazon Bedrockvia OpenRouter | openai/gpt-oss-20b unknown | 131,072 | 117,964 | $0.07 / $0.15 | Yes | Yes | No | View |
| CoreWeavevia OpenRouter | openai/gpt-oss-20b fp4 | 131,072 | 117,964 | $0.03 / $0.13 | Yes | Yes | Yes | View |
| Darkbloomvia OpenRouter | openai/gpt-oss-20b fp8 | 131,072 | 32,768 | $0.018 / $0.09 | Yes | Yes | Yes | View |
| DeepInfravia OpenRouter | openai/gpt-oss-20b bf16 | 131,072 | 117,964 | $0.03 / $0.14 | Yes | Yes | Yes | View |
| DekaLLMvia OpenRouter | openai/gpt-oss-20b bf16 | 131,072 | 117,964 | $0.029 / $0.14 | Yes | Yes | Yes | View |
| Googlevia OpenRouter | openai/gpt-oss-20b unknown | 131,072 | 32,768 | $0.07 / $0.25 | Yes | No | Yes | View |
| Groqvia OpenRouter | openai/gpt-oss-20b unknown | 131,072 | 65,536 | $0.075 / $0.30 | Yes | Yes | Yes | View |
| Novitavia OpenRouter | openai/gpt-oss-20b fp4 | 131,072 | 32,768 | $0.04 / $0.15 | Yes | No | Yes | View |
| Parasailvia OpenRouter | openai/gpt-oss-20b fp4 | 131,072 | 117,964 | $0.03 / $0.15 | Yes | Yes | Yes | View |
| PrimeIntellectvia OpenRouter | openai/gpt-oss-20b unknown | 131,072 | 32,768 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| SiliconFlowvia OpenRouter | openai/gpt-oss-20b fp8 | 131,072 | 8,192 | $0.04 / $0.18 | Yes | No | Yes | View |
| Opper | gpt-oss-20b | 128,000 | 8,192 | $0.11622 / $0.488124 | Yes | Yes | Yes | View |
| Pioneer | openai/gpt-oss-20b | 131,072 | 8,192 | $0.07 / $0.30 | Yes | Yes | Yes | View |
| QVAC | gpt-oss-20b | 131,072 | 32,768 | $0.00 / $0.00 | Yes | Yes | Yes | View |
| STACKIT | openai/gpt-oss-20b | 131,072 | 8,192 | $0.18 / $0.29 | Yes | Yes | No | View |
| Tempr Gateway | groq/openai/gpt-oss-20b | 131,072 | 65,536 | $0.075 / $0.30 | Yes | Yes | Yes | View |
| Vercel AI Gateway | openai/gpt-oss-20b | 131,072 | 8,192 | $0.03 / $0.14 | Yes | Yes | Yes | View |
Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.
Pricing
Compare token rates and cache charges across available routes and hosting configurations.
Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.
openai/gpt-oss-20bObserved - Input
- $0.018
- Output
- $0.09
- Cache read
- $0.009
Hosting prices through OpenRouter All endpoints and cache rates
| Endpoint | Input / 1M | Output / 1M | Cache read / 1M | Cache write / 1M |
|---|---|---|---|---|
| AkashMLakashml/fp4 | $0.02 | $0.10 | — | — |
| Amazon Bedrockamazon-bedrock | $0.07 | $0.15 | — | — |
| Amazon Bedrockamazon-bedrock/eu-west-1 | $0.07 | $0.15 | — | — |
| CoreWeavecoreweave/fp4 | $0.03 | $0.13 | $0.03 | — |
| Darkbloomdarkbloom/fp8 | $0.018 | $0.09 | $0.009 | — |
| DeepInfradeepinfra/bf16 | $0.03 | $0.14 | — | — |
| DekaLLMdekallm/bf16 | $0.029 | $0.14 | $0.029 | — |
| Googlegoogle-vertex/us-central1 | $0.07 | $0.25 | — | — |
| Groqgroq | $0.075 | $0.30 | $0.0375 | — |
| Novitanovita/fp4 | $0.04 | $0.15 | — | — |
| Parasailparasail/fp4 | $0.03 | $0.15 | $0.02 | — |
| PrimeIntellectprimeintellect | $0.07 | $0.30 | — | — |
| SiliconFlowsiliconflow/fp8 | $0.04 | $0.18 | — | — |
Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.
Performance
Review reported latency and output-throughput percentiles from the latest observed rolling window.
Time to first token · milliseconds
All latency percentiles 13 endpoints
| Endpoint | P50 | P75 | P90 | P99 |
|---|---|---|---|---|
| AkashMLakashml/fp4 | 654 ms | 1,227.25 ms | 3,575.7 ms | 15,007.86 ms |
| Amazon Bedrockamazon-bedrock | 477 ms | 618 ms | 786 ms | 2,259.8 ms |
| Amazon Bedrockamazon-bedrock/eu-west-1 | 689 ms | 896 ms | 959 ms | 1,482.94 ms |
| CoreWeavecoreweave/fp4 | 2,696 ms | 3,792.25 ms | 5,032.8 ms | 8,134.09 ms |
| Darkbloomdarkbloom/fp8 | 3,738.5 ms | 6,412 ms | 10,592 ms | 25,852.35 ms |
| DeepInfradeepinfra/bf16 | 229 ms | 341 ms | 642 ms | 1,777.09 ms |
| DekaLLMdekallm/bf16 | 693 ms | 1,334 ms | 2,918 ms | 11,730.4 ms |
| Googlegoogle-vertex/us-central1 | 474 ms | 1,094.25 ms | 1,480.2 ms | 2,135.05 ms |
| Groqgroq | 717 ms | 1,244 ms | 1,836.9 ms | 4,003.09 ms |
| Novitanovita/fp4 | 633 ms | 756.75 ms | 962.4 ms | 2,406.51 ms |
| Parasailparasail/fp4 | 1,047 ms | 1,288 ms | 1,593 ms | 3,147.25 ms |
| PrimeIntellectprimeintellect | 479 ms | 694 ms | 1,005 ms | 2,897 ms |
| SiliconFlowsiliconflow/fp8 | 1,023.5 ms | 1,227.25 ms | 1,644.5 ms | 3,294.95 ms |
Output throughput · tokens per second
All throughput percentiles 13 endpoints
| Endpoint | P50 | P75 | P90 | P99 |
|---|---|---|---|---|
| AkashMLakashml/fp4 | 67 t/s | 91 t/s | 122 t/s | 214.9 t/s |
| Amazon Bedrockamazon-bedrock | 248 t/s | 419 t/s | 561 t/s | 904.26 t/s |
| Amazon Bedrockamazon-bedrock/eu-west-1 | 147.5 t/s | 181 t/s | 190.5 t/s | 254.8 t/s |
| CoreWeavecoreweave/fp4 | 246 t/s | 262 t/s | 276 t/s | 299.09 t/s |
| Darkbloomdarkbloom/fp8 | 20 t/s | 33 t/s | 48 t/s | 78.09 t/s |
| DeepInfradeepinfra/bf16 | 97 t/s | 112 t/s | 122 t/s | 140 t/s |
| DekaLLMdekallm/bf16 | 33 t/s | 40 t/s | 47.7 t/s | 70 t/s |
| Googlegoogle-vertex/us-central1 | 226 t/s | 295 t/s | 368 t/s | 453.08 t/s |
| Groqgroq | 405 t/s | 601 t/s | 742 t/s | 902.09 t/s |
| Novitanovita/fp4 | 157 t/s | 200 t/s | 238 t/s | 271 t/s |
| Parasailparasail/fp4 | 60 t/s | 85 t/s | 103 t/s | 124 t/s |
| PrimeIntellectprimeintellect | 180 t/s | 228 t/s | 251.4 t/s | 309.7 t/s |
| SiliconFlowsiliconflow/fp8 | 87 t/s | 124 t/s | 151.1 t/s | 171.81 t/s |
Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.
Uptime
Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.
Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.
Benchmarks
Review attributed evaluation results on each source's original scale.
openai/gpt-oss-20bObserved | Design Arena | Category | Elo | Win rate | Rank |
|---|---|---|---|---|
| models | dataviz | 940 | 39.7% | 115 |
| models | website | 859 | 27.9% | 129 |
Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.
Detailed evaluations
| Evaluation | Configuration | Source | Results | Details |
|---|---|---|---|---|
| models · website | GPT OSS 20Bopenai/gpt-oss-20b | Design Arena | 859 Elo | Details |
| tau bench verified airline | OpenAI: gpt-oss-20bopenai/gpt-oss-20b | OpenRouter | 53.391% | Details |
| models · dataviz | GPT OSS 20Bopenai/gpt-oss-20b | Design Arena | 940 Elo | Details |
| gpqa diamond | OpenAI: gpt-oss-20bopenai/gpt-oss-20b | OpenRouter | 64.399% | Details |
| Composite indices | gpt-oss-20b (high)openai/gpt-oss-20b | Artificial Analysis | Intelligence 9 · Coding 20.7 · Agentic 1.2 | Details |
Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.
Apps & session costs
Compare observed median session costs for applications using this model, grouped by conversation length.
| Application | Model route | Turns | Median cost / session | Details |
|---|---|---|---|---|
| Hermes Agent | openai/gpt-oss-20b | 1 turn | $0.000181 | Details |
| Hermes Agent | openai/gpt-oss-20b | 2–9 turns | $0.000992 | Details |
Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.
Activity
Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.
25,656,687,360 tokens
33,587,477,092 tokens
25,477,604,271 tokens
35,143,442,747 tokens
23,073,011,842 tokens
24,416,482,274 tokens
31,265,428,309 tokens
32,139,371,330 tokens
30,274,340,719 tokens
17,906,818,322 tokens
20,007,120,278 tokens
26,190,729,104 tokens
27,018,485,849 tokens
25,741,792,725 tokens
23,722,184,565 tokens
19,339,262,034 tokens
17,078,278,590 tokens
15,830,224,184 tokens
19,935,560,396 tokens
24,180,929,218 tokens
23,832,008,061 tokens
19,790,140,301 tokens
18,516,279,783 tokens
27,153,001,354 tokens
29,465,089,337 tokens
55,479,524,701 tokens
63,771,547,020 tokens
64,909,103,373 tokens
67,189,102,806 tokens
70,046,628,489 tokens
83,153,659,746 tokens
70,098,750,528 tokens
67,630,538,366 tokens
62,875,649,305 tokens
46,673,792,787 tokens
39,545,040,835 tokens
43,863,747,369 tokens
45,685,685,197 tokens
44,771,168,619 tokens
36,297,363,641 tokens
Daily observations · 40 days
| Date (UTC) | Reported tokens | Reported routes |
|---|---|---|
| Sep 9, 2026 | 36,297,363,641 | 1 |
| Sep 8, 2026 | 44,771,168,619 | 1 |
| Sep 7, 2026 | 45,685,685,197 | 1 |
| Sep 6, 2026 | 43,863,747,369 | 1 |
| Sep 5, 2026 | 39,545,040,835 | 1 |
| Sep 4, 2026 | 46,673,792,787 | 1 |
| Sep 3, 2026 | 62,875,649,305 | 1 |
| Sep 2, 2026 | 67,630,538,366 | 1 |
| Sep 1, 2026 | 70,098,750,528 | 1 |
| Aug 31, 2026 | 83,153,659,746 | 1 |
| Aug 30, 2026 | 70,046,628,489 | 1 |
| Aug 29, 2026 | 67,189,102,806 | 1 |
| Aug 28, 2026 | 64,909,103,373 | 1 |
| Aug 27, 2026 | 63,771,547,020 | 1 |
| Aug 26, 2026 | 55,479,524,701 | 1 |
| Aug 25, 2026 | 29,465,089,337 | 1 |
| Aug 24, 2026 | 27,153,001,354 | 1 |
| Aug 22, 2026 | 18,516,279,783 | 1 |
| Aug 16, 2026 | 19,790,140,301 | 1 |
| Aug 12, 2026 | 23,832,008,061 | 1 |
| Aug 11, 2026 | 24,180,929,218 | 1 |
| Aug 10, 2026 | 19,935,560,396 | 1 |
| Aug 9, 2026 | 15,830,224,184 | 1 |
| Aug 8, 2026 | 17,078,278,590 | 1 |
| Aug 7, 2026 | 19,339,262,034 | 1 |
| Aug 5, 2026 | 23,722,184,565 | 1 |
| Aug 4, 2026 | 25,741,792,725 | 1 |
| Aug 3, 2026 | 27,018,485,849 | 1 |
| Aug 2, 2026 | 26,190,729,104 | 1 |
| Aug 1, 2026 | 20,007,120,278 | 1 |
| Jul 26, 2026 | 17,906,818,322 | 1 |
| Jul 16, 2026 | 30,274,340,719 | 1 |
| Jul 15, 2026 | 32,139,371,330 | 1 |
| Jul 14, 2026 | 31,265,428,309 | 1 |
| Jul 13, 2026 | 24,416,482,274 | 1 |
| Jul 11, 2026 | 23,073,011,842 | 1 |
| Jul 10, 2026 | 35,143,442,747 | 1 |
| Jul 9, 2026 | 25,477,604,271 | 1 |
| Jul 8, 2026 | 33,587,477,092 | 1 |
| Jul 7, 2026 | 25,656,687,360 | 1 |
Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.
Technical details & API
Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.
openai/gpt-oss-20bObserved - Context
- 131,072
- Maximum output
- 32,768
- Tokenizer
- GPT
- Input
- text
- Output
- text
- Moderation
- No
- Knowledge cutoff
- 2024-06-30
- Added to OpenRouter
- 2025-08-05
https://openrouter.ai/api/v1openai/gpt-oss-20bopenai/gpt-oss-20bopenai/gpt-oss-20bReasoning
- Required
- Yes
- Enabled by default
- Not reported
- Default effort
- medium
- Supported efforts
- high, medium, low
Supported parameters
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Frequently asked questions
Answers based on the model metadata and serving offers currently in the catalog.
How much does GPT OSS 20B cost?
Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.
Which providers offer GPT OSS 20B?
There are 42 cataloged offers across 19 providers. Compare model IDs, prices and limits in Providers.
What is the context window?
The cataloged model context is 131,072 tokens. Each serving endpoint may apply a different limit.
Does it support tools and structured output?
Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.
Resources
Catalog history · 2 recorded revisions
Metadata observations since this model was first cataloged. These are not model release versions.
- ·
31b4f2cdfeb5 - ·
c69dfbe20065