Qwen3.8 Flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

alibaba/qwen3.8-flash
Organization
Alibaba
Family
qwen
Providers
27
Context
1,000,000
Output limit
131,072
Knowledge
—
Release
2026-08-26
Updated
2026-08-26
Weights
Closed
Input
text, image, video
Output
text
Capabilities
Tools, Reasoning, Structured

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

29 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
302.AI
qwen3.8-flash
1,000,000131,072$0.18 / $0.564YesYesYes
View
AIHubMix
qwen3.8-flash
1,000,000131,072$0.1126 / $0.380025YesYesYes
View
Alibaba
qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
Alibaba (China)
qwen3.8-flash
1,000,000131,072$0.11875 / $0.40073YesYesYes
View
Alibaba Token Plan
qwen3.8-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Alibaba Token Plan (China)
qwen3.8-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Charm Hyper
qwen3.8-flash
1,000,000128,000$0.15 / $0.47YesYesYes
View
CrossModel
qwen/qwen3.8-flash
1,000,000131,072$0.13 / $0.43YesYesYes
View
Deep Infra
Qwen/Qwen3.8-Flash
1,000,000131,072$0.113 / $0.382NoYesYes
View
DevPass (LLM Gateway)
qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
Eden AI
qwen/qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
EmpirioLabs AI
qwen3-8-flash
1,000,000131,072$0.16 / $0.47YesYesYes
View
GMI Cloud
Qwen/Qwen3.8-Flash
1,048,575131,072$0.16 / $0.47YesYesYes
View
Kilo Gateway
qwen/qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
LLM Gateway
novita/qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
LLM Gateway
alibaba/qwen3.8-flash
983,616131,072$0.15 / $0.47YesYesYes
View
Merge Gateway
qwen/qwen3.8-flash
977,000128,000$0.15 / $0.47YesYesYes
View
NaN
qwen3.8-flash
262,144131,072$0.00 / $0.00YesYesYes
View
NanoGPT
qwen/qwen3.8-flash
991,808131,072$0.14 / $0.42YesYesYes
View
Ofox
bailian/qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
Ofox
qwen/qwen3.8-flash
1,000,000131,072$0.11 / $0.39YesYesYes
View
OpenCode Go
qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
OpenCode Zen
qwen3.8-flash
1,000,000131,072$0.15 / $0.47YesYesYes
View
Alibabavia OpenRouter
qwen/qwen3.8-flash
unknown
1,000,000131,072$0.15 / $0.47YesYesYes
View
Requesty
qwen3.8-flash
1,048,576131,072$0.16 / $0.47YesYesYes
View
SCNet Token Plan
Qwen3.8-Flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Vancine
qwen3.8-flash
1,000,000131,072$0.12 / $0.38YesYesYes
View
Venice AI
qwen-3-8-flash
1,000,000131,072$0.14 / $0.49YesYesYes
View
Vercel AI Gateway
alibaba/qwen3.8-flash
991,000128,000$0.15 / $0.47YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers29Across 27 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 1 lowest-priced configuration
InputOutput
Alibabaalibaba
$0.15$0.47
qwen/qwen3.8-flashObserved
Input
$0.15
Output
$0.47
Cache read
$0.016
Cache write
$0.20
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
Alibabaalibaba$0.15$0.47$0.016$0.20

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 1 endpoint
P50P75P90P99
Alibabaalibaba
1,748.5–10,549.9 ms
All latency percentiles 1 endpoints
EndpointP50P75P90P99
Alibabaalibaba1,748.5 ms2,660 ms4,070 ms10,549.9 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 1 endpoint
P50P75P90P99
Alibabaalibaba
60–100 t/s
All throughput percentiles 1 endpoints
EndpointP50P75P90P99
Alibabaalibaba60 t/s78 t/s88 t/s100 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 1 endpoint
Endpoint5 minutes30 minutes24 hours
Alibabaalibaba

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
gpqa diamondQwen: Qwen3.8 Flash
88.552%
tau bench verified airlineQwen: Qwen3.8 Flash
71.333%
EvaluationConfigurationSourceResultsDetails
gpqa diamondQwen: Qwen3.8 Flashqwen/qwen3.8-flash-20260826OpenRouter88.552%
Details
tau bench verified airlineQwen: Qwen3.8 Flashqwen/qwen3.8-flash-20260826OpenRouter71.333%
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Claude Codeqwen/qwen3.8-flash-2026082650+ turns
$0.535062
Details
Claude Codeqwen/qwen3.8-flash-2026082610–49 turns
$0.063005
Details
Hermes Agentqwen/qwen3.8-flash-202608262–9 turns
$0.003935
Details
Codexqwen/qwen3.8-flash-202608262–9 turns
$0.004954
Details
Claude Codeqwen/qwen3.8-flash-202608262–9 turns
$0.008346
Details
Codexqwen/qwen3.8-flash-202608261 turn
$0.000879
Details
Claude Codeqwen/qwen3.8-flash-202608261 turn
$0.003875
Details
Kilo Codeqwen/qwen3.8-flash-2026082610–49 turns
$0.04096
Details
Hermes Agentqwen/qwen3.8-flash-2026082610–49 turns
$0.031692
Details
Hermes Agentqwen/qwen3.8-flash-2026082650+ turns
$0.205829
Details
Hermes Agentqwen/qwen3.8-flash-202608261 turn
$0.000274
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage1,728,221,115,219 reported tokens · 27 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 27 days
Date (UTC)Reported tokensReported routes
Oct 3, 202647,264,120,1761
Oct 2, 202689,774,703,3141
Oct 1, 202685,041,599,4521
Sep 30, 202671,062,223,9351
Sep 29, 202687,184,008,4091
Sep 28, 202669,726,092,3821
Sep 27, 202664,184,693,3281
Sep 26, 202652,527,469,3061
Sep 25, 202688,557,968,6431
Sep 24, 202675,469,250,3901
Sep 23, 202673,662,708,2371
Sep 22, 202684,259,131,2101
Sep 21, 202670,076,539,9061
Sep 20, 202662,268,094,7571
Sep 19, 202656,721,757,5891
Sep 18, 202664,882,982,0981
Sep 17, 202662,340,399,4801
Sep 16, 202670,750,478,1141
Sep 15, 202668,620,347,2911
Sep 14, 202674,534,474,9141
Sep 13, 202680,566,273,7331
Sep 12, 202653,097,161,1901
Sep 11, 202650,817,088,5601
Sep 6, 202633,167,525,5021
Sep 5, 202636,771,457,3161
Aug 30, 202629,581,316,3881
Aug 29, 202625,311,249,5991

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
qwen/qwen3.8-flashObserved
Context
1,000,000
Maximum output
131,072
Tokenizer
Qwen
Input
text, image, video
Output
text
Moderation
No
Added to OpenRouter
2026-08-26
Base URLhttps://openrouter.ai/api/v1
Model IDqwen/qwen3.8-flash
Version IDqwen/qwen3.8-flash-20260826
Weights IDQwen/Qwen3.8-Flash-Next

Reasoning

Required
No
Enabled by default
Yes
Default effort
—
Supported efforts
—

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logprobs
  • max_tokens
  • presence_penalty
  • reasoning
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Qwen3.8 Flash cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Qwen3.8 Flash?

There are 29 cataloged offers across 27 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 1,000,000 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · 21661782240f
  • · 87bc3b81a17a