Qwen3.7 Flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.

alibaba/qwen3.7-flash
Organization
Alibaba
Family
qwen
Providers
12
Context
1,000,000
Output limit
131,072
Knowledge
—
Release
2026-07-15
Updated
2026-07-15
Weights
Closed
Input
text, image, video
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

13 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
AIHubMix
qwen3.7-flash
991,00064,000$0.0282 / $0.1128YesYesYes
View
Alibaba
qwen3.7-flash
1,000,000131,072$0.03 / $0.13YesYesYes
View
Alibaba (China)
qwen3.7-flash
1,000,000131,072$0.02962 / $0.1185YesYesYes
View
Charm Hyper
qwen3.7-flash
1,000,00064,000$0.20 / $0.80YesYesYes
View
CrossModel
qwen/qwen3.7-flash
1,000,00065,536$0.04 / $0.13YesYesYes
View
EmpirioLabs AI
qwen3-7-flash
1,000,00065,536$0.03 / $0.13YesYesYes
View
Kilo Gateway
qwen/qwen3.7-flash
1,000,00065,536$0.03 / $0.13YesYesNo
View
LLM Gateway
alibaba/qwen3.7-flash
983,61665,536$0.03 / $0.13YesYesNo
View
NanoGPT
qwen/qwen3.7-flash
991,80865,536$0.03 / $0.13YesYesYes
View
NanoGPT
qwen/qwen3.7-flash:thinking
983,61665,536$0.03 / $0.13YesYesYes
View
Alibabavia OpenRouter
qwen/qwen3.7-flash
unknown
1,000,00065,536$0.03 / $0.13YesYesNo
View
OrcaRouter
qwen/qwen3.7-flash
1,000,000131,072$0.03 / $0.13YesYesYes
View
Vercel AI Gateway
alibaba/qwen3.7-flash
991,00064,000$0.03 / $0.13YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.0282Per 1M input tokens
Output from$0.1128Per 1M output tokens
Available offers13Across 12 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 1 lowest-priced configuration
InputOutput
Alibabaalibaba
$0.03$0.13
qwen/qwen3.7-flashObserved
Input
$0.03
Output
$0.13
Cache read
$0.006
Cache write
$0.038
Additional fees and pricing conditions

overrides · USD per token for token prices; time windows use UTC.

[
  {
    "prompt": "0.0000001",
    "completion": "0.0000004",
    "input_cache_read": "0.00000002",
    "input_cache_write": "0.000000125",
    "min_prompt_tokens": 32000
  },
  {
    "prompt": "0.0000002",
    "completion": "0.0000008",
    "input_cache_read": "0.00000004",
    "input_cache_write": "0.00000025",
    "min_prompt_tokens": 256000
  }
]
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
Alibabaalibaba$0.03$0.13$0.006$0.038

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 1 endpoint
P50P75P90P99
Alibabaalibaba
624–4,832.6 ms
All latency percentiles 1 endpoints
EndpointP50P75P90P99
Alibabaalibaba624 ms819 ms1,545 ms4,832.6 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 1 endpoint
P50P75P90P99
Alibabaalibaba
78–149 t/s
All throughput percentiles 1 endpoints
EndpointP50P75P90P99
Alibabaalibaba78 t/s107 t/s126 t/s149 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 1 endpoint
Endpoint5 minutes30 minutes24 hours
Alibabaalibaba

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Claude Codeqwen/qwen3.7-flash-202607272–9 turns
$0.002247
Details
Hermes Agentqwen/qwen3.7-flash-202607271 turn
$0.000383
Details
Hermes Agentqwen/qwen3.7-flash-2026072710–49 turns
$0.026999
Details
Hermes Agentqwen/qwen3.7-flash-202607272–9 turns
$0.001329
Details
Claude Codeqwen/qwen3.7-flash-2026072710–49 turns
$0.034886
Details
Hermes Agentqwen/qwen3.7-flash-2026072750+ turns
$0.223373
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage1,803,042,799,851 reported tokens · 44 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 44 days
Date (UTC)Reported tokensReported routes
Oct 3, 202647,936,207,7491
Oct 2, 202654,987,015,3171
Oct 1, 202665,122,446,9511
Sep 29, 202653,064,199,5611
Sep 28, 202666,036,338,8591
Sep 27, 202666,671,634,4211
Sep 26, 202657,437,604,7971
Sep 25, 202666,890,138,4361
Sep 24, 202666,670,818,3221
Sep 23, 202668,423,870,0311
Sep 22, 202659,716,283,9011
Sep 21, 202655,183,797,2301
Sep 20, 202661,020,712,0341
Sep 19, 202658,698,340,7981
Sep 18, 202661,930,342,0251
Sep 17, 202652,118,079,6741
Sep 16, 202662,624,589,7091
Sep 15, 202650,850,643,2981
Sep 14, 202653,582,123,3981
Sep 13, 202642,198,216,5601
Sep 12, 202644,530,832,6941
Sep 11, 202648,082,456,4091
Sep 10, 202642,062,634,1021
Sep 9, 202634,786,923,5011
Aug 30, 202624,193,735,7401
Aug 29, 202624,321,568,0641
Aug 27, 202626,327,854,4621
Aug 26, 202626,368,988,9371
Aug 22, 202620,423,440,1751
Aug 20, 202630,476,116,0791
Aug 19, 202630,690,265,2391
Aug 18, 202630,466,319,1601
Aug 16, 202622,543,941,4501
Aug 15, 202624,276,998,0891
Aug 11, 202620,638,211,8881
Aug 10, 202619,814,104,2431
Aug 9, 202615,799,103,4861
Aug 8, 202616,582,240,1451
Aug 7, 202619,074,665,1071
Aug 4, 202626,596,207,0441
Aug 2, 202615,881,464,0031
Aug 1, 202616,181,765,9681
Jul 30, 202626,254,246,4311
Jul 29, 202625,505,314,3641

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
qwen/qwen3.7-flashObserved
Context
1,000,000
Maximum output
65,536
Tokenizer
Qwen
Input
text, image, video
Output
text
Moderation
No
Added to OpenRouter
2026-07-27
Base URLhttps://openrouter.ai/api/v1
Model IDqwen/qwen3.7-flash
Version IDqwen/qwen3.7-flash-20260727

Reasoning

Required
No
Enabled by default
Yes
Default effort
—
Supported efforts
—

Supported parameters

  • include_reasoning
  • logprobs
  • max_tokens
  • presence_penalty
  • reasoning
  • response_format
  • seed
  • temperature
  • tool_choice
  • tools
  • top_logprobs
  • top_p

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Qwen3.7 Flash cost?

Listed rates start at $0.0282 per million input tokens and $0.1128 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Qwen3.7 Flash?

There are 13 cataloged offers across 12 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 1,000,000 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Catalog history · 3 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · 33a15c130a04
  • · 0820451f4867
  • · 5d50b1c408c0