Qwen3.6 35B-A3B

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license.

alibaba/qwen3.6-35b-a3b
Organization
Alibaba
Family
qwen
Providers
34
Context
262,144
Output limit
65,536
Knowledge
—
Release
2026-04-17
Updated
2026-04-17
Weights
Open
Input
text, image, video, audio
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

49 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
302.AI
qwen3.6-35b-a3b
262,14465,536$0.283 / $1.705YesYesYes
View
AIHubMix
qwen3.6-35b-a3b
262,14465,536$0.254 / $1.524YesYesYes
View
AKI.IO
qwen3.6-35b
256,00032,768$0.15 / $0.50YesYesYes
View
Blue Claw
Qwen/Qwen3.6-35B-A3B-FP8
131,07265,536— / —YesYesYes
View
CoreWeave
Qwen/Qwen3.6-35B-A3B
262,144262,144$0.25 / $1.25YesYesYes
View
Cortecs
qwen3.6-35b-a3b
262,00032,768$0.167 / $0.557YesYesYes
View
DevPass (LLM Gateway)
qwen3.6-35b-a3b
262,14465,536$0.248 / $1.485YesYesYes
View
EmpirioLabs AI
qwen3-6-35b-a3b
131,07216,384$0.07 / $0.42YesYesYes
View
engy
qwen3.6-35b-a3b
208,1928,192$0.045 / $0.30YesYesYes
View
evroc
Qwen/Qwen3.6-35B-A3B
262,14465,536$0.345 / $1.38YesYesYes
View
GreenPT
qwen3.6-35b-a3b
262,14432,768$0.342 / $2.052YesYesYes
View
Hetzner
Qwen/Qwen3.6-35B-A3B-FP8
262,144262,144$0.00 / $0.00YesYesYes
View
Hugging Face
Qwen/Qwen3.6-35B-A3B
262,14465,536$0.15 / $0.95YesYesYes
View
InferX
Qwen3.6-35B-A3B-FP8
262,00065,536$0.00 / $0.00YesYesYes
View
InferX
Qwen3.6-35B-A3B-fp8-no-thinking
262,00065,536$0.00 / $0.00NoYesYes
View
Kilo Gateway
qwen/qwen3.6-35b-a3b
262,144235,929$0.15 / $1.00YesYesYes
View
LLM Gateway
alibaba/qwen3.6-35b-a3b
262,14465,536$0.375 / $2.25YesYesNo
View
LLM Gateway
novita/qwen3.6-35b-a3b
262,14464,000$0.248 / $1.485YesYesNo
View
LLMTR
qwen3-6-35b
16,38416,384$5.00 / $10.00YesNoYes
View
Merge Gateway
qwen/qwen3.6-35b-a3b
262,14465,536$0.248 / $1.485YesYesYes
View
NaN
qwen3.6
262,14465,536$0.00 / $0.00YesYesYes
View
NanoGPT
qwen/qwen3.6-35b-a3b
262,14416,384$0.112 / $0.80NoYesNo
View
NanoGPT
qwen/qwen3.6-35b-a3b:thinking
262,14416,384$0.112 / $0.80YesYesNo
View
NanoGPT
TEE/qwen3.6-35b-a3b
262,144262,144$0.20 / $1.27NoNoNo
View
NEAR AI Cloud
Qwen/Qwen3.6-35B-A3B-FP8
262,1448,192$0.17 / $1.10YesYesYes
View
Neuralwatt
qwen3.6-35b-fast
131,056131,056$0.29 / $1.15NoYesYes
View
Neuralwatt
qwen3.6-35b
131,056131,056$0.29 / $1.15YesYesYes
View
Neuralwatt
qwen3.6-35b-flex
131,056131,056$0.1885 / $0.7475YesYesYes
View
AkashMLvia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,144235,929$0.10 / $0.90YesYesYes
View
AtlasCloudvia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,14465,536$0.186 / $1.11375YesYesNo
View
CoreWeavevia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,144235,929$0.25 / $1.25YesNoYes
View
Darkbloomvia OpenRouter
qwen/qwen3.6-35b-a3b
fp4
262,14432,768$0.05 / $0.70YesYesYes
View
DeepInfravia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,14465,536$0.10 / $0.95YesYesYes
View
Parasailvia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,144235,929$0.15 / $1.00YesYesYes
View
Phalavia OpenRouter
qwen/qwen3.6-35b-a3b
unknown
262,144235,929$0.20 / $1.27YesYesYes
View
SiliconFlowvia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
262,144235,929$0.24 / $1.80YesNoYes
View
Venicevia OpenRouter
qwen/qwen3.6-35b-a3b
fp8
256,00065,536$0.10 / $1.00YesYesNo
View
Opper
qwen3.6-35b-a3b
262,14465,536$0.248 / $1.485YesYesYes
View
OrcaRouter
qwen/qwen3.6-35b-a3b
262,14465,536$0.248 / $1.485YesYesYes
View
Pioneer
Qwen/Qwen3.6-35B-A3B
262,144131,072$0.14 / $1.00YesYesYes
View
QVAC
qwen3.6-35b-a3b
262,14465,536$0.00 / $0.00YesYesYes
View
SaladCloud AI Gateway
qwen3.6-35b-a3b
262,144262,144$0.09 / $0.60YesYesYes
View
Scaleway
qwen3.6-35b-a3b
128,00016,384$0.25 / $1.50YesYesYes
View
SiliconFlow
Qwen/Qwen3.6-35B-A3B
262,144262,144$0.20 / $1.60NoYesYes
View
Umans AI
umans-flash
262,14432,768$0.15 / $1.00YesYesYes
View
Umans AI Coding Plan
umans-flash
262,144262,144$0.00 / $0.00YesYesYes
View
Umans AI Coding Plan
umans-qwen3.6-35b-a3b
262,144262,144$0.00 / $0.00YesYesYes
View
Venice AI
qwen3-6-35b-a3b
256,00065,536$0.10 / $1.00YesYesYes
View
Zenifra
alibaba/qwen3.6-35b-a3b
262,14465,536$0.19 / $0.48YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers49Across 34 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 9 lowest-priced configurations
InputOutput
Darkbloomdarkbloom/fp4
$0.05$0.70
AkashMLakashml/fp8
$0.10$0.90
DeepInfradeepinfra/fp8
$0.10$0.95
Venicevenice/fp8
$0.10$1.00
Parasailparasail/fp8
$0.15$1.00
AtlasCloudatlas-cloud/fp8
$0.186$1.11375
Phalaphala
$0.20$1.27
SiliconFlowsiliconflow/fp8
$0.24$1.80
CoreWeavecoreweave/fp8
$0.25$1.25
qwen/qwen3.6-35b-a3bObserved
Input
$0.15
Output
$1.00
Cache read
$0.05
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
AkashMLakashml/fp8$0.10$0.90$0.05—
AtlasCloudatlas-cloud/fp8$0.186$1.11375$0.186—
CoreWeavecoreweave/fp8$0.25$1.25$0.25—
Darkbloomdarkbloom/fp4$0.05$0.70$0.025—
DeepInfradeepinfra/fp8$0.10$0.95$0.10—
Parasailparasail/fp8$0.15$1.00$0.05—
Phalaphala$0.20$1.27$0.056—
SiliconFlowsiliconflow/fp8$0.24$1.80$0.15—
Venicevenice/fp8$0.10$1.00——

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AkashMLakashml/fp8
422–1,612.9 ms
AtlasCloudatlas-cloud/fp8
1,386–5,627.64 ms
CoreWeavecoreweave/fp8
463.5–5,703.59 ms
Darkbloomdarkbloom/fp4
944–12,661.22 ms
DeepInfradeepinfra/fp8
833–6,648.7 ms
Parasailparasail/fp8
557–1,647.54 ms
All latency percentiles 9 endpoints
EndpointP50P75P90P99
AkashMLakashml/fp8422 ms537 ms703 ms1,612.9 ms
AtlasCloudatlas-cloud/fp81,386 ms1,805.25 ms2,418.8 ms5,627.64 ms
CoreWeavecoreweave/fp8463.5 ms1,905.75 ms3,071.7 ms5,703.59 ms
Darkbloomdarkbloom/fp4944 ms1,841 ms4,860.6 ms12,661.22 ms
DeepInfradeepinfra/fp8833 ms1,376.75 ms2,462 ms6,648.7 ms
Parasailparasail/fp8557 ms693 ms847 ms1,647.54 ms
Phalaphala1,793.5 ms2,184 ms2,635.2 ms11,831.15 ms
SiliconFlowsiliconflow/fp81,456 ms3,268.75 ms4,807.5 ms8,490.12 ms
Venicevenice/fp8625 ms730.5 ms914.6 ms1,685.68 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AkashMLakashml/fp8
134–364.78 t/s
AtlasCloudatlas-cloud/fp8
84.5–183 t/s
CoreWeavecoreweave/fp8
136–197.46 t/s
Darkbloomdarkbloom/fp4
55–140 t/s
DeepInfradeepinfra/fp8
129–213 t/s
Parasailparasail/fp8
76–145.8 t/s
All throughput percentiles 9 endpoints
EndpointP50P75P90P99
AkashMLakashml/fp8134 t/s180.5 t/s249.8 t/s364.78 t/s
AtlasCloudatlas-cloud/fp884.5 t/s116 t/s139.7 t/s183 t/s
CoreWeavecoreweave/fp8136 t/s167.5 t/s178 t/s197.46 t/s
Darkbloomdarkbloom/fp455 t/s80 t/s102 t/s140 t/s
DeepInfradeepinfra/fp8129 t/s160 t/s185 t/s213 t/s
Parasailparasail/fp876 t/s93 t/s108 t/s145.8 t/s
Phalaphala40 t/s52 t/s102 t/s199 t/s
SiliconFlowsiliconflow/fp875.5 t/s94.75 t/s111.4 t/s146.35 t/s
Venicevenice/fp8150 t/s199 t/s226 t/s279.93 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 9 endpoints
Endpoint5 minutes30 minutes24 hours
AkashMLakashml/fp8
AtlasCloudatlas-cloud/fp8
CoreWeavecoreweave/fp8
Darkbloomdarkbloom/fp4
DeepInfradeepinfra/fp8
Parasailparasail/fp8
Phalaphala
SiliconFlowsiliconflow/fp8
Venicevenice/fp8

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
qwen/qwen3.6-35b-a3bObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 50
Coding41.9
Agentic13.1
Intelligence18.2

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
tau bench verified airlineQwen: Qwen3.6 35B A3B
71.1%
gpqa diamondQwen: Qwen3.6 35B A3B
79.477%
EvaluationConfigurationSourceResultsDetails
Composite indicesQwen3.6 35B A3B (Reasoning)qwen/qwen3.6-35b-a3b-20260415Artificial AnalysisIntelligence 18.2 · Coding 41.9 · Agentic 13.1
Details
tau bench verified airlineQwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b-20260415OpenRouter71.1%
Details
gpqa diamondQwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b-20260415OpenRouter79.477%
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Hermes Agentqwen/qwen3.6-35b-a3b-202604152–9 turns
$0.007487
Details
Hermes Agentqwen/qwen3.6-35b-a3b-2026041510–49 turns
$0.063617
Details
Hermes Agentqwen/qwen3.6-35b-a3b-202604151 turn
$0.001946
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage154,576,743,616 reported tokens · 6 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 6 days
Date (UTC)Reported tokensReported routes
Aug 16, 202634,020,667,6711
Aug 13, 202635,160,314,7061
Aug 10, 202625,655,613,7721
Aug 9, 202626,659,873,3151
Aug 8, 202615,948,658,4711
Jul 18, 202617,131,615,6811

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
qwen/qwen3.6-35b-a3bObserved
Context
262,144
Maximum output
235,929
Tokenizer
Qwen
Input
text, image, video
Output
text
Moderation
No
Added to OpenRouter
2026-04-27
Base URLhttps://openrouter.ai/api/v1
Model IDqwen/qwen3.6-35b-a3b
Version IDqwen/qwen3.6-35b-a3b-20260415
Weights IDQwen/Qwen3.6-35B-A3B

Reasoning

Required
No
Enabled by default
Yes
Default effort
—
Supported efforts
—

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Default parameters

top k
20
top p
0.95
temperature
1

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does Qwen3.6 35B-A3B cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer Qwen3.6 35B-A3B?

There are 49 cataloged offers across 34 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 262,144 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Resources

Catalog history · 2 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · 33a7a88c5273
  • · 5eeaa36e40c5