GLM-5.3-Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

zhipuai/glm-5.3-flash
Organization
Zhipu AI
Family
glm-flash
Providers
71
Context
1,000,000
Output limit
131,072
Knowledge
—
Release
2026-08-26
Updated
2026-08-26
Weights
Open
Input
text, image, video, pdf
Output
text
Capabilities
Tools, Reasoning, Structured, Temperature

Providers

Compare every cataloged offer for this model, including its provider, route, token limits, price, and capabilities.

127 offers · USD per million tokens
ProviderModel IDContextOutputInput / output · 1MReasoningToolsStructuredDetails
302.AI
glm-5.3-flash
1,000,000131,072$0.075 / $0.25YesYesYes
View
above.dev
glm-5.3-flash
1,000,000131,072$0.165 / $0.55YesYesYes
View
ai&
zai-org/glm-5.3-flash
1,048,550131,072$0.15 / $0.50YesYesYes
View
AIHubMix
glm-5.3-flash
1,000,000128,000$0.11268 / $0.39438YesYesYes
View
AIHubMix
ox-alpha
1,000,000131,072$0.00 / $0.00YesYesYes
View
Baseten
zai-org/GLM-5.3-Flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
Berget.AI
zai-org/GLM-5.3-Flash
524,28816,384$0.29 / $0.58YesYesYes
View
Bothub
glm-5.3-flash
1,000,000131,072$0.12 / $0.44YesYesYes
View
Charm Hyper
glm-5.3-flash
1,048,576131,072$0.16332 / $0.5444YesYesYes
View
Cloudflare Workers AI
@cf/zai-org/glm-5.3-flash
1,048,5761,048,576$0.15 / $0.50YesYesYes
View
CoralBricks
glm-5.3-flash-fp4
1,048,576131,072$0.15 / $0.50YesYesYes
View
CoreWeave
zai-org/GLM-5.3-Flash
1,048,5761,048,576$0.15 / $0.50YesYesYes
View
Cortecs
glm-5.3-flash
1,048,5761,048,576$0.10 / $0.35YesYesYes
View
CrofAI
glm-5.3-flash
1,000,000131,072$0.07 / $0.22YesYesYes
View
CrossModel
z-ai/glm-5.3-flash
1,000,000128,000$0.15 / $0.50YesYesYes
View
Deep Infra
zai-org/GLM-5.3-Flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
DevPass (LLM Gateway)
glm-5.3-flash
1,048,576131,072$0.088 / $0.25YesYesYes
View
DigitalOcean
glm-5.3-flash
1,048,5761,048,576$0.15 / $0.50YesYesYes
View
Eden AI
zai/glm-5.3-flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
EmpirioLabs AI
glm-5-3-flash
1,000,000131,072$0.075 / $0.25YesYesYes
View
engy
glm-5.3-flash
262,14432,768$0.135 / $0.45YesYesYes
View
Fireworks AI
accounts/fireworks/models/glm-5p3-flash
1,048,573131,072$0.15 / $0.50YesYesYes
View
Fireworks AI
accounts/fireworks/routers/glm-flash-latest
1,048,573131,072$0.15 / $0.50YesYesYes
View
Friendli
zai-org/GLM-5.3-Flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
GreenPT
glm-5.3-flash
1,000,000131,072$0.127754 / $0.511016YesYesYes
View
Hugging Face
zai-org/GLM-5.3-Flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
Inco
glm-5.3-flash:fast
1,000,000131,072$0.15 / $0.50YesYesYes
View
IteraCompute
z-ai/glm-5.3-flash
1,048,576131,072$0.14 / $0.49YesYesNo
View
Kenari
glm-5-3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Kilo Gateway
z-ai/glm-5.3-flash
1,048,575943,717$0.15 / $0.50YesYesYes
View
LLM Gateway
zai/glm-5.3-flash
1,048,576131,072$0.15 / $0.50YesYesNo
View
LLM Gateway
novita/glm-5.3-flash
1,048,576131,072$0.15 / $0.50YesYesNo
View
LLM Gateway
fireworks/glm-5.3-flash
1,048,576131,072$0.15 / $0.50YesYesNo
View
LLM Gateway
scx-ai-gp/glm-5.3-flash
1,048,576131,072$0.088 / $0.25YesYesNo
View
LLM Gateway
gonka24/glm-5.3-flash
200,00016,384$0.15 / $0.30YesYesYes
View
LLM Gateway
consensusprotocol/glm-5.3-flash
1,048,576131,072$0.10 / $0.25YesYesNo
View
LLM Gateway
inference.net/glm-5.3-flash
1,048,576128,000$0.09 / $0.28YesYesYes
View
LLM Gateway
runware/glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
LLM Gateway
together-ai/glm-5.3-flash
1,048,576943,717$0.15 / $0.50YesYesYes
View
Melious
glm-5.3-flash
1,000,000131,072$0.11592 / $0.46368YesYesYes
View
Merge Gateway
zai/glm-5.3-flash
1,000,000131,072$0.075 / $0.25YesYesYes
View
Modal
zai-org/GLM-5.3-Flash
1,000,000131,072$0.45 / $1.50YesYesYes
View
NaN
glm5.3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
NanoGPT
z-ai/glm-5.3-flash
1,048,576131,072$0.10 / $0.30YesYesYes
View
NanoGPT
TEE/glm-5.3-flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
Nebius Token Factory
zai-org/GLM-5.3-Flash
1,024,0001,024,000$0.15 / $0.50YesYesYes
View
Neon
glm-5-3-flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
Neuralwatt
glm-5.3-flash
1,048,5601,048,560$0.15 / $0.50YesYesNo
View
Neuralwatt
glm-5.3-flash-flex
1,048,5601,048,560$0.0975 / $0.325YesYesNo
View
Nvidia
z-ai/glm-5.3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Ofox
z-ai/glm-5.3-flashx
1,000,000131,072$0.15 / $0.50YesYesYes
View
Ofox
z-ai/glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
Ollama Cloud
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
OpenCode Go
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
OpenCode Zen
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
AtlasCloudvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576131,072$0.15 / $0.50YesYesNo
View
BaseTenvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576131,072$0.15 / $0.50YesYesYes
View
Cloudflarevia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.30 / $1.00YesYesYes
View
CoreWeavevia OpenRouter
z-ai/glm-5.3-flash
nvfp4
1,048,576943,718$0.15 / $0.50YesYesYes
View
Crusoevia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576943,718$0.15 / $0.50YesYesYes
View
Decartvia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576943,718$0.1125 / $0.375YesYesYes
View
DeepInfravia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576131,072$0.075 / $0.25YesYesYes
View
DekaLLMvia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.10 / $1.00YesYesYes
View
DigitalOceanvia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.15 / $0.50YesYesYes
View
Fireworksvia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.15 / $0.50YesYesYes
View
Fireworksvia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.225 / $0.75YesYesYes
View
Friendlivia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,718$0.15 / $0.50YesYesYes
View
GMICloudvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576943,718$0.09 / $0.30YesYesNo
View
Inceptronvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576943,718$0.225 / $0.45YesYesYes
View
InferenceNetvia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576262,144$0.044 / $0.60YesYesYes
View
Io Netvia OpenRouter
z-ai/glm-5.3-flash
fp8
262,12465,536$0.1125 / $0.375YesYesNo
View
Modalvia OpenRouter
z-ai/glm-5.3-flash
nvfp4
1,048,576943,718$0.15 / $0.50YesYesYes
View
Morphvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576943,718$0.16 / $0.444YesYesYes
View
Near AIvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576943,718$0.105 / $0.35YesYesYes
View
NextBitvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576128,000$0.165 / $0.55YesYesYes
View
Novitavia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576131,072$0.084 / $0.28YesYesNo
View
OpenInferencevia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576943,718$0.04576 / $0.65YesYesYes
View
Parasailvia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576943,718$0.15 / $0.50YesYesYes
View
Phalavia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576131,072$0.12 / $0.40YesYesYes
View
PrimeIntellectvia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576943,717$0.15 / $0.50YesYesYes
View
Rekavia OpenRouter
z-ai/glm-5.3-flash
unknown
262,144235,929$0.15 / $0.50YesYesYes
View
Relacevia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576131,072$0.0352 / $0.50YesYesNo
View
Sail Researchvia OpenRouter
z-ai/glm-5.3-flash
fp4
1,048,576131,072$0.045 / $0.60YesYesYes
View
SiliconFlowvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576262,144$0.15 / $0.50YesYesNo
View
StreamLakevia OpenRouter
z-ai/glm-5.3-flash
fp8
1,024,000128,000$0.08385 / $0.2795YesYesNo
View
Togethervia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,575943,717$0.15 / $0.50YesYesYes
View
Venicevia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576131,072$0.15 / $0.50YesYesYes
View
Wafervia OpenRouter
z-ai/glm-5.3-flash
unknown
1,048,576131,072$0.10 / $0.50YesYesYes
View
Z.AIvia OpenRouter
z-ai/glm-5.3-flash
fp8
1,048,576131,072$0.15 / $0.50YesYesNo
View
Opper
glm-5.3-flash
1,000,000128,000$0.20 / $0.50YesYesNo
View
OrcaRouter
z-ai/glm-5.3-flash-free
1,000,000128,000$0.00 / $0.00YesYesYes
View
OrcaRouter
z-ai/glm-5.3-flash
1,000,000128,000$0.075 / $0.25YesYesYes
View
Pareto Inference
z-ai/glm-5.3-flash
1,000,000131,072$0.03 / $0.10YesYesYes
View
Privatemode AI
glm-5.3-flash
1,000,000131,072$0.2311 / $0.7511YesYesYes
View
Privatemode AI
glm-flash-latest
1,000,000131,072$0.2311 / $0.7511YesYesYes
View
Requesty
glm-5.3-flash@eu
1,000,000262,144$0.20 / $0.60YesYesYes
View
Requesty
glm-5.3-flash
1,000,000262,144$0.20 / $0.60YesYesYes
View
RunInfra
zai-org/GLM-5.3-Flash
1,048,57632,768$0.10 / $0.40YesYesYes
View
SCNet Token Plan
GLM-5.3-Flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
SiliconFlow
zai-org/GLM-5.3-Flash
1,049,000262,000$0.15 / $0.50YesYesYes
View
Synthetic
hf:zai-org/GLM-5.3-Flash
524,28865,536$0.15 / $0.50YesYesYes
View
Tempr Gateway
zai/glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
Tempr Gateway
zai/glm-5.3-flashx
1,000,000131,072$0.37 / $1.25YesYesYes
View
TensorX
z-ai/glm-5.3-flash
1,048,57664,000$0.20 / $0.50YesYesYes
View
Tinfoil
glm-5-3-flash
1,048,576131,072$0.40 / $1.25YesYesYes
View
Together AI
zai-org/GLM-5.3-Flash
1,048,575400,000$0.15 / $0.50YesYesYes
View
TokenGo
z-ai/glm-5.3-flash
1,000,000131,072$0.075 / $0.025YesYesYes
View
Umans AI
umans-glm-5.3-flash
1,048,576131,071$0.15 / $0.50YesYesYes
View
Umans AI
umans-coder
1,048,576131,071$0.15 / $0.50YesYesYes
View
Umans AI Coding Plan
umans-coder
1,048,576131,071$0.00 / $0.00YesYesYes
View
Umans AI Coding Plan
umans-glm-5.3-flash
1,048,576131,071$0.00 / $0.00YesYesYes
View
Vancine
glm-5.3-flash
1,000,000131,072$0.12 / $0.40YesYesYes
View
Venice AI
z-ai-glm-5-3-flash
1,048,576131,072$0.15 / $0.50YesYesYes
View
Vercel AI Gateway
zai/glm-5.3-flashx
1,000,000131,072$0.37 / $1.25YesYesYes
View
Vercel AI Gateway
zai/glm-5.3-flash
1,000,000131,000$0.15 / $0.50YesYesYes
View
Vivgrid
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
Volcengine Ark
glm-5-3-flash-260828
1,000,000131,072$0.11875 / $0.41563YesYesYes
View
Volcengine Ark Coding Plan
glm-5.3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
Vultr
glm-5.3-flash
1,048,576131,072$0.10 / $0.35YesYesYes
View
Z.AI
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
Z.AI
glm-5.3-flashx
1,000,000131,072$0.37 / $1.25YesYesYes
View
Z.AI Coding Plan
glm-5.3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View
ZenMux
z-ai/glm-5.3-flash
1,000,000128,000$0.15 / $0.50YesYesYes
View
ZenMux
z-ai/glm-5.3-flashx
1,000,000128,000$0.375 / $1.25YesYesYes
View
Zhipu AI
glm-5.3-flash
1,000,000131,072$0.15 / $0.50YesYesYes
View
Zhipu AI
glm-5.3-flashx
1,000,000131,072$0.37 / $1.25YesYesYes
View
Zhipu AI Coding Plan
glm-5.3-flash
1,000,000131,072$0.00 / $0.00YesYesYes
View

Prices and limits apply to each offer. A dash means the value is not provided. OpenRouter offers show the hosting provider and routing channel separately.

Pricing

Compare token rates and cache charges across available routes and hosting configurations.

USD per million tokens
Input from$0.00Per 1M input tokens
Output from$0.00Per 1M output tokens
Available offers127Across 71 providers

Lowest listed input and output prices across the offers above; they may belong to different providers. These are base rates, before conditional pricing, discounts or additional fees.

Endpoint price comparisonLatest observed snapshot · 10 lowest-priced configurations
InputOutput
Relacerelace
$0.0352$0.50
InferenceNetinference-net/fp4
$0.044$0.60
Sail Researchsail-research/us
$0.045$0.60
OpenInferenceopen-inference/fp4
$0.04576$0.65
DeepInfradeepinfra/fp4
$0.075$0.25
StreamLakestreamlake/fp8
$0.08385$0.2795
Novitanovita/fp8
$0.084$0.28
GMICloudgmicloud/fp8
$0.09$0.30
DekaLLMdekallm
$0.10$1.00
Waferwafer
$0.10$0.50
z-ai/glm-5.3-flashObserved
Input
$0.15
Output
$0.50
Cache read
$0.03
Hosting prices through OpenRouter All endpoints and cache rates
EndpointInput / 1MOutput / 1MCache read / 1MCache write / 1M
AtlasCloudatlas-cloud/fp8$0.15$0.50$0.03—
BaseTenbaseten/fp8$0.15$0.50$0.03—
Cloudflarecloudflare$0.30$1.00$0.03—
CoreWeavecoreweave/nvfp4$0.15$0.50$0.05—
Crusoecrusoe/fp4$0.15$0.50$0.03—
Decartdecart/fp4$0.1125$0.375$0.0225—
DeepInfradeepinfra/fp4$0.075$0.25$0.015—
DekaLLMdekallm$0.10$1.00$0.04—
DigitalOceandigitalocean$0.15$0.50$0.03—
Fireworksfireworks$0.15$0.50$0.03—
Fireworksfireworks/us$0.225$0.75$0.045—
Friendlifriendli$0.15$0.50$0.03—
GMICloudgmicloud/fp8$0.09$0.30$0.018—
Inceptroninceptron/fp8$0.225$0.45$0.08—
InferenceNetinference-net/fp4$0.044$0.60$0.043—
Io Netio-net/fp8$0.1125$0.375$0.0225—
Modalmodal/nvfp4$0.15$0.50$0.03—
Morphmorph/fp8$0.16$0.444$0.04—
Near AInear-ai/fp8$0.105$0.35$0.0245—
NextBitnextbit/fp8$0.165$0.55$0.033—
Novitanovita/fp8$0.084$0.28$0.0168—
OpenInferenceopen-inference/fp4$0.04576$0.65$0.04576—
Parasailparasail/fp4$0.15$0.50$0.03—
Phalaphala/fp8$0.12$0.40$0.024—
PrimeIntellectprimeintellect$0.15$0.50——
Rekareka$0.15$0.50$0.03—
Relacerelace$0.0352$0.50$0.0352—
Sail Researchsail-research/us$0.045$0.60$0.0285—
SiliconFlowsiliconflow/fp8$0.15$0.50$0.03—
StreamLakestreamlake/fp8$0.08385$0.2795$0.01677—
Togethertogether$0.15$0.50$0.03—
Venicevenice$0.15$0.50$0.03—
Waferwafer$0.10$0.50$0.08—
Z.AIz-ai/fp8$0.15$0.50$0.03—

Each row is a hosting configuration. Open an offer in Providers for its full pricing conditions and additional fees. Endpoint observations may differ from the routing catalog's base price.

Performance

Review reported latency and output-throughput percentiles from the latest observed rolling window.

Latest 30-minute observation

Time to first token · milliseconds

Latency distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AtlasCloudatlas-cloud/fp8
2,454–22,423.52 ms
Cloudflarecloudflare
1,593–16,082.2 ms
CoreWeavecoreweave/nvfp4
984–6,293.41 ms
Crusoecrusoe/fp4
779–19,156.85 ms
Decartdecart/fp4
1,335–9,030.15 ms
DeepInfradeepinfra/fp4
1,878–15,011.43 ms
All latency percentiles 33 endpoints
EndpointP50P75P90P99
AtlasCloudatlas-cloud/fp82,454 ms3,595 ms7,842.2 ms22,423.52 ms
Cloudflarecloudflare1,593 ms2,642 ms4,963.8 ms16,082.2 ms
CoreWeavecoreweave/nvfp4984 ms1,323.25 ms1,959.9 ms6,293.41 ms
Crusoecrusoe/fp4779 ms1,648.75 ms3,850 ms19,156.85 ms
Decartdecart/fp41,335 ms1,888.25 ms2,582.9 ms9,030.15 ms
DeepInfradeepinfra/fp41,878 ms3,222.25 ms5,529.6 ms15,011.43 ms
DekaLLMdekallm1,119.5 ms2,006.5 ms3,852 ms12,952.75 ms
DigitalOceandigitalocean692 ms1,323.75 ms2,607.6 ms10,935.81 ms
Fireworksfireworks1,846 ms3,569 ms6,245 ms22,369.05 ms
Fireworksfireworks/us984 ms3,256.5 ms6,627 ms22,276 ms
Friendlifriendli1,372.5 ms2,309 ms3,941.9 ms9,976.27 ms
GMICloudgmicloud/fp82,933 ms5,479.75 ms9,185.7 ms20,115.77 ms
Inceptroninceptron/fp82,438 ms5,558 ms8,447.4 ms17,618.8 ms
InferenceNetinference-net/fp41,025 ms2,340.25 ms4,109.7 ms14,777.57 ms
Io Netio-net/fp81,343 ms2,101 ms3,114 ms15,353.95 ms
Modalmodal/nvfp4377 ms535 ms991 ms4,269.81 ms
Morphmorph/fp8580 ms838 ms1,809 ms17,582.87 ms
Near AInear-ai/fp81,574.5 ms3,114.25 ms6,909.8 ms21,933.33 ms
NextBitnextbit/fp82,792 ms3,459 ms5,013.6 ms8,507.7 ms
Novitanovita/fp81,723 ms2,601 ms4,095.6 ms13,938.78 ms
OpenInferenceopen-inference/fp43,487 ms4,757 ms14,657.8 ms21,353.08 ms
Parasailparasail/fp41,390.5 ms3,501 ms6,543 ms15,967.6 ms
Phalaphala/fp84,134 ms5,500.5 ms7,903 ms18,726.96 ms
PrimeIntellectprimeintellect1,042 ms2,531 ms4,804 ms11,504.28 ms
Rekareka1,005 ms1,517 ms3,251.9 ms13,328.81 ms
Relacerelace1,107 ms1,966.25 ms3,346.9 ms11,668.45 ms
Sail Researchsail-research/us1,404.5 ms2,029.75 ms3,046 ms23,993.89 ms
SiliconFlowsiliconflow/fp81,323 ms2,189 ms4,233.9 ms13,153.07 ms
StreamLakestreamlake/fp81,557 ms3,483.5 ms7,975.2 ms22,682.62 ms
Togethertogether663 ms1,286.25 ms2,759.8 ms21,399.49 ms
Venicevenice1,726 ms2,937 ms5,642 ms13,965.8 ms
Waferwafer584 ms959.5 ms2,412 ms38,334.9 ms
Z.AIz-ai/fp83,108 ms3,830.25 ms6,055.3 ms21,026.84 ms

Output throughput · tokens per second

Throughput distributionRolling 30-minute snapshot · 6 endpoints
P50P75P90P99
AtlasCloudatlas-cloud/fp8
69–143 t/s
Cloudflarecloudflare
45–82 t/s
CoreWeavecoreweave/nvfp4
95–264 t/s
Crusoecrusoe/fp4
98–240.29 t/s
Decartdecart/fp4
51–105 t/s
DeepInfradeepinfra/fp4
32–67 t/s
All throughput percentiles 33 endpoints
EndpointP50P75P90P99
AtlasCloudatlas-cloud/fp869 t/s92 t/s127 t/s143 t/s
Cloudflarecloudflare45 t/s54 t/s63 t/s82 t/s
CoreWeavecoreweave/nvfp495 t/s129 t/s173 t/s264 t/s
Crusoecrusoe/fp498 t/s126 t/s162 t/s240.29 t/s
Decartdecart/fp451 t/s65 t/s81 t/s105 t/s
DeepInfradeepinfra/fp432 t/s42 t/s51 t/s67 t/s
DekaLLMdekallm56 t/s97 t/s149 t/s239.04 t/s
DigitalOceandigitalocean36 t/s40 t/s41 t/s47 t/s
Fireworksfireworks66 t/s93 t/s115 t/s157 t/s
Fireworksfireworks/us81 t/s114 t/s145 t/s205.68 t/s
Friendlifriendli109 t/s147 t/s187 t/s253 t/s
GMICloudgmicloud/fp838 t/s48 t/s82 t/s188 t/s
Inceptroninceptron/fp827 t/s95.75 t/s151.1 t/s308.91 t/s
InferenceNetinference-net/fp452 t/s78 t/s104 t/s162 t/s
Io Netio-net/fp841 t/s46 t/s49 t/s55.28 t/s
Modalmodal/nvfp493 t/s120 t/s157 t/s248.09 t/s
Morphmorph/fp859 t/s79 t/s112 t/s160.36 t/s
Near AInear-ai/fp829 t/s35 t/s42 t/s59 t/s
NextBitnextbit/fp830 t/s34 t/s41 t/s92.52 t/s
Novitanovita/fp837 t/s53 t/s77 t/s139 t/s
OpenInferenceopen-inference/fp413 t/s19 t/s20 t/s21 t/s
Parasailparasail/fp466 t/s146 t/s236 t/s382.18 t/s
Phalaphala/fp848 t/s67 t/s118 t/s187 t/s
PrimeIntellectprimeintellect74 t/s105 t/s138 t/s294.45 t/s
Rekareka77 t/s110 t/s144 t/s224 t/s
Relacerelace50 t/s67 t/s86 t/s140.09 t/s
Sail Researchsail-research/us32 t/s42 t/s51 t/s62 t/s
SiliconFlowsiliconflow/fp869 t/s76 t/s104 t/s182 t/s
StreamLakestreamlake/fp837 t/s45 t/s53 t/s69 t/s
Togethertogether84 t/s109 t/s134 t/s192 t/s
Venicevenice40 t/s52 t/s95.9 t/s158.38 t/s
Waferwafer56 t/s65 t/s70 t/s103 t/s
Z.AIz-ai/fp862 t/s74 t/s89 t/s181 t/s

Percentiles describe the observed request distribution. These are snapshots of a rolling window, not a historical time series.

Uptime

Compare successful request rates across the latest 5-minute, 30-minute, and 24-hour observations.

Successful requests by window
Availability by endpointCurrent overlapping windows · 34 endpoints
Endpoint5 minutes30 minutes24 hours
AtlasCloudatlas-cloud/fp8
BaseTenbaseten/fp8
Cloudflarecloudflare
CoreWeavecoreweave/nvfp4
Crusoecrusoe/fp4
Decartdecart/fp4
DeepInfradeepinfra/fp4
DekaLLMdekallm
DigitalOceandigitalocean
Fireworksfireworks
Fireworksfireworks/us
Friendlifriendli
GMICloudgmicloud/fp8
Inceptroninceptron/fp8
InferenceNetinference-net/fp4
Io Netio-net/fp8
Modalmodal/nvfp4
Morphmorph/fp8
Near AInear-ai/fp8
NextBitnextbit/fp8
Novitanovita/fp8
OpenInferenceopen-inference/fp4
Parasailparasail/fp4
Phalaphala/fp8
PrimeIntellectprimeintellect
Rekareka
Relacerelace
Sail Researchsail-research/us
SiliconFlowsiliconflow/fp8
StreamLakestreamlake/fp8
Togethertogether
Venicevenice
Waferwafer
Z.AIz-ai/fp8

Reported by OpenRouter at the time of collection; rate-limited requests are excluded. A dash means no measurement was reported. This is not a live availability check.

Benchmarks

Review attributed evaluation results on each source's original scale.

External evaluations
z-ai/glm-5.3-flashObserved
Benchmark profileArtificial Analysis · shared source scale from 0 to 80
Coding71.5
Agentic50.9
Intelligence41.8
Design ArenaCategoryEloWin rateRank
models3d1,33456.8%11
modelsasciiart1,27654.8%12
modelscodecategories1,28949.5%17
modelsdataviz1,27250.5%24
modelsgamedev1,29447.2%17
modelssvg1,29057.1%8
modelsuicomponent1,32555%11
modelswebsite1,28048.5%21

Sources: Artificial Analysis and Design Arena, via OpenRouter. Scores retain their original scales and are not OpenAI Suite ratings.

Detailed evaluations

Evaluation scoresAccuracy or item F1 · higher is better
gpqa diamondGLM 5.3 Flash
90.909%
tau bench verified airlineGLM 5.3 Flash
70%
tau bench verified airlineZ.ai: GLM 5.3 Flash
75.6%
gpqa diamondZ.ai: GLM 5.3 Flash
85.69%
EvaluationConfigurationSourceResultsDetails
gpqa diamondGLM 5.3 Flashz-ai/glm-5.3-flashOpenRouter90.909%
Details
tau bench verified airlineGLM 5.3 Flashz-ai/glm-5.3-flashOpenRouter70%
Details
models · datavizGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,272 Elo
Details
models · uicomponentGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,325 Elo
Details
tau bench verified airlineZ.ai: GLM 5.3 Flashz-ai/glm-5.3-flash-20260826OpenRouter75.6%
Details
models · codecategoriesGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,289 Elo
Details
gpqa diamondZ.ai: GLM 5.3 Flashz-ai/glm-5.3-flash-20260826OpenRouter85.69%
Details
Composite indicesGLM 5.3 Flashz-ai/glm-5.3-flash-20260826Artificial AnalysisIntelligence 41.8 · Coding 71.5 · Agentic 50.9
Details
models · asciiartGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,276 Elo
Details
models · gamedevGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,294 Elo
Details
models · svgGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,290 Elo
Details
models · 3dGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,334 Elo
Details
models · websiteGLM-5.3-Flashz-ai/glm-5.3-flash-20260826Design Arena1,280 Elo
Details

Sources: Artificial Analysis, Design Arena and OpenRouter. Evaluation configurations and scales differ; results are not interchangeable. Updated Oct 4, 2026.

Apps & session costs

Compare observed median session costs for applications using this model, grouped by conversation length.

30-day sample · through Sep 27, 2026
ApplicationModel routeTurnsMedian cost / sessionDetails
Claude Codez-ai/glm-5.3-flash-2026082650+ turns
$0.439775
Details
Codexz-ai/glm-5.3-flash-202608262–9 turns
$0.004655
Details
Codexz-ai/glm-5.3-flash-202608261 turn
$0.000667
Details
Claude Codez-ai/glm-5.3-flash-202608261 turn
$0.002529
Details
Codexz-ai/glm-5.3-flash-2026082650+ turns
$0.052235
Details
Hermes Agentz-ai/glm-5.3-flash-2026082610–49 turns
$0.036837
Details
Kilo Codez-ai/glm-5.3-flash-2026082650+ turns
$0.356633
Details
Kilo Codez-ai/glm-5.3-flash-202608261 turn
$0.001367
Details
Kilo Codez-ai/glm-5.3-flash-202608262–9 turns
$0.008718
Details
Codexz-ai/glm-5.3-flash-2026082610–49 turns
$0.024903
Details
Kilo Codez-ai/glm-5.3-flash-2026082610–49 turns
$0.042624
Details
Claude Codez-ai/glm-5.3-flash-2026082610–49 turns
$0.054309
Details
Hermes Agentz-ai/glm-5.3-flash-202608262–9 turns
$0.005416
Details
Claude Codez-ai/glm-5.3-flash-202608262–9 turns
$0.00779
Details
Hermes Agentz-ai/glm-5.3-flash-2026082650+ turns
$0.258017
Details
Hermes Agentz-ai/glm-5.3-flash-202608261 turn
$0.001104
Details

Source: OpenRouter, as of Oct 2, 2026. Published under CC BY 4.0. These are observed costs per session, not token prices or a forecast for your workload. Apps are compared separately.

Activity

Daily prompt and completion tokens reported for this model's routes in OpenRouter's top 50. Missing days have no published value.

Last 90 complete UTC days
Daily token usage69,377,382,242,750 reported tokens · 39 days with observations
Jul 6, 2026Oct 3, 2026
Daily observations · 39 days
Date (UTC)Reported tokensReported routes
Oct 3, 20261,174,693,172,7981
Oct 2, 20261,650,868,242,7381
Oct 1, 20261,479,789,246,1901
Sep 30, 20261,421,695,513,6971
Sep 29, 20261,503,521,835,0691
Sep 28, 20261,317,250,703,6561
Sep 27, 20261,018,265,702,2941
Sep 26, 20261,282,011,293,2471
Sep 25, 20262,126,728,706,4211
Sep 24, 20261,964,153,145,1551
Sep 23, 20262,409,694,696,5231
Sep 22, 20263,153,115,211,4861
Sep 21, 20264,362,717,938,3671
Sep 20, 20262,515,123,497,6991
Sep 19, 20262,685,459,930,9031
Sep 18, 20261,944,494,628,1521
Sep 17, 20261,964,603,439,9791
Sep 16, 20261,798,851,266,8351
Sep 15, 20261,585,636,122,9861
Sep 14, 20261,577,050,969,6221
Sep 13, 20261,408,200,876,9291
Sep 12, 20261,343,229,501,2051
Sep 11, 20261,735,864,603,4291
Sep 10, 20261,821,255,463,1741
Sep 9, 20261,951,528,306,5041
Sep 8, 20261,808,250,248,9711
Sep 7, 20261,821,445,258,4381
Sep 6, 20261,472,192,943,0901
Sep 5, 20261,466,818,656,1131
Sep 4, 20261,922,687,542,0421
Sep 3, 20261,888,512,106,6631
Sep 2, 20261,791,420,461,4391
Sep 1, 20261,872,948,241,6681
Aug 31, 20261,979,092,565,4771
Aug 30, 20261,536,762,258,6431
Aug 29, 20261,322,948,887,4851
Aug 28, 20261,556,774,182,6601
Aug 27, 20261,380,964,831,8771
Aug 26, 2026360,760,043,1261

Source: OpenRouter (openrouter.ai/rankings), as of Oct 4, 2026. CC BY 4.0. Totals include only individually published routes. Tokenizers differ by provider; the aggregated “other” category is never assigned to a model.

Technical details & API

Explore specifications, supported parameters, reasoning controls, and API identifiers for each OpenRouter route.

Current route specifications
z-ai/glm-5.3-flashObserved
Context
1,048,576
Maximum output
943,717
Tokenizer
Other
Input
text, image, video
Output
text
Moderation
No
Added to OpenRouter
2026-08-26
Base URLhttps://openrouter.ai/api/v1
Model IDz-ai/glm-5.3-flash
Version IDz-ai/glm-5.3-flash-20260826
Weights IDzai-org/GLM-5.3-Flash

Reasoning

Required
Yes
Enabled by default
Yes
Default effort
max
Supported efforts
max, high, low

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • parallel_tool_calls
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Default parameters

top p
0.95
temperature
1

Frequently asked questions

Answers based on the model metadata and serving offers currently in the catalog.

How much does GLM-5.3-Flash cost?

Listed rates start at $0.00 per million input tokens and $0.00 per million output tokens. Prices vary by provider and configuration; see Pricing for details.

Which providers offer GLM-5.3-Flash?

There are 127 cataloged offers across 71 providers. Compare model IDs, prices and limits in Providers.

What is the context window?

The cataloged model context is 1,000,000 tokens. Each serving endpoint may apply a different limit.

Does it support tools and structured output?

Tool calling: Yes. Structured output: Yes. Support can vary by endpoint.

Catalog history · 3 recorded revisions

Metadata observations since this model was first cataloged. These are not model release versions.

  • · c51cc41cc7b0
  • · 051fd4dcc9dc
  • · 5834840b0b36