GLM 5.3 FlashX pricing
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture... Live index: 1 priced offer. Best input $0.370 per million tokens from Openrouter. Best output $1.25 per million tokens from Openrouter.
Pricing across providers
All figures are list prices per million tokens unless a column says otherwise. 1 offer is listed for GLM 5.3 FlashX. Best input in this view: Openrouter.
| Provider | Input / 1M | Output / 1M | Cached input | Batch |
|---|---|---|---|---|
O Openrouter | $0.370 | $1.25 | — | — |
Input vs output · 1M tokens
Cost calculator
The calculator uses the same dollars per million tokens as the table. Adjust sliders to see how GLM 5.3 FlashX cost scales with traffic.
0.037000¢ / req
0.062500¢ / req
Model specifications
Quick spec sheet for GLM 5.3 FlashX before you dive back into pricing. Reported under Z Ai.
- Context window
- 1,048,576 tokens
- Max output
- 131,072 tokens
- Vision (images)
- Yes
- Tool / function calling
- Yes
- Streaming
- Yes
- Released
- Sep 2026
- Primary provider
- Z Ai
- Model family
- N/A
Compare GLM 5.3 FlashX
Open a pair page to see GLM 5.3 FlashX next to another model with a shared provider matrix. 6 shortcuts below.
Locked
Compare with
Pick a model on both sides.
Popular GLM 5.3 FlashX comparisons
- GLM 5.3 FlashX vs GPT-4o
Compare pricing side by side
- GLM 5.3 FlashX vs Claude Sonnet 4.6
Compare pricing side by side
- GLM 5.3 FlashX vs Gemini 2.0 Flash
Compare pricing side by side
- GLM 5.3 FlashX vs Llama 3.1 70B
Compare pricing side by side
- GLM 5.3 FlashX vs Mistral 7B
Compare pricing side by side
- GLM 5.3 FlashX vs Amazon Nova Lite V1 0
Compare pricing side by side
Frequently asked questions
Read these after the table if you want plain language around GLM 5.3 FlashX rates. The short model note from our index: GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Related models
Models close to GLM 5.3 FlashX by family, provider, and context window — each with live pricing in our catalog.