Z.ai · live price
GLM 5.3 Flash API pricing
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Input
- $0.15
- per 1M tokens
- Output
- $0.5
- per 1M tokens
- Cached input
- $0.03
- per 1M tokens
- Context
- 1.0M
- 944K max output
List price from OpenRouter, which passes Z.ai's price through without markup. Refreshed hourly. Released August 26, 2026. Model id z-ai/glm-5.3-flash.
What GLM 5.3 Flash costs in practice
1,000 chat messages
$0.40
1,000 input + 500 output tokens each
1,000 RAG questions
$1.45
8,000 input + 500 output tokens each
100 coding-agent steps
$0.55
30,000 input + 2,000 output tokens each, no caching
GLM 5.3 Flash price history
GLM 5.3 Flash price by host
- OpenInference$0.02 in · $0.3 out · 1.0M ctx · 99.9% uptime 24h
- DeepInfra$0.075 in · $0.25 out · 1.0M ctx · 99.5% uptime 24h
- Novita$0.084 in · $0.28 out · 1.0M ctx · 99.4% uptime 24h
- StreamLake$0.087 in · $0.29 out · 1.0M ctx · 98.5% uptime 24h
- GMICloud$0.09 in · $0.3 out · 1.0M ctx · 99.4% uptime 24h
- Near AI$0.12 in · $0.4 out · 1.0M ctx · 98.3% uptime 24h
- Phala$0.12 in · $0.4 out · 1.0M ctx · 99.0% uptime 24h
- Relace$0.035 in · $0.5 out · 1.0M ctx · 99.7% uptime 24h
- Decart$0.128 in · $0.425 out · 1.0M ctx · 98.1% uptime 24h
- InferenceNet$0.042 in · $0.6 out · 1.0M ctx · 99.6% uptime 24h
- Sail Research$0.045 in · $0.6 out · 1.0M ctx · 99.6% uptime 24h
- Modal$0.15 in · $0.5 out · 1.0M ctx · 99.7% uptime 24h
- BaseTen$0.15 in · $0.5 out · 1.0M ctx · 98.6% uptime 24h
- Crusoe$0.15 in · $0.5 out · 1.0M ctx · 99.3% uptime 24h
- CoreWeave$0.15 in · $0.5 out · 1.0M ctx · 99.6% uptime 24h
- AtlasCloud$0.15 in · $0.5 out · 1.0M ctx · 91.9% uptime 24h
- Fireworks$0.15 in · $0.5 out · 1.0M ctx · 99.7% uptime 24h
- Friendli$0.15 in · $0.5 out · 1.0M ctx · 99.4% uptime 24h
- SiliconFlow$0.15 in · $0.5 out · 1.0M ctx · 99.1% uptime 24h
- DigitalOcean$0.15 in · $0.5 out · 1.0M ctx · 98.7% uptime 24h
- Together$0.15 in · $0.5 out · 1.0M ctx · 99.8% uptime 24h
- Reka$0.15 in · $0.5 out · 262K ctx · 99.7% uptime 24h
- Parasail$0.15 in · $0.5 out · 1.0M ctx · 99.7% uptime 24h
- Venice$0.15 in · $0.5 out · 1.0M ctx · 96.4% uptime 24h
- Io Net$0.15 in · $0.5 out · 262K ctx · 97.9% uptime 24h
- Z.AI$0.15 in · $0.5 out · 1.0M ctx · 99.8% uptime 24h
- Inceptron$0.225 in · $0.45 out · 1.0M ctx · 99.6% uptime 24h
- NextBit$0.165 in · $0.55 out · 1.0M ctx · 99.9% uptime 24h
- Morph$0.2 in · $0.7 out · 1.0M ctx · 96.1% uptime 24h
- Cloudflare$0.3 in · $1 out · 1.0M ctx · 99.1% uptime 24h
- Wafer$1 in · $0.75 out · 1.0M ctx · 99.9% uptime 24h
Where to buy Z.ai credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
GLM 5.3 Flash cost calculator
Cheaper alternatives to GLM 5.3 Flash
| Model | Input / 1M | Output / 1M |
|---|---|---|
| OpenAI: GPT-6 Luna Pro | $0.1 | $0.5 |
| OpenAI: GPT-6 Luna | $0.1 | $0.5 |
| Qwen: Qwen3.8 Omni Flash | $0.15 | $0.47 |
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Qwen: Qwen3.8 Flash | $0.15 | $0.47 |
| Qwen: Qwen3.7 Flash | $0.03 | $0.13 |
Other Z.ai models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| GLM 5.3 | $0.222 | $3.39 |
| GLM 5.3 FlashX | $0.37 | $1.25 |
| GLM 5.3 Prime | $2.80 | $8.80 |
| GLM 5.2 | $1.40 | $4.40 |
| GLM 5.1 | $0.965 | $3.03 |
| GLM 5V Turbo | $1.20 | $4 |