Z.ai · live price
GLM 4.5 Air API pricing
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
- Input
- $0.13
- per 1M tokens
- Output
- $0.85
- per 1M tokens
- Cached input
- $0.025
- per 1M tokens
- Context
- 131K
- 98K max output
List price from OpenRouter, which passes Z.ai's price through without markup. Refreshed hourly. Released July 25, 2025. Model id z-ai/glm-4.5-air.
What GLM 4.5 Air costs in practice
1,000 chat messages
$0.56
1,000 input + 500 output tokens each
1,000 RAG questions
$1.47
8,000 input + 500 output tokens each
100 coding-agent steps
$0.56
30,000 input + 2,000 output tokens each, no caching
GLM 4.5 Air price history
GLM 4.5 Air price by host
- Novita$0.13 in · $0.85 out · 131K ctx · 99.9% uptime 24h
- SiliconFlow$0.14 in · $0.86 out · 131K ctx · 99.1% uptime 24h
- Z.AI$0.2 in · $1.10 out · 131K ctx · 99.6% uptime 24h
Where to buy Z.ai credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
GLM 4.5 Air cost calculator
Cheaper alternatives to GLM 4.5 Air
| Model | Input / 1M | Output / 1M |
|---|---|---|
| OpenAI: GPT-6 Luna Pro | $0.1 | $0.5 |
| OpenAI: GPT-6 Luna | $0.1 | $0.5 |
| Qwen: Qwen3.8 Omni Flash | $0.15 | $0.47 |
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Qwen: Qwen3.8 Flash | $0.15 | $0.47 |
| Qwen: Qwen3.7 Flash | $0.03 | $0.13 |