Z.ai · live price
GLM 4.7 Flash API pricing
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
- Input
- $0.061
- per 1M tokens
- Output
- $0.4
- per 1M tokens
- Cached input
- –
- per 1M tokens
- Context
- 200K
- 118K max output
List price from OpenRouter, which passes Z.ai's price through without markup. Refreshed hourly. Released January 19, 2026. Model id z-ai/glm-4.7-flash.
What GLM 4.7 Flash costs in practice
1,000 chat messages
$0.26
1,000 input + 500 output tokens each
1,000 RAG questions
$0.68
8,000 input + 500 output tokens each
100 coding-agent steps
$0.26
30,000 input + 2,000 output tokens each, no caching
GLM 4.7 Flash price history
GLM 4.7 Flash price by host
- Venice$0.06 in · $0.4 out · 128K ctx · 97.0% uptime 24h
- Cloudflare$0.061 in · $0.4 out · 131K ctx · 98.9% uptime 24h
- Novita$0.07 in · $0.4 out · 200K ctx · 65.8% uptime 24h
Where to buy Z.ai credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
GLM 4.7 Flash cost calculator
Cheaper alternatives to GLM 4.7 Flash
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Qwen: Qwen3.7 Flash | $0.03 | $0.13 |
| DeepSeek: DeepSeek V4 Flash 0423 | $0.042 | $0.084 |
| Google: Gemma 4 26B A4B | $0.09 | $0.3 |
| Qwen: Qwen3.5-9B | $0.1 | $0.15 |
| Qwen: Qwen3.5-Flash | $0.065 | $0.26 |
| Mistral: Ministral 3 3B 2512 | $0.1 | $0.1 |
Other Z.ai models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| GLM 5 | $0.6 | $1.92 |
| GLM 4.7 | $0.6 | $2.20 |
| GLM 4.6V | $0.3 | $0.9 |
| GLM 5 Turbo | $1.20 | $4 |
| GLM 5V Turbo | $1.20 | $4 |
| GLM 5.1 | $0.965 | $3.03 |