Thinking Machines · live price
Inkling API pricing
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
- Input
- $1
- per 1M tokens
- Output
- $4.05
- per 1M tokens
- Cached input
- $0.17
- per 1M tokens
- Context
- 524K
- 472K max output
List price from OpenRouter, which passes Thinking Machines's price through without markup. Refreshed hourly. Released July 17, 2026. Model id thinkingmachines/inkling.
What Inkling costs in practice
1,000 chat messages
$3.03
1,000 input + 500 output tokens each
1,000 RAG questions
$10.03
8,000 input + 500 output tokens each
100 coding-agent steps
$3.81
30,000 input + 2,000 output tokens each, no caching
Inkling price history
Inkling price by host
- DeepInfra$0.95 in · $4.05 out · 524K ctx · 99.9% uptime 24h
- Together$1 in · $4.05 out · 524K ctx · 99.8% uptime 24h
Where to buy Thinking Machines credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
Inkling cost calculator
Cheaper alternatives to Inkling
| Model | Input / 1M | Output / 1M |
|---|---|---|
| OpenAI: GPT-6 Luna Pro | $0.1 | $0.5 |
| OpenAI: GPT-6 Luna | $0.1 | $0.5 |
| Qwen: Qwen3.8 Omni Flash | $0.15 | $0.47 |
| Z.ai: GLM 5.3 FlashX | $0.37 | $1.25 |
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Google: Gemini 3.8 Flash | $0.75 | $3.75 |
Other Thinking Machines models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Inkling Small | $0.45 | $1.20 |