Checked 2026-10-01
How to pay less for AI APIs
- 01−50%
Batch what doesn't need an instant answer
OpenAI, Anthropic and Google all price batch requests at half the standard rate, input and output. Results come back asynchronously: within 24 hours at OpenAI, most Anthropic batches in under an hour.
- 02up to −90%
Cache the prompt prefix you repeat
Anthropic bills cache reads at 0.1× the input price, with writes at 1.25× (5-minute cache) or 2× (1-hour). OpenAI discounts cached input by up to 95% depending on the model, automatically. Gemini has a separate cached rate plus an hourly storage fee. Put system prompts, tools and documents first, and keep them identical between calls.
- 03batch rates
Use OpenAI Flex for slow-but-not-batch work
Flex processing bills tokens at Batch API rates with slower, variable latency, and stacks with caching discounts. It's in beta, on a limited set of models.
Sources:OpenAI
- 04−50%
Run DeepSeek off-peak
DeepSeek's off-peak rates are half the peak rates: 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, outside Chinese holidays.
Sources:DeepSeek
- 05avoid ~2×
Skip the fast lane unless you need it
OpenAI's Fast mode (formerly Priority processing) costs more per token than standard. Use it only where latency earns its price.
Sources:OpenAI
- 06often −80% or more
Route easy requests to a smaller model
Small models cost a fraction of frontier ones per token. Classify, extract and summarize with a small model; keep the large one for hard reasoning. Compare live prices in the calculator.
Open the calculator - 07live
Buy the credits for less
Everything above lowers the token bill. Separately, the price of the credits themselves varies by where you buy them: OpenRouter adds a top-up fee, discount books sell below face value within their depth.
See today's prices