Nous · live price
Hermes 4 405B API pricing
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...
- Input
- $1
- per 1M tokens
- Output
- $3
- per 1M tokens
- Cached input
- –
- per 1M tokens
- Context
- 131K
- 118K max output
List price from OpenRouter, which passes Nous's price through without markup. Refreshed hourly. Released August 26, 2025. Model id nousresearch/hermes-4-405b.
What Hermes 4 405B costs in practice
1,000 chat messages
$2.50
1,000 input + 500 output tokens each
1,000 RAG questions
$9.50
8,000 input + 500 output tokens each
100 coding-agent steps
$3.60
30,000 input + 2,000 output tokens each, no caching
Hermes 4 405B price history
Hermes 4 405B price by host
- Nebius$1 in · $3 out · 131K ctx · 100.0% uptime 24h
Where to buy Nous credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
Hermes 4 405B cost calculator
Cheaper alternatives to Hermes 4 405B
| Model | Input / 1M | Output / 1M |
|---|---|---|
| OpenAI: GPT-6 Luna Pro | $0.1 | $0.5 |
| OpenAI: GPT-6 Luna | $0.1 | $0.5 |
| Qwen: Qwen3.8 Omni Flash | $0.15 | $0.47 |
| Z.ai: GLM 5.3 FlashX | $0.37 | $1.25 |
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Qwen: Qwen3.8 Flash | $0.15 | $0.47 |
Other Nous models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Hermes 3 70B Instruct | $0.7 | $0.7 |
| Hermes 3 405B Instruct | $1 | $1 |