NVIDIA · live price
Nemotron 3 Ultra API pricing
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
- Input
- $0.6
- per 1M tokens
- Output
- $2.40
- per 1M tokens
- Cached input
- $0.12
- per 1M tokens
- Context
- 262K
- 183K max output
List price from OpenRouter, which passes NVIDIA's price through without markup. Refreshed hourly. Released June 4, 2026. Model id nvidia/nemotron-3-ultra-550b-a55b.
What Nemotron 3 Ultra costs in practice
1,000 chat messages
$1.80
1,000 input + 500 output tokens each
1,000 RAG questions
$6.00
8,000 input + 500 output tokens each
100 coding-agent steps
$2.28
30,000 input + 2,000 output tokens each, no caching
Nemotron 3 Ultra price history
Nemotron 3 Ultra price by host
- DeepInfra$0.5 in · $2.20 out · 262K ctx · 99.9% uptime 24h
- BaseTen$0.6 in · $2.40 out · 203K ctx · 99.8% uptime 24h
- Venice$0.625 in · $3.13 out · 256K ctx · 97.9% uptime 24h
Where to buy NVIDIA credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
Nemotron 3 Ultra cost calculator
Cheaper alternatives to Nemotron 3 Ultra
| Model | Input / 1M | Output / 1M |
|---|---|---|
| OpenAI: GPT-6 Luna Pro | $0.1 | $0.5 |
| OpenAI: GPT-6 Luna | $0.1 | $0.5 |
| Qwen: Qwen3.8 Omni Flash | $0.15 | $0.47 |
| Z.ai: GLM 5.3 FlashX | $0.37 | $1.25 |
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Qwen: Qwen3.8 Flash | $0.15 | $0.47 |
Other NVIDIA models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Nemotron 3.5 Content Safety | $0.2 | $0.2 |
| Nemotron 3.5 Lightning | $0.059 | $0.17 |
| Nemotron 3 Super | $0.08 | $0.45 |
| Nemotron 3 Nano 30B A3B | $0.05 | $0.2 |