getcheapaicredits

NVIDIA · live price

Nemotron 3 Ultra API pricing

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Mid-range ($1–5 per 1M)Tool callingReasoning
Input
$0.6
per 1M tokens
Output
$2.40
per 1M tokens
Cached input
$0.12
per 1M tokens
Context
262K
183K max output

List price from OpenRouter, which passes NVIDIA's price through without markup. Refreshed hourly. Released June 4, 2026. Model id nvidia/nemotron-3-ultra-550b-a55b.

What Nemotron 3 Ultra costs in practice

At list price, before any batch or caching discount.

1,000 chat messages

$1.80

1,000 input + 500 output tokens each

1,000 RAG questions

$6.00

8,000 input + 500 output tokens each

100 coding-agent steps

$2.28

30,000 input + 2,000 output tokens each, no caching

Nemotron 3 Ultra price history

No price change since we started tracking on 2026-10-01. We record prices daily; changes will show here.

Nemotron 3 Ultra price by host

DeepInfra serves it below list price right now. Route to it to pay less on every token.

Where to buy NVIDIA credits for less

Nemotron 3 Ultra costs the same per token everywhere it's sold through OpenRouter. What changes is the price of the credits. Live quotes for $100 of usage:
#Where you buyYou pay
1
OrbioCardCheapest right now
$85.10
incl. $2.60 card processing
2
smaaartCard
$95.00
+ card processing and tax at checkout (not in this number)
3
Direct from the labCard
$100.00
no top-up fee
4
OpenRouterCrypto
$105.00
fees included
5
OpenRouterCard
$105.50
fees included

Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare

Nemotron 3 Ultra cost calculator

Enter your monthly tokens to price your own usage.
Usage at list price$54.00/ month

Cheaper alternatives to Nemotron 3 Ultra

Other labs' models in the same price range or the one below, cheaper per token.
ModelInput / 1MOutput / 1M
OpenAI: GPT-6 Luna Pro$0.1$0.5
OpenAI: GPT-6 Luna$0.1$0.5
Qwen: Qwen3.8 Omni Flash$0.15$0.47
Z.ai: GLM 5.3 FlashX$0.37$1.25
DeepSeek: DeepSeek V4.1 Flash$0.03$0.5
Qwen: Qwen3.8 Flash$0.15$0.47

Other NVIDIA models

ModelInput / 1MOutput / 1M
Nemotron 3.5 Content Safety$0.2$0.2
Nemotron 3.5 Lightning$0.059$0.17
Nemotron 3 Super$0.08$0.45
Nemotron 3 Nano 30B A3B$0.05$0.2
All NVIDIA API prices

Nemotron 3 Ultra pricing FAQ