Meta · live price
Llama 4 Scout API pricing
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
- Input
- $0.1
- per 1M tokens
- Output
- $0.3
- per 1M tokens
- Cached input
- –
- per 1M tokens
- Context
- 1.3M
- 16K max output
List price from OpenRouter, which passes Meta's price through without markup. Refreshed hourly. Released April 5, 2025. Model id meta-llama/llama-4-scout.
What Llama 4 Scout costs in practice
1,000 chat messages
$0.25
1,000 input + 500 output tokens each
1,000 RAG questions
$0.95
8,000 input + 500 output tokens each
100 coding-agent steps
$0.36
30,000 input + 2,000 output tokens each, no caching
Llama 4 Scout price history
Llama 4 Scout price by host
- DeepInfra$0.1 in · $0.3 out · 328K ctx · 100.0% uptime 24h
- Novita$0.18 in · $0.59 out · 131K ctx · 99.9% uptime 24h
- Google$0.25 in · $0.7 out · 1.3M ctx
Where to buy Meta credits for less
Disclosure: this site is run by the team behind smaaart, one of the options compared. Tables are ranked by price only, from live quotes and published fees, including when smaaart isn't first. How we compare
Llama 4 Scout cost calculator
Cheaper alternatives to Llama 4 Scout
| Model | Input / 1M | Output / 1M |
|---|---|---|
| DeepSeek: DeepSeek V4.1 Flash | $0.03 | $0.5 |
| Qwen: Qwen3.7 Flash | $0.03 | $0.13 |
| DeepSeek: DeepSeek V4 Flash 0423 | $0.042 | $0.084 |
| Google: Gemma 4 26B A4B | $0.09 | $0.3 |
| Qwen: Qwen3.5-9B | $0.1 | $0.15 |
| Qwen: Qwen3.5-Flash | $0.065 | $0.26 |
Other Meta models
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Llama 4 Maverick | $0.188 | $0.652 |
| Llama Guard 4 12B | $0.18 | $0.18 |
| Llama 3.3 70B Instruct | $0.1 | $0.32 |
| Llama 3.2 1B Instruct | $0.027 | $0.201 |
| Llama 3.2 3B Instruct | $0.05 | $0.33 |
| Llama 3.1 70B Instruct | $0.4 | $0.4 |