Llama 3.1 405B: price on every cloud
Open weights · 405 billion parameters · 131,072-token context · licence llama3.1
Per token, by cloud
| Cloud | Input, per 1M tokens | Output, per 1M tokens | Cheapest region |
|---|---|---|---|
| AWS | $2.40 | $2.40 | us-east-2 · 2 regions |
| OCI | $10.68 | $10.68 | global |
| Azure | — | — | not sold per token |
| Core42 | — | — | not sold per token |
| Alibaba Cloud | — | — | not sold per token |
| Google Cloud | — | — | not sold per token |
Run it yourself
It needs 1014.6 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| AWS | p5en.48xlarge | 8 × H200 141 GB | $63.30 | $46,206 |
| Azure | Standard_ND96isr_MI300X_v5 | 8 × MI300X 192 GB | $68.64 | $50,107 |
| Google Cloud | a3-ultragpu-8g | 8 × H200 141 GB | $84.81 | $61,909 |
At the cheapest per-token price, running it on AWS p5en.48xlarge pays off above roughly 19.3 billion tokens a month.
More from Meta
- Llama 3.3 70B
- Llama 4 Maverick 17B
- Llama 4 Scout 17B
- Llama 3.1 70B
- Llama 3.1 8B
- Llama-3.2-1B-Instruct (meta-llama)
Similar size, open weights
Inside the UAE: sold in-country by OCI.
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.