Llama 3.1 70B: price on every cloud
Open weights · 70 billion parameters · 131,072-token context · licence llama3.1
Per token, by cloud
| Cloud | Input, per 1M tokens | Output, per 1M tokens | Cheapest region |
|---|---|---|---|
| AWS | $0.72 | $0.72 | us-east-1 · 3 regions |
| Azure | — | — | not sold per token |
| OCI | — | — | not sold per token |
| Core42 | — | — | not sold per token |
| Alibaba Cloud | — | — | not sold per token |
| Google Cloud | — | — | not sold per token |
Run it yourself
It needs 176.4 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Alibaba Cloud | ecs.gn8v-2x.8xlarge | 2 × GPU H 96 GB | $14.38 | $10,500 |
| Google Cloud | a2-ultragpu-4g | 4 × A100 80 GB | $20.28 | $14,801 |
| AWS | p4d.24xlarge | 8 × A100 40 GB | $21.96 | $16,029 |
| Azure | Standard_ND96asr_v4 | 8 × A100 40 GB | $27.20 | $19,854 |
At the cheapest per-token price, running it on Alibaba Cloud ecs.gn8v-2x.8xlarge pays off above roughly 14.6 billion tokens a month.
More from Meta
- Llama 3.3 70B
- Llama 4 Maverick 17B
- Llama 4 Scout 17B
- Llama 3.1 405B
- Llama 3.1 8B
- Llama-3.2-1B-Instruct (meta-llama)
Similar size, open weights
- OpenAI gpt-oss 120B
- Qwen3.6-35B-A3B-NVFP4
- Qwen3.6-35B-A3B
- NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- Qwen3.5-122B-A10B-NVFP4
- Qwen3-Coder-Next-FP8 (Qwen)
Inside the UAE: no cloud sells it in-country.
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.