Llama 3.1 8B: price on every cloud
Open weights · 8 billion parameters · 131,072-token context · licence llama3.1
Per token, by cloud
| Cloud | Input, per 1M tokens | Output, per 1M tokens | Cheapest region |
|---|---|---|---|
| AWS | $0.22 | $0.22 | us-east-1 · 3 regions |
| Azure | — | — | not sold per token |
| OCI | — | — | not sold per token |
| Core42 | — | — | not sold per token |
| Alibaba Cloud | — | — | not sold per token |
| Google Cloud | — | — | not sold per token |
Run it yourself
It needs 20.1 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Google Cloud | g2-standard-4 | 1 × L4 24 GB | $0.7068 | $516 |
| AWS | g6.xlarge | 1 × L4 24 GB | $0.8048 | $588 |
| Azure | Standard_NC36ds_xl_RTXPRO6000BSE_v6 | 1 × RTX PRO 6000 24 GB | $1.44 | $1,053 |
| Alibaba Cloud | ecs.gn7i-c8g1.2xlarge | 1 × A10 24 GB | $1.68 | $1,224 |
At the cheapest per-token price, running it on Google Cloud g2-standard-4 pays off above roughly 2.3 billion tokens a month.
More from Meta
- Llama 3.3 70B
- Llama 4 Maverick 17B
- Llama 4 Scout 17B
- Llama 3.1 405B
- Llama 3.1 70B
- Llama-3.2-1B-Instruct (meta-llama)
Similar size, open weights
Inside the UAE: no cloud sells it in-country.
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.