Llama-3.1-405B-FP8: GPU memory and hosting cost
Open weights · 405.87 billion parameters · licence llama3.1
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 507.3 GB of GPU memory at fp8, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| AWS | p4de.24xlarge | 8 × A100 80 GB | $27.45 | $20,036 |
| Azure | Standard_ND96amsr_A100_v4 | 8 × A100 80 GB | $32.77 | $23,922 |
| Google Cloud | a2-ultragpu-8g | 8 × A100 80 GB | $40.55 | $29,602 |
| Alibaba Cloud | ecs.gn8v-8x.16xlarge | 8 × GPU H 96 GB | $58.49 | $42,698 |
More from Meta
Similar size, open weights
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.