LLM Opex Every way to run an LLM, every cloud

Models

Llama 3.1 405B: price on every cloud

Open weights · 405 billion parameters · 131,072-token context · licence llama3.1

Per token, by cloud

US dollars per million tokens, each cloud's cheapest region.
CloudInput, per 1M tokensOutput, per 1M tokensCheapest region
AWS$2.40$2.40us-east-2 · 2 regions
OCI$10.68$10.68global
Azure——not sold per token
Core42——not sold per token
Alibaba Cloud——not sold per token
Google Cloud——not sold per token

Run it yourself

It needs 1014.6 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AWSp5en.48xlarge8 × H200 141 GB$63.30$46,206
AzureStandard_ND96isr_MI300X_v58 × MI300X 192 GB$68.64$50,107
Google Clouda3-ultragpu-8g8 × H200 141 GB$84.81$61,909

At the cheapest per-token price, running it on AWS p5en.48xlarge pays off above roughly 19.3 billion tokens a month.

More from Meta

Similar size, open weights

Inside the UAE: sold in-country by OCI.

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.