LLM Opex Every way to run an LLM, every cloud

Models

Llama 3.3 70B: price on every cloud

Open weights · 70 billion parameters · 131,072-token context · licence llama3.3

Per token, by cloud

US dollars per million tokens, each cloud's cheapest region.
CloudInput, per 1M tokensOutput, per 1M tokensCheapest region
Azure$0.71$0.71australiaeast · 42 regions
AWS$0.72$0.72us-east-1 · 3 regions
Google Cloud$0.72$0.72global
OCI$0.72$0.72global · 7 regions
Core42——not sold per token
Alibaba Cloud——not sold per token

Run it yourself

It needs 176.4 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
Alibaba Cloudecs.gn8v-2x.8xlarge2 × GPU H 96 GB$14.38$10,500
Google Clouda2-ultragpu-4g4 × A100 80 GB$20.28$14,801
AWSp4d.24xlarge8 × A100 40 GB$21.96$16,029
AzureStandard_ND96asr_v48 × A100 40 GB$27.20$19,854

At the cheapest per-token price, running it on Alibaba Cloud ecs.gn8v-2x.8xlarge pays off above roughly 14.8 billion tokens a month.

More from Meta

Similar size, open weights

Inside the UAE: sold in-country by Azure.

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.