LLM Opex Every way to run an LLM, every cloud

Models

Qwen3 32B: price on every cloud

Open weights · 32 billion parameters · 40,960-token context · licence apache-2.0

Per token, by cloud

US dollars per million tokens, each cloud's cheapest region.
CloudInput, per 1M tokensOutput, per 1M tokensCheapest region
AWS$0.075$0.30eu-north-1 · 13 regions
Alibaba Cloud$0.16$0.64ap-southeast-1 · 3 regions
Azure$0.30$1.20australiaeast · 38 regions
OCI——not sold per token
Core42——not sold per token
Google Cloud——not sold per token

Run it yourself

It needs 81.9 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
Google Cloudg4-standard-481 × RTX PRO 6000 96 GB$4.50$3,285
AzureStandard_NC144ds_xl_RTXPRO6000BSE_v61 × RTX PRO 6000 96 GB$6.38$4,657
Alibaba Cloudecs.gn8v.4xlarge1 × GPU H 96 GB$7.31$5,337
AWSp4d.24xlarge8 × A100 40 GB$21.96$16,029

At the cheapest per-token price, running it on Google Cloud g4-standard-48 pays off above roughly 25.0 billion tokens a month.

More from Qwen

Similar size, open weights

Inside the UAE: sold in-country by Azure.

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.