Qwen3 32B: price on every cloud
Open weights · 32 billion parameters · 40,960-token context · licence apache-2.0
Per token, by cloud
| Cloud | Input, per 1M tokens | Output, per 1M tokens | Cheapest region |
|---|---|---|---|
| AWS | $0.075 | $0.30 | eu-north-1 · 13 regions |
| Alibaba Cloud | $0.16 | $0.64 | ap-southeast-1 · 3 regions |
| Azure | $0.30 | $1.20 | australiaeast · 38 regions |
| OCI | — | — | not sold per token |
| Core42 | — | — | not sold per token |
| Google Cloud | — | — | not sold per token |
Run it yourself
It needs 81.9 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Google Cloud | g4-standard-48 | 1 × RTX PRO 6000 96 GB | $4.50 | $3,285 |
| Azure | Standard_NC144ds_xl_RTXPRO6000BSE_v6 | 1 × RTX PRO 6000 96 GB | $6.38 | $4,657 |
| Alibaba Cloud | ecs.gn8v.4xlarge | 1 × GPU H 96 GB | $7.31 | $5,337 |
| AWS | p4d.24xlarge | 8 × A100 40 GB | $21.96 | $16,029 |
At the cheapest per-token price, running it on Google Cloud g4-standard-48 pays off above roughly 25.0 billion tokens a month.
More from Qwen
Similar size, open weights
- OpenAI gpt-oss 20B
- Mistral Small
- gemma-4-26B-A4B-it
- gemma-4-31b-it
- Qwen3.6-35B-A3B-NVFP4
- dolphin-2.9.1-yi-1.5-34b
Inside the UAE: sold in-country by Azure.
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.