Qwen2.5-VL-32B-Instruct: GPU memory and hosting cost
Open weights · 33.45 billion parameters · 128,000-token context · licence apache-2.0
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 83.6 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Google Cloud | g4-standard-48 | 1 × RTX PRO 6000 96 GB | $4.50 | $3,285 |
| Azure | Standard_NC144ds_xl_RTXPRO6000BSE_v6 | 1 × RTX PRO 6000 96 GB | $6.38 | $4,657 |
| Alibaba Cloud | ecs.gn8v.4xlarge | 1 × GPU H 96 GB | $7.31 | $5,337 |
| AWS | p4d.24xlarge | 8 × A100 40 GB | $21.96 | $16,029 |
More from Qwen
Similar size, open weights
- OpenAI gpt-oss 20B
- Mistral Small
- gemma-4-26B-A4B-it
- gemma-4-31b-it
- Qwen3.6-35B-A3B-NVFP4
- dolphin-2.9.1-yi-1.5-34b
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.