Qwen3-VL-235B-A22B-Instruct-FP8-dynamic: GPU memory and hosting cost
Open weights · 235.76 billion parameters · Mixture of experts · 262,144-token context · licence apache-2.0
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 294.7 GB of GPU memory at fp8, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Google Cloud | a2-ultragpu-4g | 4 × A100 80 GB | $20.28 | $14,801 |
| AWS | p4d.24xlarge | 8 × A100 40 GB | $21.96 | $16,029 |
| Azure | Standard_ND96asr_v4 | 8 × A100 40 GB | $27.20 | $19,854 |
| Alibaba Cloud | ecs.gn8v-4x.8xlarge | 4 × GPU H 96 GB | $29.25 | $21,349 |
More from Other
- dolphin-2.9.1-yi-1.5-34b
- JiRackUltra_1b
- OTel-2.0-LLM-31B-IT
- moondream2
- TinyLlama-1.1B-Chat-v1.0
- MiniCPM5-2B
Similar size, open weights
- Llama 4 Maverick 17B
- Llama 3.1 405B
- Qwen3 235B A22B
- GLM-5.3-Flash
- DeepSeek-V4-Flash-0731
- DeepSeek-V4-Flash
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.