DeepSeek-V4-Flash: GPU memory and hosting cost
Open weights · 290.94 billion parameters · Mixture of experts · 1,048,576-token context · licence mit
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 727.4 GB of GPU memory at fp16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Alibaba Cloud | ecs.gn8v-8x.16xlarge | 8 × GPU H 96 GB | $58.49 | $42,698 |
| AWS | p5en.48xlarge | 8 × H200 141 GB | $63.30 | $46,206 |
| Azure | Standard_ND96isr_MI300X_v5 | 8 × MI300X 192 GB | $68.64 | $50,107 |
| Google Cloud | a3-ultragpu-8g | 8 × H200 141 GB | $84.81 | $61,909 |
More from DeepSeek
- DeepSeek V3.1
- DeepSeek-V4-Flash-0731
- DeepSeek-V3.2
- DeepSeek-V3-0324
- DeepSeek-R1
- DeepSeek-R1-Distill-Qwen-1.5B
Similar size, open weights
- Llama 4 Maverick 17B
- Llama 3.1 405B
- Qwen3 235B A22B
- GLM-5.3-Flash
- MiniMax-M2.7
- Qwen3-Coder-480B-A35B-Instruct-FP8
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.