GLM-4.7-Flash (unsloth): GPU memory and hosting cost
Open weights · 31.22 billion parameters · Mixture of experts · 202,752-token context · licence mit
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 78 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Azure | Standard_NC24ads_A100_v4 | 1 × A100 80 GB | $3.67 | $2,681 |
| Google Cloud | g4-standard-48 | 1 × RTX PRO 6000 96 GB | $4.50 | $3,285 |
| AWS | p5.4xlarge | 1 × H100 80 GB | $6.88 | $5,022 |
| Alibaba Cloud | ecs.gn8v.4xlarge | 1 × GPU H 96 GB | $7.31 | $5,337 |
More from Unsloth
- Llama-3.2-1B-Instruct (unsloth)
- Llama-3.2-3B-Instruct (unsloth)
- Meta-Llama-3.1-8B-Instruct (unsloth)
- Qwen2.5-7B-Instruct (unsloth)
- Qwen2.5-Coder-7B-Instruct (unsloth)
- Qwen3-Coder-Next-FP8 (unsloth)
Similar size, open weights
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.