NVIDIA-Nemotron-3-Super-120B-A12B-FP8: GPU memory and hosting cost
Open weights · 123.61 billion parameters · Mixture of experts · 262,144-token context · licence other
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 154.5 GB of GPU memory at fp8, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| Google Cloud | a2-ultragpu-2g | 2 × A100 80 GB | $10.14 | $7,400 |
| Alibaba Cloud | ecs.gn8v-2x.8xlarge | 2 × GPU H 96 GB | $14.38 | $10,500 |
| AWS | p4d.24xlarge | 8 × A100 40 GB | $21.96 | $16,029 |
| Azure | Standard_ND96asr_v4 | 8 × A100 40 GB | $27.20 | $19,854 |
More from NVIDIA
- Qwen3.6-35B-A3B-NVFP4
- NVIDIA-Nemotron-3-Nano-4B-BF16
- Gemma-4-31B-IT-NVFP4
- Qwen2.5-VL-7B-Instruct-NVFP4
- NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- Qwen3.5-122B-A10B-NVFP4
Similar size, open weights
- OpenAI gpt-oss 120B
- Llama 3.3 70B
- Llama 4 Scout 17B
- Qwen3 235B A22B
- Llama 3.1 70B
- Qwen3-Coder-Next-FP8 (Qwen)
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.