Llama-3.2-3B (unsloth): GPU memory and hosting cost
Open weights · 3.21 billion parameters · 131,072-token context · licence llama3.2
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 8 GB of GPU memory at bf16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| AWS | g4dn.xlarge | 1 × T4 16 GB | $0.526 | $384 |
| Azure | Standard_NC4as_T4_v3 | 1 × T4 16 GB | $0.526 | $384 |
| Google Cloud | g2-standard-4 | 1 × L4 24 GB | $0.7068 | $516 |
| Alibaba Cloud | ecs.gn6i-c4g1.xlarge | 1 × T4 16 GB | $1.12 | $816 |
More from Unsloth
- Llama-3.2-1B-Instruct (unsloth)
- Llama-3.2-3B-Instruct (unsloth)
- Meta-Llama-3.1-8B-Instruct (unsloth)
- Qwen2.5-7B-Instruct (unsloth)
- Qwen2.5-Coder-7B-Instruct (unsloth)
- Qwen3-Coder-Next-FP8 (unsloth)
Similar size, open weights
- Qwen3-4B
- Qwen2.5-3B-Instruct (Qwen)
- Qwen3-4B-Instruct-2507 (Qwen)
- Qwen3-1.7B (Qwen)
- NVIDIA-Nemotron-3-Nano-4B-BF16
- JiRackUltra_1b
List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.