LLM Opex Every way to run an LLM, every cloud

Models

Llama-3.2-3B-Instruct-FP8-dynamic: GPU memory and hosting cost

Open weights · 3.61 billion parameters · 131,072-token context · licence llama3.2

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 4.5 GB of GPU memory at fp8, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AWSg4dn.xlarge1 × T4 16 GB$0.526$384
AzureStandard_NC4as_T4_v31 × T4 16 GB$0.526$384
Google Cloudg2-standard-41 × L4 24 GB$0.7068$516
Alibaba Cloudecs.gn6i-c4g1.xlarge1 × T4 16 GB$1.12$816

More from Other

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.