LLM Opex Every way to run an LLM, every cloud

Models

gpt-oss-120b-Eagle3-v3: GPU memory and hosting cost

Open weights · 0.79 billion parameters · 131,072-token context · licence other

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 2 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AzureStandard_NV6ads_A10_v51 × A10 4 GB$0.454$331
AWSg4dn.xlarge1 × T4 16 GB$0.526$384
Google Cloudg2-standard-41 × L4 24 GB$0.7068$516
Alibaba Cloudecs.gn6i-c4g1.xlarge1 × T4 16 GB$1.12$816

More from NVIDIA

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.