LLM Opex Every way to run an LLM, every cloud

Models

gpt-neox-20b: GPU memory and hosting cost

Open weights · 20.74 billion parameters · 2,048-token context · licence apache-2.0

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 51.8 GB of GPU memory at fp16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AzureStandard_NC24ads_A100_v41 × A100 80 GB$3.67$2,681
Google Cloudg4-standard-481 × RTX PRO 6000 96 GB$4.50$3,285
Alibaba Cloudecs.gn9gc.4xlarge1 × L20N 72 GB$6.71$4,898
AWSp5.4xlarge1 × H100 80 GB$6.88$5,022

More from EleutherAI

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.