LLM Opex Every way to run an LLM, every cloud

Models

gemma-2-9b: GPU memory and hosting cost

Open weights · 9.24 billion parameters · licence gemma

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 46.2 GB of GPU memory at fp32, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AWSg6e.xlarge1 × L40S 48 GB$1.86$1,359
Alibaba Cloudecs.gn8is.2xlarge1 × L20 48 GB$2.26$1,648
AzureStandard_NC72ds_xl_RTXPRO6000BSE_v61 × RTX PRO 6000 48 GB$2.83$2,066
Google Cloudg4-standard-481 × RTX PRO 6000 96 GB$4.50$3,285

More from Google

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.