LLM Opex Every way to run an LLM, every cloud

Models

granite-3.2-8b-instruct: GPU memory and hosting cost

Open weights · 8.17 billion parameters · 131,072-token context · licence apache-2.0

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 20.4 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
Google Cloudg2-standard-41 × L4 24 GB$0.7068$516
AWSg6.xlarge1 × L4 24 GB$0.8048$588
AzureStandard_NC36ds_xl_RTXPRO6000BSE_v61 × RTX PRO 6000 24 GB$1.44$1,053
Alibaba Cloudecs.gn7i-c8g1.2xlarge1 × A10 24 GB$1.68$1,224

More from IBM

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.