LLM Opex Every way to run an LLM, every cloud

Models

GLM-5.1-FP8: GPU memory and hosting cost

Open weights · 753.91 billion parameters · Mixture of experts · 202,752-token context · licence mit

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 942.4 GB of GPU memory at fp8, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AWSp5en.48xlarge8 × H200 141 GB$63.30$46,206
AzureStandard_ND96isr_MI300X_v58 × MI300X 192 GB$68.64$50,107
Google Clouda3-ultragpu-8g8 × H200 141 GB$84.81$61,909

More from Z.ai

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.