LLM Opex Every way to run an LLM, every cloud

Models

Mistral-Medium-3.5-128B: GPU memory and hosting cost

Open weights · 127.7 billion parameters · 262,144-token context · licence other

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 159.6 GB of GPU memory at fp8, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
Google Clouda2-ultragpu-2g2 × A100 80 GB$10.14$7,400
Alibaba Cloudecs.gn8v-2x.8xlarge2 × GPU H 96 GB$14.38$10,500
AWSp4d.24xlarge8 × A100 40 GB$21.96$16,029
AzureStandard_ND96asr_v48 × A100 40 GB$27.20$19,854

More from Mistral

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.