LLM Opex Every way to run an LLM, every cloud

Which GPU instance or cluster can actually hold it?

A model fits when its weights, plus the memory its attention cache needs for your context length and the requests you serve at once, are within the GPU memory of an instance or a cluster. Anything too small is left out; what fits is listed per cloud with its price.

Memory needed and the cheapest hardware that fits, by model

At the precision the weights ship in, with the default allowance for the attention cache; list prices, on demand.
ModelParametersPrecisionGPU memory neededCheapest instance that fits
OpenAI gpt-oss 120B116.83 Bint473 GBAzure Standard_NC24ads_A100_v4: 1 × A100, $3.67 an hour
OpenAI gpt-oss 20B20.91 Bint413.1 GBAWS g4dn.xlarge: 1 × T4, $0.526 an hour
Llama 3.3 70B70.55 Bbf16176.4 GBAlibaba Cloud ecs.gn8v-2x.8xlarge: 2 × GPU H, $14.38 an hour
Llama 4 Maverick 17B401.58 Bbf161003.9 GBAWS p5en.48xlarge: 8 × H200, $63.30 an hour
Llama 4 Scout 17B108.64 Bbf16271.6 GBGoogle Cloud a2-ultragpu-4g: 4 × A100, $20.28 an hour
Qwen3 32B32.76 Bbf1681.9 GBGoogle Cloud g4-standard-48: 1 × RTX PRO 6000, $4.50 an hour
DeepSeek V3.1684.53 Bfp8855.7 GBAWS p5en.48xlarge: 8 × H200, $63.30 an hour
Llama 3.1 405B405.85 Bbf161014.6 GBAWS p5en.48xlarge: 8 × H200, $63.30 an hour
Mistral Small23.57 Bbf1658.9 GBAzure Standard_NC24ads_A100_v4: 1 × A100, $3.67 an hour
Qwen3 235B A22B235.09 Bbf16587.7 GBAWS p4de.24xlarge: 8 × A100, $27.45 an hour
Llama 3.1 70B70.55 Bbf16176.4 GBAlibaba Cloud ecs.gn8v-2x.8xlarge: 2 × GPU H, $14.38 an hour
Llama 3.1 8B8.03 Bbf1620.1 GBGoogle Cloud g2-standard-4: 1 × L4, $0.7068 an hour
Phi-414.66 Bbf1636.6 GBAWS g6e.xlarge: 1 × L40S, $1.86 an hour
Mistral 7B7.25 Bbf1618.1 GBGoogle Cloud g2-standard-4: 1 × L4, $0.7068 an hour

Compare what each cloud charges for a model.