LLM Opex Every way to run an LLM, every cloud

Who sells this model, where, and at what price?

Pick a model and see what every cloud charges for it, side by side: per token, on a dedicated cluster, or self-hosted on a GPU instance the model actually fits on. Each price is the cloud's own public list price.

List price per million tokens, by model and cloud

Input and output price in US dollars per million tokens, each cloud's cheapest region, collected 5 October 2026. A dash means the cloud does not sell the model per token.
ModelAWSAzureOCICore42Alibaba CloudGoogle Cloud
OpenAI gpt-oss 120B$0.075 in / $0.30 out$0.15 in / $0.60 out$0.15 in / $0.60 out$0.15 in / $0.37 out—$0.09 in / $0.36 out
OpenAI gpt-oss 20B$0.035 in / $0.15 out$0.07 in / $0.30 out$0.07 in / $0.30 out$0.10 in / $0.30 out—$0.07 in / $0.25 out
Llama 3.3 70B$0.72 in / $0.72 out$0.71 in / $0.71 out$0.72 in / $0.72 out——$0.72 in / $0.72 out
Llama 4 Maverick 17B$0.24 in / $0.97 out$0.25 in / $1.00 out$0.72 in / $0.72 out——$0.35 in / $1.15 out
Llama 4 Scout 17B$0.17 in / $0.66 out—$0.72 in / $0.72 out——$0.25 in / $0.70 out
Qwen3 32B$0.075 in / $0.30 out$0.30 in / $1.20 out——$0.16 in / $0.64 out—
Cohere Command A—$0.80 in / $3.20 out$6.24 in / $6.24 out———
DeepSeek V3.1$0.29 in / $0.84 out————$0.60 in / $1.70 out
Llama 3.1 405B$2.40 in / $2.40 out—$10.68 in / $10.68 out———
Mistral Small$1.00 in / $3.00 out——$0.12 in / $0.36 out——
GPT-4o—$0.60 in / $5.00 out—$2.50 in / $10.00 out——
GPT-5—$0.05 in / $0.40 out—$1.25 in / $10.00 out——
Qwen3 235B A22B$0.11 in / $0.44 out———$0.287 in / $1.15 out—
Amazon Nova Lite$0.06 in / $0.24 out—————
Amazon Nova Pro$0.40 in / $1.60 out—————
Claude Haiku 4.5$1.10 in / $5.50 out—————
Gemini 2.5 Flash—————$0.30 in / $2.50 out
Gemini 2.5 Flash Lite—————$0.10 in / $0.40 out
Gemini 2.5 Pro—————$1.25 in / $10.00 out
Llama 3.1 70B$0.72 in / $0.72 out—————
Llama 3.1 8B$0.22 in / $0.22 out—————
Phi-4—$0.075 in / $0.30 out————
Mistral 7B$0.15 in / $0.20 out—————
Mistral Large$0.25 in / $0.75 out—————

The page compares one model in full, with regions, dedicated clusters and self-hosting. Work out a monthly cost or see which hardware fits a model.