LLM Opex Every way to run an LLM, every cloud

Clouds

Google Cloud: LLM prices, GPUs and regions

Google Cloud sells 9 of the compared models per token and prices 5 kinds of GPU for hosting your own, across 2 regions.

Models sold per token

ModelInput, per 1M tokensOutput, per 1M tokensCheapest region
DeepSeek V3.1$0.60$1.70global
Gemini 2.5 Flash$0.30$2.50global
Gemini 2.5 Flash Lite$0.10$0.40global
Gemini 2.5 Pro$1.25$10.00global
Llama 3.3 70B$0.72$0.72global
Llama 4 Maverick 17B$0.35$1.15global
Llama 4 Scout 17B$0.25$0.70global
OpenAI gpt-oss 120B$0.09$0.36global
OpenAI gpt-oss 20B$0.07$0.25global

GPUs for hosting your own

On demand.
GPUFrom, per GPU-hourCheapest instanceGPUs in itRegion
L4$0.7068g2-standard-41 × 24 GBus-central1
A100$3.48a2-megagpu-16g16 × 40 GBus-central1
RTX PRO 6000$4.50g4-standard-481 × 96 GBus-central1
H200$10.60a3-ultragpu-8g8 × 141 GBus-central1
H100$11.06a3-highgpu-8g8 × 80 GBus-central1

Regions with prices

global, us-central1

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.