LLM Opex Every way to run an LLM, every cloud

Models

Phi-4: price on every cloud

Open weights · 14 billion parameters · 16,384-token context · licence mit

Per token, by cloud

US dollars per million tokens, each cloud's cheapest region.
CloudInput, per 1M tokensOutput, per 1M tokensCheapest region
Azure$0.075$0.30centralus · 13 regions
AWS——not sold per token
OCI——not sold per token
Core42——not sold per token
Alibaba Cloud——not sold per token
Google Cloud——not sold per token

Run it yourself

It needs 36.6 GB of GPU memory at bf16, the precision it ships in.

CloudCheapest instance that fitsGPUsAn hourA month, around the clock
AWSg6e.xlarge1 × L40S 48 GB$1.86$1,359
Alibaba Cloudecs.gn8is.2xlarge1 × L20 48 GB$2.26$1,648
AzureStandard_NC72ds_xl_RTXPRO6000BSE_v61 × RTX PRO 6000 48 GB$2.83$2,066
Google Clouda2-highgpu-1g1 × A100 40 GB$3.67$2,682

At the cheapest per-token price, running it on AWS g6e.xlarge pays off above roughly 10.4 billion tokens a month.

More from Microsoft

Similar size, open weights

Inside the UAE: no cloud sells it in-country.

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.