LLM Opex Every way to run an LLM, every cloud

Models

Qwen3.8-2.4T-A95B-FP8: GPU memory and hosting cost

Open weights · 2446.18 billion parameters · Mixture of experts · 262,144-token context · licence other

No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.

Run it yourself

It needs 3057.7 GB of GPU memory at fp8, the precision it ships in.

More from Qwen

Similar size, open weights

List prices collected 5 October 2026, refreshed every three hours. No negotiated rates or taxes.