How LLM Arena works

← back to the app A ten-chapter walkthrough · 9 minutes at 1× · the text stays; play runs the diagrams
AWSBedrock · EC2 AzureFoundry · VMs OCIGenAI · GPUs Core42Compass AlibabaModel Studio GoogleVertex · GCE One model, one workload public list prices only, refreshed every 3 hours per tokenmanaged inference self-hosteda GPU instance you rent dedicated clusterOCI AI units Calculator · Compare · Fit · Trending · Data
AWS bulk price listpricing.us-east-1.amazonaws.com Azure Retail Prices APIprices.azure.com OCI price list APIapexapps.oracle.com/cetools Core42 Compass docspricing · models · retirements Alibaba price pageModel Studio · ECS API Google pricing pagesVertex AI · Compute Engine Oracle GenAI docsmodels · clusters · imports Hugging Face APIsizes · native precision collectorsone per cloudno credentials needed normalize$ per 1M tokens$ per hour catalogues, not priceswhat OCI offers · how big each model is raw_unit + raw_price kept on every row
prices — the live table AWS · 19,964 rows Azure · 36,058 rows OCI · 122 rows Core42 · 104 rows Alibaba · 1,322 rows Google · ~330 rows exactly one snapshot per cloud prices_archive the newest 30 per cloud, indexed by snapshot only A collection writes a new snapshot… …and only marks it current once every row is in. A failed or empty run never replaces the previous one. Restore from Admin › Data an archived snapshot swaps into the live table in one transaction; new collections wait in the archive until you return. cron: 0 */3 * * * → every 3 hours
every meter 55,000+ rows unit normalizationper 1K → per 1M · chars → tokensunknown unit = "other", never guessed model aliases"Llama 3.3 70B Instruct" · "Large Meta"→ one canonical model kept apartservice tiers stay separatetraining meters stay out current_priceswhat every pageand answer reads OCI: the price list is not a menuit still bills retired models — what OCI offers comes from Oracle's docs
needs = parameters × bytes per weight × 1.25 gpt-oss-120b · 116.8B · int4 → 73 GB Llama 3.3 70B · 70.6B · bf16 → 176 GB native precision comes from the repo's quantization, not its storage type U8 storage + quant_method: mxfp4 → int4 (0.5 B/weight), not int8 (1 B/weight) GPU memory · the model's need drawn as a line 73 GB 1× A100 80 GBfits alone 4× A10 24 GB · PCIe96 GB in total — not offered each token would cross a 64 GB/s bus 8× H100 80 GB · NVSwitchmay be split across them OCI clustershape from Oracle's docsprice = AI-unit rate × units
1M100M10B1T tokens / month $1M$1K$1 per token: linear self-hosted: a staircase — whole instances break-even · 3.3B / month your workload drag it: every figure follows Export PDF → for the customer
price feedsGPU column · NC/ND/NV sizesper-GPU-hour meters not in the catalogueg3.16xlarge · NC128…RTXPRO6000OCI RTX PRO 6000 evidenceprovider page · GPU maker pageNVSwitch · NVLink · PCIe counted a person decidesGPU · count · memory · joined by · the source pageNVLink on a multi-GPU box must be confirmed explicitly hardware_overridesmerged into the catalogue at once priced from the next refresh default: unknown → not shown; unverified → PCIe;PCIe box → only models that fit in one of its GPUs
every 3 hours00:00 · 03:00 · 06:00 … UTC 1 · collectsix clouds, in turn 2 · new snapshotold one → archive 3 · diff → Trendingcuts, rises, new models 4 · cataloguesHugging Face · Oracle docs Admin · run nowsame job, live log Admin · hardwareon demand, 2–3 min
gpt-oss-120bOCI Generative AI · Frankfurt tool definitionscompare · fit · cost… tools.pythe numbers, from the data tool results the conversationearlier answers go back with a follow-up every number = a tool call · tokens counted per question ▶ Demoevery page
Dataauto-refresh on/off · retention · purgerefresh now · restore a snapshot Vendors · Modelsvendor on/off · per token · self-hosted · dedicatedhide a model Websitenotice · OCI banner · calculator defaultsTrending switches · Ask budget Hardwarefind unknown GPU hardwareNVLink evidence · add or override audit logwho · what · when — every change, live on the next request hide, switch off, restore — never type a price in admin headerwhich server, which build DEV v1.185 · chris-dev PROD v1.186 · main
0:00 / 0:00

Where to see each part live

  • Data tab: how many prices each cloud contributes, when each was last collected, and the sources.
  • Admin › Data: switch the automatic refresh off or on, set the retention and purge, run a refresh, restore a snapshot.
  • Admin header: DEV or PROD, the version, the commit and the branch this server runs.
  • Admin › Website: every switch for the pages visitors see, grouped by page; the OCI position banner is under Every page.
  • Admin › Vendors and Models: vendors, their hosting modes and models on or off.
  • Calculator, Compare, Fit: the Location step with the country found for you; Export PDF in the second panel.
  • Calculator: the OCI position banner, egress on the self-hosted and dedicated cards.
  • Trending (Opex): Download newsletter, the issue as a PDF. Demo, in the header of every page that has one, including the landing page.
  • Admin › Hardware: check for new hardware, add or override entries.
  • Fit: the memory rule and the interconnect tag on each instance.

The four environments

  • Dubai production: LLM Opex, chris-prod, from main.
  • Dubai development: LLM Opex, chris-dev, from branch chris-dev.
  • Frankfurt production: LLM Arena, oci-prod, from main.
  • Frankfurt development: LLM Arena, oci-dev, from branch oci-dev.
  • A change is built on a development server, tested there, then promoted: ./deploy.sh promote chris or promote oci.

Rules that never change

  • No number is ever stated from memory; all arithmetic lives in one place (tools.py) with a test.
  • A collection that fails never replaces the previous snapshot.
  • A unit that cannot be normalized is kept as other, never guessed.
  • Unverified hardware is PCIe; unknown hardware is not shown.

Keyboard

  • Space play / pause · ← → chapters · 1–9, 0 jump

Calculation reference

Every number on every page is computed in one module, tools.py (with pricing/egress.py and pricing/benchmarks.py), and tested. These are the formulas, with their constants and a worked example. Prices are list prices; a month is 730 hours for hourly options and 30 days for per-token volumes unless a "days" setting says otherwise.

1 · Memory a model needs

VRAM (GB) = parameters (billions) × bytes per weight × (1 + overhead)

Bytes per weight: fp32 4 · fp16 2 · bf16 2 · fp8 1 · int8 1 · int4 0.5 (MXFP4 and NVFP4 count as int4). Overhead is the KV-cache and activation allowance: 25% by default, adjustable on the Fit page.

An instance fits when VRAM ≤ GPU count × memory per GPU. A model larger than one GPU is offered only on instances whose GPUs are joined by NVLink, NVSwitch or Infinity Fabric. When nothing fits, shortfall = needed − largest instance.

Example: 120B at int4 = 120 × 0.5 × 1.25 = 75 GB; 70B at fp16 = 70 × 2 × 1.25 = 175 GB.

With a use case (Chat, RAG, Agents, Summarise and batch, or your own numbers) the flat allowance is replaced by a worked one, when the model's architecture is known:

VRAM = weights × 1.10 + KV cache KV cache (GB) = bytes per token × tokens kept × requests at once ÷ 10⁹ bytes per token and layer = 2 × KV heads × head size × 2 (grouped-query attention) = (kv_lora_rank + rope dim) × 2 (DeepSeek's latent attention)

The 1.10 is a working margin (activations, CUDA graphs, fragmentation); the 2 is bytes per cached value (16-bit). Layers that attend over a sliding window keep only the last window tokens, and the context is held to the model's own limit. The presets: Chat 8K × 16 · RAG 32K × 16 · Agents and coding 128K × 4 · Summarise and batch 16K × 32. A model whose layers and heads are not readable (a gated repository) keeps the flat allowance, and the page says so.

Example: gpt-oss-120b (36 layers, 18 full attention, 8 KV heads of 64) for RAG: 18 × 2 × 8 × 64 × 2 = 36,864 bytes a token × 32,768 × 16 ≈ 19.4 GB of KV cache, so 58.4 × 1.10 + 19.4 = 83.7 GB, which no longer fits one 80 GB GPU.

2 · Per-token monthly cost

input tokens/day = output tokens/day × input:output ratio monthly = days × ( input/day ÷ 10⁶ × $in + output/day ÷ 10⁶ × $out )

Defaults: 30 days, 2,000,000 output tokens a day, ratio 3. The chart's per-token line is cost(t) = t × ($out + ratio × $in) ÷ 10⁶ for t output tokens a month: a straight line.

Example: 2M out/day, ratio 3 → 6M in/day → 180 × $in + 60 × $out a month (with $in and $out per 1M).

3 · Which price is used per token

score(region, tier) = $out + ratio × $in (lowest wins)

Rows are grouped by region and service tier; only groups that price both directions compete, so input and output always come from the same region and tier. Cached-input and batch prices come from the chosen group, else the cheapest accepted row. Characters are converted at 4 per token ($/1M = price × 10⁶ ÷ tokens per unit; 10,000 characters = 2,500 tokens). A cloud with no price in the chosen region is "not sold there", never priced from another.

Example: $0.0010 per 10,000 characters = 0.001 × 10⁶ ÷ 2,500 = $0.40 per 1M tokens.

4 · Self-hosted and hourly instances

capacity per instance = tokens/s × 3600 × 730 × busy needed = output tokens/day × days instances = max(1, ⌈ needed ÷ (capacity × days/30) ⌉) monthly = $/hour × 730 × instances × days/30 effective $/1M out = monthly ÷ (tokens ÷ 10⁶)

busy is the GPU busy time you set (50% by default) for self-hosting, and 100% for a dedicated cluster, which is sold as capacity. The chart draws an hourly option as a staircase: n × $/hour × 730 × days/30 + egress, where n steps up each time the volume outgrows one instance.

The instance chosen is the cheapest that fits, one per instance type at its cheapest region. A 0.5× / 1× / 2× throughput band is shown beside each self-hosted figure.

Example: 2,500 tokens/s at 50% busy = 2,500 × 3,600 × 730 × 0.5 = 3.285 × 10⁹ tokens a month per instance.

5 · Throughput: which figure is used

  • Yours, if you set one for that option.
  • Oracle's benchmark, dedicated clusters only: the scenario whose input:output ratio is nearest (Random Length 480/300, Chat 1, Generation Heavy 0.1, RAG 10, by log distance), at the peak output tokens/s across concurrency levels.
  • NVIDIA's TensorRT-LLM measurement for the model and GPU: a dense model within ×1.5 of a measured size, a mixture-of-experts model within ×2 of a measured active size; the measured precision at least as heavy as yours. tokens/s = per-group tokens/s × ⌊GPUs ÷ tensor-parallel⌋.
  • An estimate, labelled as one: log-log interpolation between measured sizes (size = parameters, or √(active × total) for mixture-of-experts), beyond them per GPU = ref × size_ref ÷ size; a precision factor have ÷ want bytes; a GPU not measured is scaled from the H100 by bandwidth ÷ 3,350 GB/s; total = per GPU × GPU count.
  • The admin default, 2,500 tokens/s, when nothing else applies.

6 · OCI dedicated AI cluster

$/hour = $ per AI unit-hour × AI units in the shape × replicas

The rate is Oracle's single "AI Unit Per Hour" meter for the vendor; an imported model uses the "Model Import" meter. The monthly figure then follows formula 4 with busy = 100%. Oracle's 744 unit-hour minimum commitment on pretrained clusters is shown on the card; it is not part of the monthly arithmetic. Imported models carry none. The suggested shape is Oracle's minimum import shape, else the cheapest published shape, else the cheapest that passes the memory rule.

7 · Egress

bytes per output token = B + ratio × E ÷ 2,000 GB = output tokens/month × bytes per output token ÷ 10⁹ + extra GB cost = Σ over tiers of (min(GB, to) − from) × $/GB

B (bytes per output token) and E (response envelope): streamed 200 and 1,500 · text 5 and 1,000 · same region 0 and 0. One request per 2,000 input tokens, so requests = input tokens ÷ 2,000. The Calculator's tick box uses streamed.

Tiers are cumulative and the free allowance is a $0 first tier: AWS 100 GB, Azure 100 GB, OCI 10 TB (10,240 GB), Alibaba 200 GB, Google about 1 GB. Core42 publishes no egress price. The zone is the one the region belongs to; with "anywhere" it is the zone that costs least at that volume, and a dedicated cluster is limited to the zones its shape is sold in.

Egress is added to self-hosted and dedicated-cluster options only. Per-token endpoints are billed per token and carry none.

Example: streamed, 60M output tokens a month, ratio 3 → 200 + 3 × 1,500 ÷ 2,000 = 202.25 bytes per token → 12.1 GB (as text: 6.5 bytes, 0.39 GB).

8 · Break-even

no egress: tokens/month = ($/hour × 730) ÷ ($out + ratio × $in) × 10⁶ with egress: ($out + ratio × $in) × t ÷ 10⁶ = $/hour × 730 + egress cost(t)

With egress the crossing is found by bisection (80 steps; the search widens ×4 up to 10¹⁶ tokens). If the egress a token causes costs more than its per-token price there is "no break-even". On the chart, 90 log-spaced volumes from 10⁵ to 10¹² tokens a month are priced; the winner at each is the cheapest (ties go to the lower option); where the winner changes the volume is refined by 40 bisections. A staircase step can therefore appear as a break-even.

Margin: runner-up cost − winner cost at your workload.

9 · Regions

Each vendor has one choice: any (the cheapest region), in-country (the UAE) or in-country:XX (another country), or one region id. In country accepts only that vendor's regions in the country (UAE: AWS me-central-1 · Azure uaenorth, uaecentral · OCI me-dubai-1, me-abudhabi-1 · Core42 uae · Alibaba me-east-1 · Google none). The country comes from a hand-written region-to-country table; on Opex it is found in the browser from the time zone (nothing is sent) and you can change it.

OCI prices are global, so whether it sells a model in a country is Oracle's documentation's call: a region named for that country or one of its cities.

10 · Trending

price change = (new − old) ÷ old (moves under 0.5% are ignored) importance = weight(kind) + min(40, round(|change| × 100)) + bonuses

Weights: price cut or rise 60 · model added 55 · retirement 50 · announcement 30. Bonuses: +15 an OCI angle, +10 it reaches the UAE regions, and for a retirement +15 within 30 days, +5 within 90, +10 if the app prices the model.

What it means in money: an hourly meter is price × 730; a per-token meter is per-day tokens × 30 ÷ 10⁶ × price at the default workload. Withdrawn is reported only when a row is in an older snapshot and missing from both of the two newest: one skipped collection is not news. A snapshot missing a whole section is kept aside as partial. A retirement is shown only with a firm date, today or later.

11 · OCI position

Rank = 1 + the number of options whose monthly cost is lower than OCI's best by more than $0.005. A tie is another option within $0.005, shown as "ties for cheapest"; "wins" only when OCI is strictly cheapest, and the margin is runner-up − OCI.

When OCI is not cheapest, the chart's break-evens between OCI's best and the cheapest say above or below which volume OCI takes over, or "no crossover" between 100K and 1T tokens. On Compare the score is $out + ratio × $in per 1M; on Fit, the cheapest $/hour that fits.

12 · Other figures you see

  • Monthly from hourly (24×7): $/hour × 730, wherever an instance is shown with a monthly figure.
  • Days: hourly cost and capacity both scale by days ÷ 30; per-token volume scales by days.
  • Input tokens a month: output tokens/day × days × ratio.
  • Ask's own cost: (input tokens × rate_in + output tokens × rate_out) ÷ 10⁶ at OCI's $0.15 / $0.60 per 1M.