US open-weight flagship models: latency by region, and price

The same open-weight model is served by several hosts, and how fast it answers depends on where you call it from. We measure 5 endpoints — 1 flagship models from OpenAI — from 2 regions, host by host, region by region; the quickest host in one region is often not the quickest in another. Blended price runs $0.207$1.1 per 1M tokens, cheapest being gpt-oss-120b on DeepInfra.

Price against measured speed

0588117517632350$0.3$0.5$0.8$1$ / 1M tokens (blended) — pricier →tokens/s measured — faster ↑DeepInfra · gpt-oss-120b — $0.207/1M · 35 tok/s ★ Pareto frontNovita AI · gpt-oss-120b — $0.3/1M · 127 tok/s ★ Pareto frontGroq · gpt-oss-120b — $0.75/1M · 775 tok/s ★ Pareto frontTogether AI · gpt-oss-120b — $0.75/1M · 161 tok/s (dominated)Cerebras · gpt-oss-120b — $1.1/1M · 2159 tok/s ★ Pareto frontCerebrasGroqNovita AIDeepInfra
The dashed line is the Pareto front: at every price, the fastest host you can buy. Points below it are dominated — some other host is cheaper andquicker. Price is free public data; the throughput that draws this front is our own first-party measurement from eu-paris.

Model by model, region by region

gpt-oss-120b OpenAI

tested 8 Sep 2026

5 hosts · price spread 5.3× · throughput spread 61.7×

HostEU · ParisTTFT msUS · EastTTFT mstok/sBlended $/1M
Cerebras242Cerebras · EU · Parismedian, 5 runs242 msp95 (tail)413 mslatest run300 ms228Cerebras · US · Eastmedian, 2 runs228 msp95 (tail)344 mslatest run260 ms2159$1.1
Groq393Groq · EU · Parismedian, 5 runs393 msp95 (tail)492 mslatest run432 ms463Groq · US · Eastmedian, 2 runs463 msp95 (tail)1010 mslatest run491 ms775$0.75
Novita AI481Novita AI · EU · Parismedian, 3 runs481 msp95 (tail)577 mslatest run8381 mslatest run set aside — off week1761Novita AI · US · Eastmedian, 2 runs1761 msp95 (tail)6528 mslatest run2903 mslatest run set aside — off week127$0.3
Together AI704Together AI · EU · Parismedian, 3 runs704 msp95 (tail)2365 mslatest run704 ms1497Together AI · US · Eastmedian, 2 runs1497 msp95 (tail)2605 mslatest run777 ms161$0.75
DeepInfra742DeepInfra · EU · Parismedian, 3 runs742 msp95 (tail)1500 mslatest run712 ms607DeepInfra · US · Eastmedian, 2 runs607 msp95 (tail)1053 mslatest run577 ms35$0.207

One row per host, one column per region. Only latency is coloured —green fastest, redslowest, ring on the best cell — because it’s the only thing that changes with region. Each cell is the rolling median of the last few weekly runs, not a single reading; hover for the p95 tail and the latest raw run. The blended bar is split: solid = input, lighter tail = output. Context windows sit with each spec sheet below.

What's in this tier

Click a model to open the maker's spec sheet.

gpt-oss-120bOpenAI120B51624.1

Spec sheetmaker-declared

Context window
131,072 tokens
Parameters
116,829,156,672≈ 116.8B
Active experts
4 of 128 active per token
Weight precision
mxfp4
License
Apache 2.0 — open weights, permissivecuratedScope: the licence on the published weights, and nothing else. It is not the maker's acceptable-use policy, which is a separate document, and it is not the contract you sign with whichever host you call — hosts set their own terms, which this index does not read.
Knowledge cutoff
June 2024maker's docs
Modalities
text
Max output (official)
to verify
Release date
to verify≈ August 5, 2025third-party catalogue — not the maker's figure
Maker lifecycle status
to verify

openai/gpt-oss-120b · config.json @ b5c939d · read 2026-09-07 · 7 of 10 fields published by the maker

5 of the 16 priced endpoints carry a measured speed; the rest are priced but not yet called. Intelligence, where shown, is a third-party score we report but do not produce.

Prices carry a source and a verification date on each endpoint page. Latency verified 2026-09-08. See also the speed-per-dollar board, why the same model has several prices, and the methodology.

FAQ

Which flagship US open-weight model API is cheapest?

As of 2026-09-08, gpt-oss-120b on DeepInfra at $0.207 per 1M tokens blended (input + output). That is the cheapest token, which is not the same as the cheapest answer: two models do not spend the same number of tokens on the same task, so the ranking below settles where to buy a given model, not which model costs least to run. It is not the fastest either: measured from eu-paris it streams 35 tokens/s, against 2159 for the fastest endpoint in this tier.

How is the speed measured?

First-party. Each endpoint is called with the same prompt, streamed, and its time to first token and output throughput recorded — 12 samples per endpoint per run, from eu-paris, and separately from us-east. All regions run at the same hour, because figures taken at different hours are not comparable. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.

What counts as a "flagship" model here?

Each US maker's main open-weight model (Meta Llama 4 Maverick, OpenAI gpt-oss-120b) — the maker's own top open-weight model, not our ranking. The line is the makers' own — where a maker ships both a full model and a light one, it names the light one "flash". We do not rank these models on quality — this page compares price and measured speed only.