USUS open-weight flash models: latency by region, and price

The same open-weight model is served by several hosts, and how fast it answers depends on where you call it from. We measure 4 endpoints — 2 flash models from OpenAI and Meta — from 2 regions, host by host, region by region; the quickest host in one region is often not the quickest in another. Blended price runs $0.17$0.77 per 1M tokens, cheapest being gpt-oss-20b on DeepInfra.

Price against measured speed, region by region

Measured from EU · Paris2026-09-09 · 10:18 UTC
040080012001600$0.2$0.3$0.5$ / 1M tokens (blended) — pricier →tokens/s measured — faster ↑DeepInfra · gpt-oss-20b — $0.17/1M · 81 tok/s ★ Pareto frontTogether AI · gpt-oss-20b — $0.25/1M · 241 tok/s ★ Pareto frontGroq · gpt-oss-20b — $0.375/1M · 1446 tok/s ★ Pareto frontNovita AI · llama-4-scout — $0.77/1M · 30 tok/s (dominated)GroqTogether AIDeepInfra
The dashed line is the Pareto front: at every price, the fastest host you can buy. Points below it are dominated — some other host is cheaper andquicker. Price is free public data; the throughput that draws this front is our own first-party measurement from EU · Paris, 2026-09-09 · 10:18 UTC.

Model by model, region by region

gpt-oss-20b OpenAI

tested 9 Sep 2026

3 hosts · price spread 2.2× · throughput spread 17.8×

HostEU · ParisTTFT mstok/sBlended $/1M
Groq339Groq · EU · Parismedian, 3 runs339 msp95 (tail)419 mslatest run319 ms1446$0.375
DeepInfra491DeepInfra · EU · Parismedian, 1 run491 msp95 (tail)503 mslatest run491 ms81$0.17
Together AI3184Together AI · EU · Parismedian, 1 run3184 msp95 (tail)3474 mslatest run3184 ms241$0.25

llama-4-scout Meta

tested 9 Sep 2026

1 hosts

HostEU · ParisTTFT mstok/sBlended $/1M
Novita AI707Novita AI · EU · Parismedian, 1 run707 msp95 (tail)785 mslatest run707 ms30$0.77

One row per host, one column per region. Only latency is coloured —green fastest, redslowest, ring on the best cell — because it’s the only thing that changes with region. Each cell is the rolling median of the last few weekly runs, not a single reading; hover for the p95 tail and the latest raw run. The blended bar is split: solid = input, lighter tail = output. Context windows sit with each spec sheet below.

What's in this tier

Click a model to open the maker's spec sheet.

gpt-oss-20bOpenAI21B · 3.6B31215.2

Spec sheetmaker-declared

Context window
131,072 tokens
Parameters
20,914,757,184≈ 20.9B
Active experts
4 of 32 active per token
Weight precision
mxfp4
License
Apache 2.0 — open weights, permissivecuratedScope: the licence on the published weights, and nothing else. It is not the maker's acceptable-use policy, which is a separate document, and it is not the contract you sign with whichever host you call — hosts set their own terms, which this index does not read.
Knowledge cutoff
June 2024maker's docs
Modalities
text
Max output (official)
to verify
Release date
to verify≈ August 5, 2025third-party catalogue — not the maker's figure
Maker lifecycle status
to verify

openai/gpt-oss-20b · config.json @ 6cee5e8 · read 2026-09-07 · 7 of 10 fields published by the maker

llama-4-scoutMeta109B · 17B1310.3

Spec sheetmaker-declared

Context window
weights repository under access conditions
Parameters
weights repository under access conditions
Active experts
weights repository under access conditions
Weight precision
weights repository under access conditions
License
Llama Community License — open weights, conditionalcuratedScope: the licence on the published weights, and nothing else. It is not the maker's acceptable-use policy, which is a separate document, and it is not the contract you sign with whichever host you call — hosts set their own terms, which this index does not read.
Knowledge cutoff
December 2024maker's docs
Modalities
weights repository under access conditions
Max output (official)
to verify
Release date
to verify≈ April 5, 2025third-party catalogue — not the maker's figure
Maker lifecycle status
to verify

meta-llama/Llama-4-Scout-17B-16E-Instruct · weights repository @ 92f3b15 · file not readable · read 2026-09-07 · 2 of 10 fields published by the maker

4 of the 15 priced endpoints carry a measured speed; the rest are priced but not yet called. Intelligence, where shown, is a third-party score we report but do not produce.

Prices carry a source and a verification date on each endpoint page. Latency verified 2026-09-09. The other half of the range is on us open-weight flagship models. See also the speed-per-dollar board, why the same model has several prices, and the methodology.

FAQ

Which flash US open-weight model API is cheapest?

As of 2026-09-09, gpt-oss-20b on DeepInfra at $0.17 per 1M tokens blended (input + output). That is the cheapest token, which is not the same as the cheapest answer: two models do not spend the same number of tokens on the same task, so the ranking below settles where to buy a given model, not which model costs least to run. It is not the fastest either: measured from eu-paris it streams 81 tokens/s, against 1446 for the fastest endpoint in this tier.

How is the speed measured?

First-party. Each endpoint is called with the same prompt, streamed, and its time to first token and output throughput recorded — 2 samples per endpoint per run, from eu-paris, and separately from us-east. All regions run at the same hour, because figures taken at different hours are not comparable. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.

What counts as a "flash" model here?

The lighter open-weight variant each US maker ships alongside its flagship (OpenAI gpt-oss-20b, Meta Llama 4 Scout, Thinking Machines Inkling Small, NVIDIA Nemotron 3 Super) — the maker's own naming, not our ranking. The line is the makers' own — where a maker ships both a full model and a light one, it names the light one "flash". We do not rank these models on quality — this page compares price and measured speed only.