USUS open-weight flash models: latency by region, and price
The same open-weight model is served by several hosts, and how fast it answers depends on where you call it from. We measure 4 endpoints — 2 flash models from OpenAI and Meta — from 2 regions, host by host, region by region; the quickest host in one region is often not the quickest in another. Blended price runs $0.17–$0.77 per 1M tokens, cheapest being gpt-oss-20b on DeepInfra.
Price against measured speed, region by region
Model by model, region by region
gpt-oss-20b OpenAI
tested 9 Sep 2026llama-4-scout Meta
tested 9 Sep 2026One row per host, one column per region. Only latency is coloured —green fastest, redslowest, ring on the best cell — because it’s the only thing that changes with region. Each cell is the rolling median of the last few weekly runs, not a single reading; hover for the p95 tail and the latest raw run. The blended bar is split: solid = input, lighter tail = output. Context windows sit with each spec sheet below.
What's in this tier
Click a model to open the maker's spec sheet.
gpt-oss-20bOpenAI21B · 3.6B31215.2
Spec sheetmaker-declared
- Context window
- 131,072 tokens
- Parameters
- 20,914,757,184≈ 20.9B
- Active experts
- 4 of 32 active per token
- Weight precision
- mxfp4
- License
- Apache 2.0 — open weights, permissivecuratedScope: the licence on the published weights, and nothing else. It is not the maker's acceptable-use policy, which is a separate document, and it is not the contract you sign with whichever host you call — hosts set their own terms, which this index does not read.
- Knowledge cutoff
- June 2024maker's docs
- Modalities
- text
- Max output (official)
- to verify
- Release date
- to verify≈ August 5, 2025third-party catalogue — not the maker's figure
- Maker lifecycle status
- to verify
openai/gpt-oss-20b · config.json @ 6cee5e8 · read 2026-09-07 · 7 of 10 fields published by the maker
llama-4-scoutMeta109B · 17B1310.3
Spec sheetmaker-declared
- Context window
- weights repository under access conditions
- Parameters
- weights repository under access conditions
- Active experts
- weights repository under access conditions
- Weight precision
- weights repository under access conditions
- License
- Llama Community License — open weights, conditionalcuratedScope: the licence on the published weights, and nothing else. It is not the maker's acceptable-use policy, which is a separate document, and it is not the contract you sign with whichever host you call — hosts set their own terms, which this index does not read.
- Knowledge cutoff
- December 2024maker's docs
- Modalities
- weights repository under access conditions
- Max output (official)
- to verify
- Release date
- to verify≈ April 5, 2025third-party catalogue — not the maker's figure
- Maker lifecycle status
- to verify
meta-llama/Llama-4-Scout-17B-16E-Instruct · weights repository @ 92f3b15 · file not readable · read 2026-09-07 · 2 of 10 fields published by the maker
4 of the 15 priced endpoints carry a measured speed; the rest are priced but not yet called. Intelligence, where shown, is a third-party score we report but do not produce.
Prices carry a source and a verification date on each endpoint page. Latency verified 2026-09-09. The other half of the range is on us open-weight flagship models. See also the speed-per-dollar board, why the same model has several prices, and the methodology.
FAQ
Which flash US open-weight model API is cheapest?
As of 2026-09-09, gpt-oss-20b on DeepInfra at $0.17 per 1M tokens blended (input + output). That is the cheapest token, which is not the same as the cheapest answer: two models do not spend the same number of tokens on the same task, so the ranking below settles where to buy a given model, not which model costs least to run. It is not the fastest either: measured from eu-paris it streams 81 tokens/s, against 1446 for the fastest endpoint in this tier.
How is the speed measured?
First-party. Each endpoint is called with the same prompt, streamed, and its time to first token and output throughput recorded — 2 samples per endpoint per run, from eu-paris, and separately from us-east. All regions run at the same hour, because figures taken at different hours are not comparable. Method: stream, prompt=fixe, out=256 tok, samples=5/région, region=eu-paris.
What counts as a "flash" model here?
The lighter open-weight variant each US maker ships alongside its flagship (OpenAI gpt-oss-20b, Meta Llama 4 Scout, Thinking Machines Inkling Small, NVIDIA Nemotron 3 Super) — the maker's own naming, not our ranking. The line is the makers' own — where a maker ships both a full model and a light one, it names the light one "flash". We do not rank these models on quality — this page compares price and measured speed only.