Prompt caching: who offers it, and what it actually saves
Most of a real workload's input tokens repeat — a system prompt, a retrieved document, a few-shot preamble. Providers that cache those tokens bill them at a fraction of the normal input price on the next call. Across this index,364 of 539 priced endpoints publish a cached-input price, spread over 39 providers. Every figure below is computed from those published prices — nothing here is hand-written.
A worked example
Take a RAG / long context job — 8,000 input /500 output tokens per request, 1,000 requests — on Claude Fable 5.1 (Anthropic) (Anthropic), where 80% of the input is reused context and hits the cache:
Full price: $105.00 → with caching: $42.60.
−$62.40 (59% cheaper)
Cache hit rate:80% → cost$42.60(59% saved)
The higher your reuse, the more caching pays off. Below ~30% hit rate the gain is marginal; a RAG or agent loop that re-sends the same context every turn is where it compounds.
Biggest input savings from caching
Endpoints ranked by how much their cached-input price undercuts their normal input price.
Savings by provider
Median input discount and how many cached endpoints each provider publishes.
| Provider | Median input discount | Cached endpoints |
|---|---|---|
| DeepSeek | −97% | 3 |
| Mistral AI | −90% | 10 |
| Fireworks AI | −90% | 9 |
| −90% | 7 | |
| Amazon Bedrock | −90% | 7 |
| Chutes | −90% | 6 |
| Azure OpenAI | −90% | 5 |
| Anthropic | −90% | 4 |
| Minimax | −90% | 4 |
| SambaNova | −90% | 1 |
| xAI | −85% | 1 |
| Modal | −84% | 4 |
| BaseTen | −83% | 10 |
| Moonshot AI | −83% | 3 |
| Z.AI | −82% | 10 |
| Cloudflare | −82% | 8 |
| OpenAI | −82% | 6 |
| Reka | −82% | 5 |
| Morph | −82% | 4 |
| Novita AI | −81% | 27 |
| GMICloud | −81% | 14 |
| SiliconFlow | −81% | 13 |
| Alibaba | −81% | 9 |
| Together AI | −81% | 7 |
| Baidu | −81% | 7 |
| Friendli | −81% | 6 |
| DeepInfra | −80% | 29 |
| Venice | −80% | 23 |
| DigitalOcean | −80% | 15 |
| Relace | −78% | 2 |
| Phala | −73% | 14 |
| AkashML | −63% | 6 |
| Parasail | −62% | 25 |
| AtlasCloud | −59% | 21 |
| Io Net | −53% | 5 |
| Crusoe | −50% | 7 |
| Groq | −50% | 5 |
| CoreWeave | −23% | 20 |
| Cerebras | −0% | 2 |
Prices per 1M tokens (USD), from each provider's published cached-input rate. Verified 2026-09-05.
FAQ
Which LLM API providers offer prompt caching?
In this index, 364 of 539 priced endpoints publish a cached-input price, spread across 39 providers. Coverage and the size of the discount vary by provider.
How much does prompt caching actually save?
On a RAG / long context workload on Claude Fable 5.1 (Anthropic) (Anthropic) where 80% of the input is reused context, caching cuts the bill by 59%. Savings scale with your cache-hit rate — below roughly 30% reuse the gain is marginal.
What kind of workload benefits most from caching?
Anything that re-sends the same context on every call — RAG over a fixed corpus, an agent loop replaying its history, or a long system prompt or few-shot preamble. One-off, unique prompts see little benefit because there is nothing to reuse.