Llama 3.3 70B Instruct (Cloudflare) API

Cloudflare · serves llama-3.3-70b

Pricing

Input / 1M$0.29
Output / 1M$2.25
Blended / 1M$2.546
Context window131,072 tokens
Max output21,600 tokens
LicenseOpen weights · Llama Community License conditional
Knowledge cutoffDecember 2024 · source

text

Price verified 2026-09-05 (0 days ago) · sourcecorroborated

Sources

SourceInput / 1MOutput / 1MVerified
OpenRouter$0.29$2.252026-09-05
LiteLLM$0.29$2.252026-09-03
models.dev$0.29$2.252026-09-04

Two independent sources agree within 20%.

Latency measured

RegionTTFT p50TTFT p95Total p50Total p95ThroughputSamples
eu-paris0

Time-to-first-token, measured server-side. Measured 2026-09-04 (0 days ago).

Cheaper providers for llama-3.3-70b

ProviderBlended / 1MEndpoint
DeepInfra$0.42view →
Nebius$0.53view →
Novita AI$0.535view →

Compare all 12 providers serving llama-3.3-70b →

Call it

OpenAI-compatible endpoint — set CLOUDFLARE_AI_TOKEN.

curl https://api.cloudflare.com/client/v4/accounts/CLOUDFLARE_ACCOUNT_ID/ai/v1/chat/completions \
  -H "Authorization: Bearer $CLOUDFLARE_AI_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "@cf/meta/llama-3.3-70b-instruct-fp8-fast",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
from openai import OpenAI  # pip install openai

client = OpenAI(
    base_url="https://api.cloudflare.com/client/v4/accounts/CLOUDFLARE_ACCOUNT_ID/ai/v1",
    api_key="...",  # CLOUDFLARE_AI_TOKEN
)
resp = client.chat.completions.create(
    model="@cf/meta/llama-3.3-70b-instruct-fp8-fast",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)

Price unchanged since 2026-09-03 — 1 point on record.

Use this data

Raw JSON for this page: llama-3.3-70b.json. Free to reuse under CC BY 4.0 with attribution to apipriceindex.com.