Capability

Long-context models

Models that take 200,000 tokens or more in a single request — roughly a long book, a full codebase, or a year of support tickets. Big enough that chunking stops being mandatory and the architecture of your product can change.

Models
57
Vendors
6
Largest context
1.0M
With cached input
47/57

A big window is a big invoice

The window is a ceiling, not an allowance. Filling a million-token context costs a million tokens of input at that model's rate, every single call — so the question is never "does it fit" but "what does it cost each time, multiplied by your traffic". Run that number before you design around it.

Caching is what makes it affordable

If the large part of your prompt is the same between calls — a manual, a schema, a contract — a model with a cached-input rate bills that part at a fraction after the first request. That single detail decides whether a long-context design is viable at volume, and it is why the cheapest list price is often not the cheapest model.

Long context versus retrieval

Retrieval is still cheaper per call and still wins on very large corpora. Long context wins on simplicity, on questions that need the whole document at once, and on anything where a retrieval miss produces a confidently wrong answer. Most production systems end up using both — retrieval to narrow, a long window to reason over what came back.

Your plan

Vendor cost is the list price. Your price is that plus your plan margin — the same figures that appear on your invoice.

Every model here, cheapest first

ModelVendorContextIn / 1MOut / 1MCached inYour price in
GPT-5 nano
gpt-5-nano
OpenAI400K$0.05$0.4$0.005$0.054Details
GPT-4.1 nano
gpt-4.1-nano
OpenAI1M$0.1$0.4$0.025$0.108Details
Ministral 8B
ministral-8b-latest
Mistral AI262K$0.15$0.15$0.161Details
Mistral Small 3
mistral-small-latest
Mistral AI262K$0.15$0.6$0.161Details
GPT-5.4 nano
gpt-5.4-nano
OpenAI400K$0.2$1.25$0.02$0.215Details
GPT-5.6 Luna
gpt-5.6-luna
OpenAI400K$0.2$1.20$0.02$0.215Details
Gemini 3.1 Flash-Lite
gemini-3.1-flash-lite
Google AI1.0M$0.25$1.50$0.025$0.269Details
GPT-5 mini
gpt-5-mini
OpenAI400K$0.25$2$0.025$0.269Details
Codestral
codestral-latest
Mistral AI256K$0.3$0.9$0.323Details
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite
Google AI1.0M$0.3$2.50$0.03$0.323Details
GPT-4.1 mini
gpt-4.1-mini
OpenAI1M$0.4$1.60$0.1$0.43Details
Gemini 3 Flash Preview
gemini-3-flash-preview
Google AI1.0M$0.5$3$0.05$0.538Details
Gemini 3.6 Flash
gemini-3.6-flash
Google AI1.0M$0.75$3.75$0.075$0.806Details
Gemini 3.7 Flash
gemini-3.7-flash
Google AI1.0M$0.75$3.75$0.075$0.806Details
GPT-5.4 mini
gpt-5.4-mini
OpenAI400K$0.75$4.50$0.075$0.806Details
Claude Haiku 4.5
claude-haiku-4-5
Anthropic200K$1$5$0.1$1.08Details
grok-build-0.1
grok-build-0.1
xAI256K$1$2$0.2$1.08Details
o3-mini
o3-mini
OpenAI200K$1.10$4.40$0.55$1.18Details
o4-mini
o4-mini
OpenAI200K$1.10$4.40$0.275$1.18Details
Gemini 2.5 Pro
gemini-2.5-pro
Google AI1.0M$1.25$10$0.125$1.34Details
GPT-5
gpt-5
OpenAI400K$1.25$10$0.125$1.34Details
gpt-5.1
gpt-5.1
OpenAI400K$1.25$10$0.125$1.34Details
grok-4.20-0309-non-reasoning
grok-4.20-0309-non-reasoning
xAI1M$1.25$2.50$0.2$1.34Details
grok-4.20-0309-reasoning
grok-4.20-0309-reasoning
xAI1M$1.25$2.50$0.2$1.34Details
grok-4.20-multi-agent-0309
grok-4.20-multi-agent-0309
xAI1M$1.25$2.50$0.2$1.34Details
grok-4.3
grok-4.3
xAI1M$1.25$2.50$0.2$1.34Details
Gemini 3.5 Flash
gemini-3.5-flash
Google AI1.0M$1.50$9$0.15$1.61Details
gpt-5.2
gpt-5.2
OpenAI400K$1.75$14$0.175$1.88Details
GPT-5.3 Codex
gpt-5.3-codex
OpenAI400K$1.75$14$0.175$1.88Details
Claude Sonnet 5
claude-sonnet-5
Anthropic1M$2$10$0.2$2.15Details
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
Google AI1.0M$2$12$0.2$2.15Details
GPT-4.1
gpt-4.1
OpenAI1M$2$8$0.5$2.15Details
GPT-5.6 Terra
gpt-5.6-terra
OpenAI400K$2$12$0.2$2.15Details
grok-4.5
grok-4.5
xAI500K$2$6$0.3$2.15Details
grok-4.6
grok-4.6
xAI500K$2$6$0.5$2.15Details
o3
o3
OpenAI200K$2$8$0.5$2.15Details
GPT-5.4
gpt-5.4
OpenAI400K$2.50$15$0.25$2.69Details
Claude Sonnet 4.5
claude-sonnet-4-5
Anthropic200K$3$15$0.3$3.23Details
Claude Sonnet 4.6
claude-sonnet-4-6
Anthropic1M$3$15$0.3$3.23Details
Grok 4
grok-4
xAI256K$3$15$0.75$3.23Details
Sonar Pro
sonar-pro
Perplexity200K$3$15$3.23Details
GPT-5.6 Sol
gpt-5.6-sol
OpenAI400K$4$20$0.4$4.30Details
ChatGPT (chat-latest)
chat-latest
OpenAI400K$5$30$0.5$5.38Details
Claude Opus 4.5
claude-opus-4-5-20251101
Anthropic200K$5$25$0.5$5.38Details
Claude Opus 4.6
claude-opus-4-6
Anthropic1M$5$25$0.5$5.38Details
Claude Opus 4.7
claude-opus-4-7
Anthropic1M$5$25$0.5$5.38Details
Claude Opus 4.8
claude-opus-4-8
Anthropic1M$5$25$0.5$5.38Details
Claude Opus 5
claude-opus-5
Anthropic1M$5$25$0.5$5.38Details
GPT-5.5
gpt-5.5
OpenAI400K$5$30$0.5$5.38Details
Claude Fable 5
claude-fable-5
Anthropic1M$10$50$1$10.75Details
Claude Fable 5.1
claude-fable-5-1
Anthropic1M$10$50$0.25$10.75Details
gpt-5-pro
gpt-5-pro
OpenAI400K$15$120$16.13Details
gpt-5.2-pro
gpt-5.2-pro
OpenAI400K$21$168$22.58Details
GPT-5.4 Pro
gpt-5.4-pro
OpenAI400K$30$180$32.26Details
GPT-5.5 Pro
gpt-5.5-pro
OpenAI400K$30$180$32.26Details
Lyria 3 Clip (30s)
lyria-3-clip-preview
Google AI1.0MBilled per media unit — see detailsDetails
Lyria 3 Pro (full song)
lyria-3-pro-preview
Google AI1.0MBilled per media unit — see detailsDetails

Your price column is the vendor cost +7%.

One key, every model on this page.

Change the model with a routing rule instead of a deploy, and bill every call to the customer who made it.