ANALYSIS · VERIFIED 30 JUL 2026

Context window per dollar: a big window only helps if you can afford to fill it

Context window size gets quoted like a spec sheet number — bigger is better, full stop. But a 2M-token window is only useful if feeding it 2M tokens of input doesn't blow your budget on a single call. This page answers a different question than the rest of this site: not "what's the cheapest model," but "what does it actually cost to use the context window you're being sold."

What it costs to fill each model's window once

Context window size × input price per token, for every model this site tracks. Lower is better — it means more room for less money.

ModelContextInput priceCost to fill window
GPT-4o mini128K$0.15/M$0.019
Grok 4.1128K$0.20/M$0.026
DeepSeek V3128K$0.27/M$0.035
Gemini 2.5 Flash1M$0.15/M$0.15
GPT-4o128K$2.50/M$0.32
Gemini 3 Flash1M$0.50/M$0.50
Claude Haiku 4.5200K$1.00/M$0.20
Claude Sonnet 51M$2.00/M$2.00
GPT-5.5400K$5.00/M$2.00
Gemini 3.1 Pro2M$2.00/M$4.00
Claude Opus 4.81M$5.00/M$5.00
Claude Fable 51M$10.00/M$10.00

The model this actually flatters

Gemini 2.5 Flash is the standout: a 1M-token window — 8x GPT-4o mini's — for a fill cost of $0.15, still cheaper than fully using GPT-4o's much smaller 128K window ($0.32). If your workload genuinely benefits from stuffing in large documents or many retrieved chunks, Gemini 2.5 Flash gives you room to do that without the cost scaling into a different tier entirely.

The model this doesn't flatter

Gemini 3.1 Pro has the largest window this site tracks — 2M tokens — but at $2.00/M input, using all of it costs $4.00 per call. That's not necessarily wrong for a workload that needs the accuracy and headroom, but it's worth knowing the number before assuming "biggest window" means "most affordable way to work with a lot of context." Claude Opus 4.8 and Claude Fable 5 show the same pattern at the very top of the price range — large windows priced for calls where context depth matters more than cost per call.

How we'd actually decide

"Cost to fill window" = context window size × input price per token, a one-time input cost with no output tokens included. Prices verified 30 July 2026. Use the calculator for your own exact mix of input and output.