ANALYSIS · VERIFIED 30 JUL 2026
Context window per dollar: a big window only helps if you can afford to fill it
Context window size gets quoted like a spec sheet number — bigger is better, full stop. But a 2M-token window is only useful if feeding it 2M tokens of input doesn't blow your budget on a single call. This page answers a different question than the rest of this site: not "what's the cheapest model," but "what does it actually cost to use the context window you're being sold."
What it costs to fill each model's window once
Context window size × input price per token, for every model this site tracks. Lower is better — it means more room for less money.
| Model | Context | Input price | Cost to fill window |
|---|---|---|---|
| GPT-4o mini | 128K | $0.15/M | $0.019 |
| Grok 4.1 | 128K | $0.20/M | $0.026 |
| DeepSeek V3 | 128K | $0.27/M | $0.035 |
| Gemini 2.5 Flash | 1M | $0.15/M | $0.15 |
| GPT-4o | 128K | $2.50/M | $0.32 |
| Gemini 3 Flash | 1M | $0.50/M | $0.50 |
| Claude Haiku 4.5 | 200K | $1.00/M | $0.20 |
| Claude Sonnet 5 | 1M | $2.00/M | $2.00 |
| GPT-5.5 | 400K | $5.00/M | $2.00 |
| Gemini 3.1 Pro | 2M | $2.00/M | $4.00 |
| Claude Opus 4.8 | 1M | $5.00/M | $5.00 |
| Claude Fable 5 | 1M | $10.00/M | $10.00 |
The model this actually flatters
Gemini 2.5 Flash is the standout: a 1M-token window — 8x GPT-4o mini's — for a fill cost of $0.15, still cheaper than fully using GPT-4o's much smaller 128K window ($0.32). If your workload genuinely benefits from stuffing in large documents or many retrieved chunks, Gemini 2.5 Flash gives you room to do that without the cost scaling into a different tier entirely.
The model this doesn't flatter
Gemini 3.1 Pro has the largest window this site tracks — 2M tokens — but at $2.00/M input, using all of it costs $4.00 per call. That's not necessarily wrong for a workload that needs the accuracy and headroom, but it's worth knowing the number before assuming "biggest window" means "most affordable way to work with a lot of context." Claude Opus 4.8 and Claude Fable 5 show the same pattern at the very top of the price range — large windows priced for calls where context depth matters more than cost per call.
How we'd actually decide
- Need real headroom for large documents or many retrieved chunks, cost still matters: Gemini 2.5 Flash — the best ratio on this list by a wide margin.
- Context needs are modest (single documents, short conversations): the fill-cost number barely matters — pick on price per token instead.
- Need the largest window available regardless of cost: Gemini 3.1 Pro at 2M, with the understanding that using all of it is a real line item, not a rounding error.
"Cost to fill window" = context window size × input price per token, a one-time input cost with no output tokens included. Prices verified 30 July 2026. Use the calculator for your own exact mix of input and output.