USE CASE · VERIFIED 2 OCT 2026

What's the cheapest AI API for summarizing long documents?

For a workload of 500M input and 20M output tokens a month, GPT-6 Luna from OpenAI is the cheapest option at $60/mo. Close behind are Qwen3.8 Flash at $84/mo and GLM 5.3 Flash at $85/mo — both also budget-tier models with 1M context windows built for exactly this input-heavy, output-light pattern.

Worked example: 500M input, 20M output tokens/month

Long reports and transcripts in, compact summaries out. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.

GPT-6 Luna$60/mo
Qwen3.8 Flash$84/mo
GLM 5.3 Flash$85/mo
GPT-4o mini$87/mo
DeepSeek V4.1 Flash$87/mo
DeepSeek V3$143/mo
Grok 4.7$1,120/mo
Gemini 3.1 Pro$1,240/mo

Why this workload favors budget models so heavily

Summarization is almost entirely an input problem: you're feeding in 500M tokens of reports or transcripts and getting back only 20M tokens of compact summary, a 25:1 ratio. That means input price dominates your bill far more than output price. Models like GPT-6 Luna ($0.1/$0.5), Qwen3.8 Flash ($0.15/$0.47), and GLM 5.3 Flash ($0.15/$0.5) win precisely because they keep input costs near rock bottom, while expensive output pricing on flagship models barely matters when output volume is this small relative to input. If your ratio shifted toward more output per document, the rankings would compress a lot.

Budget tier isn't just cheap, it's usually enough here

Summarization doesn't typically require the deep reasoning or agentic tool-use flagship models are built for — it needs reliable extraction and compression of information already in the prompt, which budget models tend to handle fine. The jump from the cheapest model ($60/mo) to a mid-tier like Claude Sonnet 5 or GPT-6.1 Sol ($1,200/mo) is 20x, and the jump to flagship models like Claude Opus or GPT-6 Astra ($3,000–$6,000/mo) is 50–100x. Before paying that premium, it's worth testing whether a budget or mid-tier model actually produces worse summaries for your specific documents, or just feels riskier on paper. Context window matters too: at 500M input tokens a month, make sure whatever model you pick has enough context to handle your longest single documents without chunking — most options here offer 1M context, but GPT-4o, DeepSeek V3, and Claude Haiku 4.5 top out at 128K–200K.

Batching and caching change the math further

If your summarization jobs aren't real-time — overnight batch runs on transcripts, scheduled report digests, and so on — many providers offer batch processing that's often around half the normal price, which can push an already-cheap model even lower. Prompt caching is also worth checking if you're repeatedly summarizing against the same reference material, system prompt, or style guide, since cached input tokens are typically billed at a steep discount versus fresh input. Given that input volume is what drives cost on this workload, caching and batching are usually the two highest-leverage optimizations available, often more impactful than switching models.

How we'd actually decide

Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.

Frequently asked questions

What's the absolute cheapest model for this workload?

GPT-6 Luna from OpenAI, at $0.1 per 1M input tokens and $0.5 per 1M output tokens, comes to $60/mo for 500M input and 20M output tokens.

Is the cheapest model always the right choice for summarization?

Not automatically. Budget models are usually well-suited to summarization since it's an extraction and compression task rather than complex reasoning, but you should verify output quality on your actual documents. If you need more capability, mid-tier options like Gemini 3 Flash ($310/mo) or DeepSeek V4 Pro ($370/mo) cost more but are still far cheaper than flagship models.

Does context window size matter for choosing a model here?

Yes. Most of the cheapest options — GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, GPT-4o mini's sibling DeepSeek V4.1 Flash — offer up to 1M context, which handles long reports or transcripts in a single pass. Some budget options like GPT-4o mini (128K) and Claude Haiku 4.5 (200K) have smaller context windows that may require splitting very long documents.

Prices on this page change. Get told when they do.

One email the day a tracked model changes price or a new one launches. Nothing else, unsubscribe anytime.