USE CASE · VERIFIED 2 OCT 2026
What's the cheapest AI API for summarizing long documents?
For a workload of 500M input and 20M output tokens a month, GPT-6 Luna from OpenAI is the cheapest option at $60/mo. Close behind are Qwen3.8 Flash at $84/mo and GLM 5.3 Flash at $85/mo — both also budget-tier models with 1M context windows built for exactly this input-heavy, output-light pattern.
Worked example: 500M input, 20M output tokens/month
Long reports and transcripts in, compact summaries out. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.
Why this workload favors budget models so heavily
Summarization is almost entirely an input problem: you're feeding in 500M tokens of reports or transcripts and getting back only 20M tokens of compact summary, a 25:1 ratio. That means input price dominates your bill far more than output price. Models like GPT-6 Luna ($0.1/$0.5), Qwen3.8 Flash ($0.15/$0.47), and GLM 5.3 Flash ($0.15/$0.5) win precisely because they keep input costs near rock bottom, while expensive output pricing on flagship models barely matters when output volume is this small relative to input. If your ratio shifted toward more output per document, the rankings would compress a lot.
Budget tier isn't just cheap, it's usually enough here
Summarization doesn't typically require the deep reasoning or agentic tool-use flagship models are built for — it needs reliable extraction and compression of information already in the prompt, which budget models tend to handle fine. The jump from the cheapest model ($60/mo) to a mid-tier like Claude Sonnet 5 or GPT-6.1 Sol ($1,200/mo) is 20x, and the jump to flagship models like Claude Opus or GPT-6 Astra ($3,000–$6,000/mo) is 50–100x. Before paying that premium, it's worth testing whether a budget or mid-tier model actually produces worse summaries for your specific documents, or just feels riskier on paper. Context window matters too: at 500M input tokens a month, make sure whatever model you pick has enough context to handle your longest single documents without chunking — most options here offer 1M context, but GPT-4o, DeepSeek V3, and Claude Haiku 4.5 top out at 128K–200K.
Batching and caching change the math further
If your summarization jobs aren't real-time — overnight batch runs on transcripts, scheduled report digests, and so on — many providers offer batch processing that's often around half the normal price, which can push an already-cheap model even lower. Prompt caching is also worth checking if you're repeatedly summarizing against the same reference material, system prompt, or style guide, since cached input tokens are typically billed at a steep discount versus fresh input. Given that input volume is what drives cost on this workload, caching and batching are usually the two highest-leverage optimizations available, often more impactful than switching models.
How we'd actually decide
- Situation: High-volume batch summarization, cost is the main constraint — GPT-6 Luna — cheapest at $60/mo for this workload, 1M context handles long documents directly
- Situation: Want a safety margin on quality without leaving budget tier — Qwen3.8 Flash or GLM 5.3 Flash — $84–85/mo, both still 1M context, close runners-up worth A/B testing against Luna
- Situation: Summarizing within a broader agentic or multi-step pipeline — Gemini 3 Flash or DeepSeek V4 Pro — mid-tier pricing ($310–370/mo) if you need more headroom than budget models offer elsewhere in the pipeline
- Situation: Documents routinely exceed 200K tokens and you're considering Claude Haiku 4.5 — check context window first — Haiku 4.5 caps at 200K context, so very long reports may need chunking that budget models with 1M context avoid
Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.