USE CASE · VERIFIED 2 OCT 2026
What's the cheapest AI API for high-volume text classification?
For a workload of 500M input tokens and 10M output tokens a month, GPT-6 Luna from OpenAI is cheapest at $0.1/$0.5 per 1M tokens, costing $55/mo. Qwen3.8 Flash and GLM 5.3 Flash are close behind at $80/mo, with GPT-4o mini and DeepSeek V4.1 Flash tied at $81/mo.
Worked example: 500M input, 10M output tokens/month
Labelling millions of short texts: almost all input, tiny outputs. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.
Why input price dominates this workload
Classification jobs like labelling, routing, or tagging short texts are almost all input tokens with a few words of output — a tag, a score, a category. That means the input price per 1M tokens matters far more than the output price when ranking total cost. This is why GPT-6 Luna wins despite having a higher output price ratio than some rivals: at $0.1 per 1M input tokens, it's the cheapest way to push hundreds of millions of tokens through the model, and the tiny output volume barely moves the bill. If your ratio shifts toward longer outputs — summaries instead of labels — the ranking can change, so always check the per-model workload cost rather than just the sticker input price.
Budget tier is built for this job, but check the floor
Every model in the top ten for this workload is a budget-tier model, and that's not a coincidence — classification doesn't need flagship reasoning, it needs volume at low cost. GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, GPT-4o mini, and DeepSeek V4.1 Flash all land between $55 and $81/mo, and all but GPT-4o mini offer a 1M token context window, which is overkill for short-text classification but useful if you ever batch many samples into one call. The jump to mid-tier models like Gemini 3 Flash ($280/mo) or Claude Haiku 4.5 ($550/mo) is a real cost increase — worth it only if the budget models are measurably getting classifications wrong for your specific categories.
Batching and caching change the real number
The prices here are list prices for pay-per-token API calls. Many providers offer a batch processing mode for non-realtime jobs — classification is a textbook fit since you don't need the answer in 200ms — and batch discounts are often around half the standard rate. Prompt caching is the other lever: if every classification call reuses the same instructions or category list as a prefix, a provider with strong prompt caching can cut effective input cost substantially even before batching. Check both options with your shortlisted provider before locking in a model purely on list price.
How we'd actually decide
- Situation: pure lowest cost, no special requirements — GPT-6 Luna — cheapest at $55/mo for this workload
- Situation: want to avoid single-vendor lock-in to OpenAI — Qwen3.8 Flash or GLM 5.3 Flash — both $80/mo, near-identical to the cheapest option
- Situation: already building on OpenAI and want a familiar, well-documented budget model — GPT-4o mini — $81/mo, smaller 128K context but simple to adopt
- Situation: budget-tier accuracy isn't cutting it for your categories — Gemini 3 Flash or Claude Haiku 4.5 — $280/mo and $550/mo respectively, a real step up in tier
Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.