USE CASE · VERIFIED 2 OCT 2026

What's the cheapest AI API for high-volume text classification?

For a workload of 500M input tokens and 10M output tokens a month, GPT-6 Luna from OpenAI is cheapest at $0.1/$0.5 per 1M tokens, costing $55/mo. Qwen3.8 Flash and GLM 5.3 Flash are close behind at $80/mo, with GPT-4o mini and DeepSeek V4.1 Flash tied at $81/mo.

Worked example: 500M input, 10M output tokens/month

Labelling millions of short texts: almost all input, tiny outputs. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.

GPT-6 Luna$55/mo
Qwen3.8 Flash$80/mo
GLM 5.3 Flash$80/mo
GPT-4o mini$81/mo
DeepSeek V4.1 Flash$81/mo
DeepSeek V3$139/mo
Grok 4.7$1,060/mo
Gemini 3.1 Pro$1,120/mo

Why input price dominates this workload

Classification jobs like labelling, routing, or tagging short texts are almost all input tokens with a few words of output — a tag, a score, a category. That means the input price per 1M tokens matters far more than the output price when ranking total cost. This is why GPT-6 Luna wins despite having a higher output price ratio than some rivals: at $0.1 per 1M input tokens, it's the cheapest way to push hundreds of millions of tokens through the model, and the tiny output volume barely moves the bill. If your ratio shifts toward longer outputs — summaries instead of labels — the ranking can change, so always check the per-model workload cost rather than just the sticker input price.

Budget tier is built for this job, but check the floor

Every model in the top ten for this workload is a budget-tier model, and that's not a coincidence — classification doesn't need flagship reasoning, it needs volume at low cost. GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, GPT-4o mini, and DeepSeek V4.1 Flash all land between $55 and $81/mo, and all but GPT-4o mini offer a 1M token context window, which is overkill for short-text classification but useful if you ever batch many samples into one call. The jump to mid-tier models like Gemini 3 Flash ($280/mo) or Claude Haiku 4.5 ($550/mo) is a real cost increase — worth it only if the budget models are measurably getting classifications wrong for your specific categories.

Batching and caching change the real number

The prices here are list prices for pay-per-token API calls. Many providers offer a batch processing mode for non-realtime jobs — classification is a textbook fit since you don't need the answer in 200ms — and batch discounts are often around half the standard rate. Prompt caching is the other lever: if every classification call reuses the same instructions or category list as a prefix, a provider with strong prompt caching can cut effective input cost substantially even before batching. Check both options with your shortlisted provider before locking in a model purely on list price.

How we'd actually decide

Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.

Frequently asked questions

What's the absolute cheapest model for classifying millions of short texts?

GPT-6 Luna from OpenAI, at $0.1 per 1M input tokens and $0.5 per 1M output tokens. For a workload of 500M input and 10M output tokens a month, that comes to $55/mo, the lowest of all 30 tracked models.

Is the cheapest model always the right choice for classification?

Not automatically. GPT-6 Luna and the other budget models (Qwen3.8 Flash, GLM 5.3 Flash, GPT-4o mini, DeepSeek V4.1 Flash) are all priced for high-volume, low-complexity work, but if your classification task needs more nuanced judgment, a mid-tier model like Gemini 3 Flash ($280/mo) or Claude Haiku 4.5 ($550/mo) may be worth the extra cost. Price data alone can't tell you accuracy — test on your own categories first.

Does context window size matter for classification workloads?

Usually not much, since each classification call only needs the text being labelled plus short instructions. Most budget models here, including GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, and DeepSeek V4.1 Flash, offer a 1M token context window anyway, which matters more if you batch many texts into a single prompt than for one-off short classifications.

Prices on this page change. Get told when they do.

One email the day a tracked model changes price or a new one launches. Nothing else, unsubscribe anytime.