USE CASE · VERIFIED 2 OCT 2026
What's the cheapest AI API for translation and localization?
On a 50M input + 50M output token workload, GPT-6 Luna from OpenAI is the cheapest at $30/mo ($0.1/$0.5 per 1M tokens), with Qwen3.8 Flash at $31/mo and GLM 5.3 Flash at $32/mo right behind it. All three are budget-tier models with 1M context windows, which suits translation's roughly even input/output split.
Worked example: 50M input, 50M output tokens/month
Text in, same-length text out: input and output roughly equal. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.
Why translation workloads favor budget models
Translation and localization are a near-ideal case for cheap models: you're not asking the model to reason deeply, you're asking it to map text in one language to equivalent text in another, roughly token-for-token. Since output volume tracks input volume closely, output price matters just as much as input price here — unlike chat or coding workloads where output is usually much shorter. That's why models with cheap output pricing, like GPT-6 Luna at $0.5/1M out or DeepSeek V3 at $0.41/1M out, do well on this specific workload even though they'd rank differently on an output-heavy task.
Context window and batching considerations
Most of the cheapest options here — GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, DeepSeek V4.1 Flash — offer a 1M token context window, which is plenty for translating long documents or batches of strings in a single call without chunking. DeepSeek V3 and GPT-4o mini cap out at 128K context, which is still workable for most localization files but may force you to split very large documents. If your pipeline already batches requests or can tolerate async processing, check whether your provider offers batch API pricing, which is often around half the standard rate — that can push an already-cheap model even lower for high-volume localization jobs.
When to pay more than the cheapest option
The cheapest model is rarely the wrong default for bulk translation, but there are real reasons to move up a tier. If you're translating into low-resource languages, handling heavy idiom or cultural nuance, or need consistent terminology across a huge glossary, a mid-tier model like Gemini 3 Flash ($175/mo) or Claude Haiku 4.5 ($300/mo) may justify its cost through fewer post-edit fixes. Flagship models like Claude Opus 5.5 ($1,200/mo) or GPT-5.5 ($1,750/mo) are usually overkill for straight translation — that budget is better spent on human QA passes than on a bigger model, unless your content is unusually high-stakes (legal, medical, regulatory).
How we'd actually decide
- High-volume bulk translation, cost is the main constraint: GPT-6 Luna — cheapest at $30/mo for this workload, 1M context covers long documents
- Need a close runner-up with slightly different pricing shape: Qwen3.8 Flash or GLM 5.3 Flash — both near $31-32/mo with 1M context
- Working with documents over 128K tokens but want to stay ultra-cheap: DeepSeek V4.1 Flash — $38/mo with a 1M context window
- Higher-stakes content needing more careful handling: Gemini 3 Flash — $175/mo mid-tier option, still well below flagship pricing
Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.