USE CASE · VERIFIED 2 OCT 2026

What's the cheapest AI API for translation and localization?

On a 50M input + 50M output token workload, GPT-6 Luna from OpenAI is the cheapest at $30/mo ($0.1/$0.5 per 1M tokens), with Qwen3.8 Flash at $31/mo and GLM 5.3 Flash at $32/mo right behind it. All three are budget-tier models with 1M context windows, which suits translation's roughly even input/output split.

Worked example: 50M input, 50M output tokens/month

Text in, same-length text out: input and output roughly equal. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.

GPT-6 Luna$30/mo
Qwen3.8 Flash$31/mo
GLM 5.3 Flash$32/mo
DeepSeek V3$34/mo
GPT-4o mini$38/mo
DeepSeek V4.1 Flash$38/mo
Grok 4.7$400/mo
Gemini 3.1 Pro$700/mo

Why translation workloads favor budget models

Translation and localization are a near-ideal case for cheap models: you're not asking the model to reason deeply, you're asking it to map text in one language to equivalent text in another, roughly token-for-token. Since output volume tracks input volume closely, output price matters just as much as input price here — unlike chat or coding workloads where output is usually much shorter. That's why models with cheap output pricing, like GPT-6 Luna at $0.5/1M out or DeepSeek V3 at $0.41/1M out, do well on this specific workload even though they'd rank differently on an output-heavy task.

Context window and batching considerations

Most of the cheapest options here — GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, DeepSeek V4.1 Flash — offer a 1M token context window, which is plenty for translating long documents or batches of strings in a single call without chunking. DeepSeek V3 and GPT-4o mini cap out at 128K context, which is still workable for most localization files but may force you to split very large documents. If your pipeline already batches requests or can tolerate async processing, check whether your provider offers batch API pricing, which is often around half the standard rate — that can push an already-cheap model even lower for high-volume localization jobs.

When to pay more than the cheapest option

The cheapest model is rarely the wrong default for bulk translation, but there are real reasons to move up a tier. If you're translating into low-resource languages, handling heavy idiom or cultural nuance, or need consistent terminology across a huge glossary, a mid-tier model like Gemini 3 Flash ($175/mo) or Claude Haiku 4.5 ($300/mo) may justify its cost through fewer post-edit fixes. Flagship models like Claude Opus 5.5 ($1,200/mo) or GPT-5.5 ($1,750/mo) are usually overkill for straight translation — that budget is better spent on human QA passes than on a bigger model, unless your content is unusually high-stakes (legal, medical, regulatory).

How we'd actually decide

Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.

Frequently asked questions

What's the single cheapest API for translation at this volume?

GPT-6 Luna from OpenAI, at $0.1/$0.5 per 1M input/output tokens, which works out to $30/mo for 50M input + 50M output tokens.

Does a bigger context window matter for translation?

It matters if you're translating long documents in one pass rather than splitting them into chunks. Most of the cheapest models here, including GPT-6 Luna, Qwen3.8 Flash, and GLM 5.3 Flash, offer a 1M token context window, while DeepSeek V3 and GPT-4o mini are capped at 128K.

Is the cheapest model always the right choice for localization?

Not automatically. Budget models cover basic translation well, but if you need more careful handling of nuance or terminology consistency, mid-tier options like Gemini 3 Flash ($175/mo) or Claude Haiku 4.5 ($300/mo) are still far cheaper than flagship models like GPT-5.5 ($1,750/mo) or Claude Opus 5.5 ($1,200/mo), which are generally unnecessary for straight translation work.

Prices on this page change. Get told when they do.

One email the day a tracked model changes price or a new one launches. Nothing else, unsubscribe anytime.