USE CASE · VERIFIED 2 OCT 2026
What's the cheapest AI API for a customer support chatbot?
For a 100M input + 50M output token monthly workload, GPT-6 Luna is cheapest at $35/mo ($0.1/$0.5 per 1M tokens), just ahead of Qwen3.8 Flash at $38/mo and GLM 5.3 Flash at $40/mo. A whole cluster of budget models — GPT-4o mini, DeepSeek V4.1 Flash, DeepSeek V3 — land between $45-48/mo, so price alone won't separate them.
Worked example: 100M input, 50M output tokens/month
A support bot with a knowledge-base prompt and conversational replies. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.
Why support bots are a budget-tier workload
Support chatbots are a textbook case for budget models: the prompt is usually dominated by input tokens (knowledge base context, chat history, system instructions) with shorter conversational replies, which is exactly the 100M in / 50M out ratio used here. Because input tokens are cheap across almost every model on this list, the gap between the cheapest and most expensive options is driven mostly by output pricing and by how much you're paying for a 'flagship' label you may not need. Before picking a model, check whether your real traffic matches this 2:1 input-heavy ratio — if your bot does more generation (long explanations, drafted emails) the ranking can shift toward models with cheaper output rates.
Budget vs mid vs flagship: what you're actually paying for
The budget tier (GPT-6 Luna through Gemini 2.5 Flash, roughly $35-155/mo for this workload) is built for high-volume, repetitive tasks like FAQ answering and ticket triage, and most of these models now ship with 1M-token context windows, so you're not sacrificing knowledge-base size to save money. The mid tier ($165-700/mo, DeepSeek V4 Pro through GPT-4o) and flagship tier ($700-3,500/mo, Gemini 3.1 Pro through GPT-6 Astra) cost 5-100x more for this same workload — that premium only makes sense if your bot needs stronger reasoning for complex multi-step troubleshooting, tool use, or escalation logic that budget models handle poorly. Don't assume you need mid or flagship by default; test a budget model against your actual support transcripts first.
Context window, caching, and batching matter more than the sticker price
If your knowledge base is large, check context window before price: most budget models here support 1M tokens, but GPT-4o mini and DeepSeek V3 top out at 128K, which can force you into chunking or retrieval workarounds that add engineering cost. Prompt caching is also critical for support bots since the same system prompt and knowledge-base context get reused across thousands of conversations — providers that support caching can cut effective input costs substantially, often around half, so check current caching support and discounts for whichever model you shortlist rather than relying solely on the list price. Batching requests (for offline tasks like summarizing tickets overnight) can bring further savings on providers that offer a batch API, which is worth checking separately from the per-token prices listed here.
How we'd actually decide
- Situation: You're prototyping or have low volume — GPT-6 Luna — cheapest at $35/mo, low risk to test against your transcripts
- Situation: You want a safety margin on vendor/model risk without much cost increase — Qwen3.8 Flash or GLM 5.3 Flash — both $38-40/mo with 1M context, close seconds to the cheapest
- Situation: Your knowledge base is large and you're stuck with a 128K-context model — DeepSeek V4.1 Flash or GPT-6 Luna — same budget price tier but 1M context avoids chunking
- Situation: Your bot needs stronger reasoning for complex escalations — DeepSeek V4 Pro — cheapest mid-tier option at $165/mo before jumping to $350+ models like Claude Haiku 4.5
Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.