USE CASE · VERIFIED 2 OCT 2026

What's the cheapest AI API for a customer support chatbot?

For a 100M input + 50M output token monthly workload, GPT-6 Luna is cheapest at $35/mo ($0.1/$0.5 per 1M tokens), just ahead of Qwen3.8 Flash at $38/mo and GLM 5.3 Flash at $40/mo. A whole cluster of budget models — GPT-4o mini, DeepSeek V4.1 Flash, DeepSeek V3 — land between $45-48/mo, so price alone won't separate them.

Worked example: 100M input, 50M output tokens/month

A support bot with a knowledge-base prompt and conversational replies. The six cheapest of the 30 models we track, plus two flagships for scale. Open-weight models are priced at the original provider's own list price.

GPT-6 Luna$35/mo
Qwen3.8 Flash$38/mo
GLM 5.3 Flash$40/mo
GPT-4o mini$45/mo
DeepSeek V4.1 Flash$45/mo
DeepSeek V3$48/mo
Grok 4.7$500/mo
Gemini 3.1 Pro$800/mo

Why support bots are a budget-tier workload

Support chatbots are a textbook case for budget models: the prompt is usually dominated by input tokens (knowledge base context, chat history, system instructions) with shorter conversational replies, which is exactly the 100M in / 50M out ratio used here. Because input tokens are cheap across almost every model on this list, the gap between the cheapest and most expensive options is driven mostly by output pricing and by how much you're paying for a 'flagship' label you may not need. Before picking a model, check whether your real traffic matches this 2:1 input-heavy ratio — if your bot does more generation (long explanations, drafted emails) the ranking can shift toward models with cheaper output rates.

Budget vs mid vs flagship: what you're actually paying for

The budget tier (GPT-6 Luna through Gemini 2.5 Flash, roughly $35-155/mo for this workload) is built for high-volume, repetitive tasks like FAQ answering and ticket triage, and most of these models now ship with 1M-token context windows, so you're not sacrificing knowledge-base size to save money. The mid tier ($165-700/mo, DeepSeek V4 Pro through GPT-4o) and flagship tier ($700-3,500/mo, Gemini 3.1 Pro through GPT-6 Astra) cost 5-100x more for this same workload — that premium only makes sense if your bot needs stronger reasoning for complex multi-step troubleshooting, tool use, or escalation logic that budget models handle poorly. Don't assume you need mid or flagship by default; test a budget model against your actual support transcripts first.

Context window, caching, and batching matter more than the sticker price

If your knowledge base is large, check context window before price: most budget models here support 1M tokens, but GPT-4o mini and DeepSeek V3 top out at 128K, which can force you into chunking or retrieval workarounds that add engineering cost. Prompt caching is also critical for support bots since the same system prompt and knowledge-base context get reused across thousands of conversations — providers that support caching can cut effective input costs substantially, often around half, so check current caching support and discounts for whichever model you shortlist rather than relying solely on the list price. Batching requests (for offline tasks like summarizing tickets overnight) can bring further savings on providers that offer a batch API, which is worth checking separately from the per-token prices listed here.

How we'd actually decide

Worked example uses standard (non-batch, non-cached) list pricing verified 2 October 2026. Prices change; use the calculator for today's numbers with your own volume.

Frequently asked questions

What is the cheapest model for a customer support chatbot?

For this workload (100M input + 50M output tokens/month), GPT-6 Luna from OpenAI is cheapest at $35/mo, priced at $0.1/$0.5 per 1M input/output tokens.

Is the cheapest model good enough for a support bot?

It depends on your use case. The data here only covers price, not quality — budget models like GPT-6 Luna, Qwen3.8 Flash, and GLM 5.3 Flash are built for high-volume, straightforward tasks, while mid-tier ($165-700/mo) and flagship ($700-3,500/mo) models cost significantly more and are intended for harder reasoning tasks. Test against your own transcripts before committing.

Does context window size matter for choosing a model?

Yes, especially if your knowledge base is large. Most budget models in this list, including GPT-6 Luna, Qwen3.8 Flash, GLM 5.3 Flash, and DeepSeek V4.1 Flash, offer a 1M-token context window, while GPT-4o mini and DeepSeek V3 are limited to 128K, which may require more chunking or retrieval engineering.

Prices on this page change. Get told when they do.

One email the day a tracked model changes price or a new one launches. Nothing else, unsubscribe anytime.