Estimate what your OpenAI, Claude or Gemini API usage really costs — per request, per month and per year. Compare model pricing side by side and find the cheapest model for your workload.
Monthly cost for your token & request numbers above, cheapest first. The green row is your best-value option.
| Model | Input $/1M | Output $/1M | Monthly cost |
|---|
Estimates only. Token counts are approximate (~4 characters per token); actual tokenization varies by model. Rates verified September 2026 — always confirm current pricing on each provider's official page before budgeting.
Turn confusing per-token pricing into a straight monthly figure in four steps — then see which model is cheapest for your workload.
Choose OpenAI, Claude or Gemini — the current input and output rates fill in automatically (and stay editable).
Type your input and output tokens per request, or paste a sample prompt and reply to estimate them live.
Get per-request, daily, monthly and yearly figures instantly, with an input-vs-output split so you know what's driving the bill.
The table ranks every model cheapest-first for your workload — and caching and batch options show your real savings.
An LLM API cost calculator turns the confusing per-million-token pricing of OpenAI, Anthropic and Google into a straight answer: what will this actually cost me per request, per month and per year? Enter your model, your typical input and output tokens, and your monthly request volume, and it works out your real bill — then compares every model so you can see which one is cheapest for your workload.
It's the fastest way to sanity-check an AI feature's running cost before you ship it — and to catch the model choice that quietly doubles your invoice.
Every major provider bills the same way: per token, with a lower rate for input tokens (your prompt, system message and context) and a higher rate for output tokens (the model's reply). Rates are published per million tokens — so a model at "$2 / $10" charges $2 per million input and $10 per million output.
Two things surprise most teams. First, output dominates: because output runs several times the input rate, long responses cost far more than long prompts. Second, tokens aren't words — a token is about four characters, so a 750-word answer is roughly 1,000 tokens. The calculator estimates tokens from any text you paste so you don't have to guess.
If the monthly figure made you wince, here's where the savings actually are — in order of impact.
The biggest lever by far. You don't need a flagship model for classification, tagging or short summaries — a cheaper tier does the job at a fraction of the cost. Reserve the expensive models for work that genuinely needs deep reasoning. The comparison table above shows the gap for your own numbers.
If every request reuses the same system prompt or document, prompt caching bills those repeated tokens at roughly 10% of the input rate. For a support bot with a big fixed prompt, that alone can halve the bill.
Asynchronous jobs — bulk classification, content generation, data enrichment — can run through the Batch API at 50% off both input and output. Toggle it in the advanced options to see the effect.
Shorter system prompts and a sensible max-output limit stop you paying for tokens you never needed. Since output is the expensive side, capping response length is one of the fastest wins.
Want it handled in production? Devlet's AI integration team builds AI features with model routing, caching and cost controls baked in — and works with custom AI models too.
The vocabulary behind your API bill, without the jargon.
$2 / $10 charges $2 per million input tokens, $10 per million output.Everything you need to know about estimating and cutting your OpenAI, Claude and Gemini API costs.
Sanity-check the running cost of an AI feature before you ship, and pick the model that hits your budget.
Model per-request economics at scale and forecast spend as usage grows, before it hits the invoice.
Quote AI builds accurately and show clients exactly what running the model will cost each month.
The full guide to model routing, caching and batching — with the pricing traps to avoid.
How Devlet builds AI features that stay fast and affordable at scale.
See whether your pages get cited in Google's AI Overviews and AI answers.
H&U Solutions ↗Create an llms.txt file so AI models understand and reference your site correctly.
H&U Solutions ↗Building an AI product? Estimate costs here, then make sure it's visible in AI search with the AI Overview Checker by H&U Solutions ↗
Every tool is free — no signup. Estimate your AI costs here, then measure and grow with the rest of the suite.
Build clean, GA4-ready campaign tracking links with source, medium and campaign.
Open tool →Work out ROAS, break-even ROAS, CPM, CPC, CPA and your ad budget.
Open tool →Generate valid FAQ, Article & LocalBusiness JSON-LD for rich results and AI citations.
Open tool →Get a free scope & cost estimate from the Devlet AI integration team.
Talk to us →Build clean, GA4-ready campaign tracking links with source, medium and campaign.
Open tool →Work out ROAS, break-even ROAS, CPM, CPC, CPA and your ad budget.
Open tool →Generate valid FAQ, Article & LocalBusiness JSON-LD for rich results and AI citations.
Open tool →Get a free scope & cost estimate from the Devlet AI integration team.
Talk to us →We design AI integrations that stay fast and affordable at scale — the right model, caching and cost controls built in from day one.
Talk to Devlet's AI team →Enter your details and we’ll send you a complete website audit report.