Built for builders

LLM API Pricing Calculator

Compare token costs across every major LLM provider — OpenAI, Anthropic, Google, Mistral, Meta and more. Estimate your monthly spend in seconds.

  • 25+ models
  • Live calculator
  • Updated quarterly
  • No sign-up

Estimate your costs

Pick a model, dial in your traffic, and see the monthly bill update live.

Configure usage

Selected: GPT-4o mini · Context 128k tokens · Fast

800

Average prompt + system + retrieved context per call.

400

Average response length you ask the model to produce.

500

≈ 15,000 requests/month

Estimated monthly cost

$5.40

Based on 15,000 requests · 18.00M tokens

Input 33%Output 67%

Input tokens

12,000,000 tok @ $0.15/1M

$1.80

Output tokens

6,000,000 tok @ $0.6/1M

$3.60

Monthly total

$5.40

Annual projection

$64.80

Compare

All models side-by-side

Click any column to sort. Filter by provider, tier, or keyword. Prices are quoted per 1M tokens.

Notes
Qwen 3.7 Flash
Alibaba$0.03$0.131,000kFastCheapest 1M-contextVery fastThe cheapest 1M-context frontier-adjacent model on the market.
GPT-OSS 120B
Estimate
OpenAI$0.04$0.17131.072kFastSelf-hosted OSSVery fastOpenAI’s 120B open-weights model; extremely cheap when hosted on Together / Groq.
GPT-5 nano
OpenAI$0.05$0.40400kFastBulk classificationVery fastCheapest 5-series tier; strong replacement for GPT-4.1 nano.
Llama 3.2 3B
Estimate
Meta$0.05$0.33131.072kFastEdge / on-deviceVery fastEdge-friendly open-weights model; well below 5¢/M served on Groq.
Amazon Nova Lite
Amazon$0.06$0.24300kFastAWS cheap multimodalVery fastCheapest Nova tier; multi-modal at 6¢/M input.
DeepSeek V4 Flash
DeepSeek$0.09$0.181,048.576kFastCheap fast tierVery fast1M-context DeepSeek at sub-30¢/M input; production-grade fast tier.
GPT-5.6 Luna
OpenAI$0.10$0.601,050kFastHigh-volume tasksVery fastCheapest GPT-5.6 tier; strong replacement for GPT-4o mini.
Llama 3.3 70BEstimateMeta$0.10$0.32131.072kBalancedSelf-hosted balancedFastOpen weights; median price across Together, Groq and Fireworks.
GPT-4o miniOpenAI$0.15$0.60128kFastLegacy fast chatVery fastLong-serving low-cost default; superseded by GPT-5.6 Luna and GPT-5 nano.
Mistral Small 2603
Mistral$0.15$0.60262.144kFastCheap fast MistralVery fastExcellent latency/cost ratio; 262k context at 15¢/M input.
Command R
Cohere$0.15$0.60128kBalancedCheap RAGFastRetrieval-optimised Cohere at GPT-4o mini prices.
GPT-5.4 nano
OpenAI$0.20$1.25400kFastClassifiersVery fastCheapest GPT-5.4 tier; 400k context at 20¢/M input.
GPT-5 mini
OpenAI$0.25$2.00400kFastCheap chatVery fastPopular low-cost 5-series tier; 400k context at 25¢/M input.
Claude 3 Haiku
Anthropic$0.25$1.25200kFastLegacy budget ClaudeVery fastLegacy budget Haiku still on the API; useful for cost baselines vs. Haiku 4.5.
Gemini 3.1 Flash-Lite
Google$0.25$1.501,048.576kFastHigh-throughput RAGVery fastCheapest 1M-context Gemini; ideal for high-throughput RAG.
DeepSeek V3.2
DeepSeek$0.27$0.40163.84kBalancedOSS defaultFastPopular 2026 open-weights model; the reference free/paid split.
Gemini 2.5 Flash
Google$0.30$2.501,048.576kFastLegacy fast tierVery fastWidely deployed low-cost Gemini; being replaced by 3.5 Flash.
Qwen 3.7 Plus
Alibaba$0.32$1.281,000kBalancedOSS multilingualFastBalanced Qwen tier; strong multilingual + tool-use performance.
GPT-4.1 mini
OpenAI$0.40$1.601,047.576kFastCheap long-contextVery fast1M-context 4.1 mini at 40¢/M input; still competitive on RAG workloads.
DeepSeek V4 Pro
DeepSeek$0.43$0.871,048.576kFrontierCheap frontierMediumFrontier-class output at fast-tier prices; open weights.
Kimi K2 Thinking
Moonshot$0.60$2.50262.144kReasoningMulti-step toolsSlowReasoning-specialised Kimi; strong at multi-step tool use.
GPT-5.4 mini
OpenAI$0.75$4.50400kBalancedLow-latency chatFastBest cost/quality in the 5.4 line; sits between GPT-5.6 Luna and Terra.
Amazon Nova Pro
Amazon$0.80$3.20300kBalancedAWS balancedFastBalanced Bedrock model with 300k context and multi-modal input.
GPT-5.6 Terra
OpenAI$1.00$6.001,050kFrontierEveryday productionMediumBalanced GPT-5.6 tier — flagship quality without Sol pricing.
Claude Haiku 4.5
Anthropic$1.00$5.00200kFastFast Claude tierVery fastFastest Claude tier; 200k context at a fraction of Sonnet 5 pricing.
OpenAI o4-mini
OpenAI$1.10$4.40200kReasoningAgent defaultSlowSuccessor to o3-mini; same pricing with stronger multi-step tool use.
OpenAI o3-miniOpenAI$1.10$4.40200kReasoningCheap reasoningSlowCheapest reasoning-tier model with useful math and coding chops.
GPT-5.1
OpenAI$1.25$10.00400kFrontierGeneral chatMedium2025 flagship still on the API; same list price as GPT-5 at higher headroom.
GPT-5.1 Codex
OpenAI$1.25$10.00400kBalancedCoding tasksFastCoding-specialised sibling of GPT-5.1; recommended for agentic code workflows.
GPT-5
OpenAI$1.25$10.00400kFrontierCheap flagshipMediumWidely deployed 2025 flagship at 400k context.
Gemini 2.5 ProGoogle$1.25$10.001,048.576kFrontierLegacy flagshipMediumPopular 2025 flagship still shipped as a stable/priced option.
Grok 4.3
xAI$1.25$2.501,000kBalancedGeneral-purpose 1MFast1M-context Grok at Sonnet-5 pricing; strong general-purpose choice.
Qwen 3.7 Max
Alibaba$1.48$4.421,000kFrontierOSS frontierMedium1M-context frontier Qwen; open-weights alternative to Sonnet 5.
Gemini 3.5 Flash
Google$1.50$9.001,048.576kBalanced1M-context workhorseFast1M-token workhorse; drop-in for most Gemini 2.5 Flash traffic.
Mistral Medium 3.5
Mistral$1.50$7.50262.144kBalancedEU-hosted balancedFastEU-hosted balanced tier; strong multilingual + code performance.
OpenAI o3
OpenAI$2.00$8.00200kReasoningAgent loopsSlowo-series reasoning at ~1/7 the cost of o1-pro.
GPT-4.1
OpenAI$2.00$8.001,047.576kBalancedLong-context RAGFast1M-context legacy flagship; still active for teams pinned to 4-series behaviour.
Claude Sonnet 5
Anthropic$2.00$10.001,000kFrontierProduction defaultMediumBest cost/quality Claude — replaces Sonnet 4.5/4.6 for most production workloads.
Gemini 3.1 Pro
Google$2.00$12.001,048.576kFrontierFrontier RAGMedium1M-token frontier tier — the strongest Gemini available on the API.
Grok 4.5
xAI$2.00$6.00500kFrontierReal-time-awareMediumReal-time-data-aware frontier; competitive on reasoning benchmarks.
GPT-5.4
OpenAI$2.50$15.001,050kFrontierProduction defaultMediumRecommended replacement for GPT-4o; 1M context at $2.50/M input.
GPT-4oOpenAI$2.50$10.00128kFrontierLegacy multimodalMediumLegacy 2024 flagship — the model to migrate off (recommended: GPT-5.4).
Amazon Nova Premier
Amazon$2.50$12.501,000kFrontierAWS flagshipMediumBedrock-native flagship; useful when your stack is already on AWS.
Command R+
Cohere$2.50$10.00128kFrontierEnterprise RAGMediumBuilt for RAG; native citation support and strong multilingual coverage.
Claude Sonnet 4.6
Anthropic$3.00$15.001,000kFrontierLong-context chatMediumLast stable Sonnet 4.x; 1M context at $3/M input.
Claude Sonnet 4.5
Anthropic$3.00$15.001,000kFrontierLegacy 2025 flagshipMediumPopular 2025 flagship still live on the API for teams pinned to a specific behaviour.
Kimi K3
Moonshot$3.00$15.001,048.576kFrontierChinese frontierMedium1M-context reasoning-heavy Chinese frontier model.
GPT-5.6 Sol Pro
OpenAI$5.00$30.001,050kReasoningTop-tier reasoningSlowTop OpenAI reasoning tier; use for hard math, code, and research.
GPT-5.6 Sol
OpenAI$5.00$30.001,050kReasoningHard reasoningSlowGPT-5.6 reasoning tier at the standard rate.
GPT-5.5OpenAI$5.00$30.001,050kFrontierGeneral flagshipMedium1M-context flagship prior to GPT-5.6; still the widely deployed standard.
Claude Opus 5
Anthropic$5.00$25.001,000kFrontierAgentic tool useMediumTop production Opus tier; best for agentic reasoning and long-horizon tool use.
Claude Opus 4.8
Anthropic$5.00$25.001,000kFrontierStable Opus 4.xMediumLast stable Opus 4.x before Opus 5; 1M context at the same $5/M input rate.
Claude Opus 4.7
Anthropic$5.00$25.001,000kFrontierPinned productionMediumPinned production Opus for teams that value output stability over the latest weights.
Claude Fable 5
Anthropic$10.00$50.001,000kFrontierLong-context frontierMediumAnthropic’s premium tier above Opus 5; 1M context at Opus-5-fast pricing.
Claude Opus 4.1
Anthropic$15.00$75.00200kFrontierEnterprise legacyMediumLong-serving enterprise Opus; 200k context at premium $15/M pricing.
GPT-5.5 Pro
OpenAI$30.00$180.001,050kReasoningDeep analysisSlowHighest-effort GPT-5.5 tier; the pre-5.6 flagship for reasoning workloads.
GPT-5.4 Pro
OpenAI$30.00$180.001,050kReasoningHard math / codeSlowHigh-effort variant of GPT-5.4; long-context reasoning without 5.5 pricing.

Prices reflect each provider's public list price, auto-refreshed daily via the OpenRouter API. Open-weights rows are marked as estimates (varies by host: Together, Groq, Fireworks, etc.). Last updated methodology & sources for per-provider details.

How to reduce your LLM API costs

Six levers that consistently bring monthly LLM bills down by 30–70% in production.

  • Pick the right tier for the job

    Use a cheap fast model (GPT-4o mini, Gemini Flash, Claude Haiku) for routing, classification, and most production tasks. Reserve frontier models for the hard 5–10% of calls.

  • Cache prompt prefixes

    Anthropic, OpenAI, and Google all expose prompt caching. A long system prompt or RAG context replayed across requests can be served at ~10% of normal input price.

  • Cap the output, not the input

    Output tokens are 3–5× more expensive than input. Set max_tokens, use stop sequences, and ask the model for structured JSON instead of prose whenever possible.

  • Route by complexity

    Send simple queries to a fast model and only escalate to a frontier model when a confidence check or eval fails. A two-tier router cuts cost 40–70% in production.

  • Compress context before sending

    Summarise long histories, dedupe retrieved chunks, and strip boilerplate from system prompts. Most teams over-fill the context window by 2–3×.

  • Batch & stream where you can

    OpenAI and Anthropic Batch APIs run within 24h at ~50% off. Streaming doesn’t change the bill but lets you cancel mid-flight when the user navigates away.

Understanding LLM API pricing

LLM API pricing is almost always quoted per million tokens, separately for the prompt you send in (“input”) and the text the model writes back (“output”). A token is roughly 3–4 characters of English, so 1,000 tokens ≈ 750 words. The total bill for any given call is simply: input_tokens × input_rate + output_tokens × output_rate.

Three modifiers can change that base number meaningfully. Prompt caching lets you replay long system prompts or retrieved context at ≈10% of the normal input rate, which is enormous for RAG. Batch APIs (OpenAI, Anthropic) trade synchronous latency for a 50% discount on jobs that can wait up to 24 hours. And fine-tuning generally costs 1.5–3× the base inference rate, so the math only works if you’re saving on prompt length or quality at scale.

Most teams underestimate output cost. Output tokens are typically 3–5× more expensive than input tokens, and chatty models pad responses unless you cap them. The single highest-leverage change is usually setting max_tokens aggressively and asking for structured JSON instead of prose.

Methodology & sources

How this calculator stays fresh

Prices auto-refresh daily via the OpenRouter API, with each provider’s official pricing page as the underlying source of truth. Pricing last updated: .

Live pricing · 57/57 rows from OpenRouter

Calculation methodology

The estimated monthly bill you see on the right of the calculator is a deterministic function of four inputs — the selected model, input tokens per request, output tokens per request, and requests per day. The exact formula is:

monthly_cost = ((input_tokens_per_request × input_price + output_tokens_per_request × output_price) ÷ 1,000,000) × requests_per_day × 30

Assumptions

  • A 30-day month, matching how every major provider reports usage on invoices.
  • List prices published by each provider — no volume discounts, committed-use credits, or negotiated enterprise rates.
  • No prompt caching, batch discounts or fine-tuning surcharges (those levers are covered separately in the “How to reduce cost” section).
  • Token counts are treated as-billed: providers charge for reasoning traces and tool-call outputs where applicable, so numbers can undershoot in agentic workloads.
  • Open-weights rows (Llama, Mixtral, Qwen, self-hosted) are marked as estimates because effective cost depends on the host you pick and your GPU utilisation.

Caveats

  • Prices change frequently. Always confirm live rates on the official provider page linked below before contracting.
  • Region-specific rates (Vertex AI EU, Bedrock cross-region) can differ by 10–30% from the base US price shown here.
  • Fine-tuned model inference is billed at a premium (typically 1.5–3× base). This calculator prices the base model only.

Update history

A running log of every refresh — the topmost entry is regenerated automatically on each ISR revalidation; older entries capture the broader release cadence.

  1. Automated daily refresh
    • Live prices pulled from the OpenRouter API and merged onto the canonical model list.
    • Applied to the calculator, comparison table and JSON-LD dateModified in the same build.
  2. Q3 2026 refresh
    • Re-verified every row against each provider’s public pricing page.
    • Added Mistral Small 3 (Feb 2026 release) and refreshed DeepSeek V3 rates after the March discount.
    • Widened Google Gemini 2.0 Flash context note to reflect the 1M-token GA.
  3. Q2 2026 refresh
    • Re-priced OpenAI o1 and o1-mini after the June rate cut.
    • Marked Claude 3 Opus as legacy pending Anthropic’s deprecation notice.
  4. DeepSeek V3 launch update
    • Added DeepSeek V3 and DeepSeek Coder V2 at launch prices.
    • Refreshed Llama 3.1 family rates from Together AI / Groq.
  5. January 2026 refresh
    • Backfilled release years and category tags on every row.
    • Introduced the `isEstimate` flag for self-hosted and open-weights rows.
  6. Initial publish
    • First public release of the calculator with 22 models across 8 providers.
    • Established the "prices per 1M tokens" convention shown throughout the page.

Data sources & last verified dates

Live prices come from the OpenRouter API on a daily ISR schedule. The provider rows below are the underlying source of truth we (and OpenRouter) mirror — where a provider does not publish first-party API rates (Meta’s open-weights family, self-hosted rows) we cite the third-party host we sourced the number from and flag those rows as estimates in the comparison table.

ProviderOfficial pricing pageLast verified
OpenRouter API (automated feed)Live

Primary source. All prices in the table are refreshed daily from this endpoint via ISR; providers below are the underlying source of truth OpenRouter mirrors.

openrouter.ai/api/v1/models
Anthropic
platform.claude.com/docs/en/about-claude/pricing
Cohere
cohere.com/pricing
DeepSeek
api-docs.deepseek.com/quick_start/pricing
Google

Gemini API rates on AI Studio; Vertex AI enterprise rates may differ.

ai.google.dev/pricing
Meta

Meta does not publish first-party API pricing. Rates shown are the median across Together AI, Groq and Fireworks; marked as estimates.

together.ai/pricing
Mistral
mistral.ai/pricing/
OpenAI
developers.openai.com/api/docs/pricing
Self-hosted

Self-hosted rows are illustrative $/1M-token equivalents assuming a well-utilised A10/A100 pod; ignore idle capacity.

together.ai/pricing
xAI
docs.x.ai/docs/models

Pricing pages are published and controlled by each provider — we don’t mirror them. If you spot a stale rate, open the relevant source above, then let us know so we can push a mid-quarter update.

Frequently asked questions

Ready to scale with confident pricing?

Plug your real traffic into the calculator above, then dig into our directory of 1,000+ AI tools and frameworks to ship your next product.