Estimate the exact monthly cost of every OpenAI model — GPT-5.6 Sol / Terra / Luna, GPT-5.5, GPT-5.4, GPT-5, GPT-4.1, GPT-4o and the o-series reasoners. Includes Standard, Batch, Flex and Fast (priority) pricing, plus image, audio, video, embedding and tool rates.
Pick your unit, dial in a typical request size, and instantly see the per-call cost and monthly total across every current OpenAI model — sorted from cheapest to most expensive by default.
Calculate by
× 30 = 15,000 /month
Show:
Cheapest: GPT-5 Nano — $3.75/moPriciest: o1 Pro — $6,750/mo1800× spread across the range.
GPT-5 Nano
Cheapest 5-series tier; strong replacement for GPT-4.1 nano.
GPT-5
Bulk classification
Very fast
$0.0500
$0.4000
$0.0003
$3.75
GPT-4.1 Nano
Cheapest long-context model in the 4.1 line.
GPT-4.1
Cheap long-context
Very fast
$0.1000
$0.4000
$0.0003
$4.50
GPT-4o Mini
Still a very cheap workhorse — but GPT-5 mini beats it on most benchmarks.
GPT-4o
Legacy fast chat
Very fast
$0.1500
$0.6000
$0.0005
$6.75
GPT-5.6 Luna
Cheapest 5.6-series model — pick this for high-volume, low-latency work.
GPT-5.6
High-volume tasks
Very fast
$0.2000
$1.20
$0.0008
$12.00
GPT-5.4 Nano
Fastest and cheapest 5.4 tier — great for classifiers and simple structured outputs.
GPT-5.4
Classifiers
Very fast
$0.2000
$1.25
$0.0008
$12.38
GPT-4.1 Mini
Fast, long-context mini — good default for RAG.
GPT-4.1
Long-context RAG
Very fast
$0.4000
$1.60
$0.0012
$18.00
GPT-5 Mini
Chatty, cheap and fast — great replacement for GPT-4o mini.
GPT-5
Cheap chat
Very fast
$0.2500
$2.00
$0.0013
$18.75
GPT-5.4 Mini
Sub-second first-token times at a fraction of the flagship price.
GPT-5.4
Low-latency chat
Very fast
$0.7500
$4.50
$0.0030
$45.00
o3 Mini
Cheap reasoning tier for straightforward multi-step problems.
o-series
Cheap reasoning
Fast
$1.10
$4.40
$0.0033
$49.50
o4 Mini
Latest cheap reasoning tier — usually the right default for agent loops.
o-series
Agent default
Fast
$1.10
$4.40
$0.0033
$49.50
GPT-4.1
1M-token context; great for whole-repo or long-document workloads.
GPT-4.1
Long-context RAG
Fast
$2.00
$8.00
$0.0060
$90.00
o3
The pragmatic reasoning default — much cheaper than o1 with similar quality.
o-series
Agent loops
Medium
$2.00
$8.00
$0.0060
$90.00
GPT-5.1
Intermediate release — same rates as GPT-5.
GPT-5.1
General chat
Fast
$1.25
$10.00
$0.0063
$93.75
GPT-5
The original GPT-5 — still a competitive baseline.
GPT-5
Bulk classifiers
Fast
$1.25
$10.00
$0.0063
$93.75
GPT-4o
Legacy multimodal flagship. Prefer GPT-5.4 or GPT-5.6 Terra for new work.
GPT-4o
Legacy multimodal
Fast
$2.50
$10.00
$0.0075
$112.50
GPT-5.6 Terra
The everyday default — best price/quality ratio in the GPT-5.6 line.
GPT-5.6
Everyday production
Fast
$2.00
$12.00
$0.0080
$120.00
GPT-5.2
Older mid-tier — usually cheaper to move to GPT-5.4 or GPT-5.6 Terra.
GPT-5.2
Legacy chat
Fast
$1.75
$14.00
$0.0088
$131.25
GPT-5.4
The workhorse — the price/quality sweet spot for most production traffic.
Previous flagship — still competitive on price and quality.
GPT-5.5
Migration workloads
Medium
$5.00
$30.00
$0.0200
$300.00
o1
Original chain-of-thought reasoner — expensive but unbeaten on hard math/code.
o-series
Hard reasoning
Slow
$15.00
$60.00
$0.0450
$675.00
o3 Pro
Long-thinking o3 variant — useful when o3 fails on a hard problem.
o-series
Long thinking
Slow
$20.00
$80.00
$0.0600
$900.00
GPT-5 Pro
Extended-thinking tier of GPT-5.
GPT-5
Deep reasoning
Slow
$15.00
$120.00
$0.0750
$1,125
GPT-5.2 Pro
Older pro-tier reasoning model.
GPT-5.2
Legacy reasoning
Slow
$21.00
$168.00
$0.1050
$1,575
GPT-5.5 Pro
Deep-reasoning tier for hardest problems in code and analysis.
GPT-5.5
Deep analysis
Slow
$30.00
$180.00
$0.1200
$1,800
GPT-5.4 Pro
Extended-thinking variant for hard math/code — 12× cost of vanilla 5.4.
GPT-5.4
Hard math / code
Slow
$30.00
$180.00
$0.1200
$1,800
o1 Pro
Highest-effort o1 tier — use sparingly, only for problems o1 fails.
o-series
Extreme reasoning
Slow
$150.00
$600.00
$0.4500
$6,750
Totals = per-call cost × API calls/day × 30 days. Standard, short-context rates from OpenAI’s official pricing page.Batch (‑50%), Flex (‑50%) and Fast/Priority (×2) tiers are shown in the tables below.
OpenAI model families
Which OpenAI model should you actually pick?
OpenAI now ships six active text-model families plus the o-series reasoners. Here’s the one-line brief for each, the workloads they were built for, and the price range you’ll pay.
GPT-5.6
The newest flagship family (Sol / Terra / Luna).
Best for: Any new project — Sol for hard problems, Terra for the default, Luna for high-volume tasks.
GPT-5.6 Sol$5.00 in · $30.00 out /1M
GPT-5.6 Terra$2.00 in · $12.00 out /1M
GPT-5.6 Luna$0.200 in · $1.20 out /1M
Range: $0.200–$5.00 input · $1.20–$30.00 output per 1M
GPT-5.5
Previous flagship — competitive on price and quality.
Best for: Migration targets from GPT-4-Turbo / GPT-4o that don’t yet need 5.6.
GPT-5.5$5.00 in · $30.00 out /1M
GPT-5.5 Pro$30.00 in · $180.00 out /1M
Range: $5.00–$30.00 input · $30.00–$180.00 output per 1M
GPT-5.4
The price/quality sweet spot for most production traffic.
Best for: Chatbots, summarisation, structured extraction, coding assistants.
GPT-5.4$2.50 in · $15.00 out /1M
GPT-5.4 Mini$0.750 in · $4.50 out /1M
GPT-5.4 Nano$0.200 in · $1.25 out /1M
GPT-5.4 Pro$30.00 in · $180.00 out /1M
Range: $0.200–$30.00 input · $1.25–$180.00 output per 1M
GPT-5
Original GPT-5 line — great cheap-ish workhorses.
Best for: High-volume classifiers and lightweight chat.
GPT-5$1.25 in · $10.00 out /1M
GPT-5 Mini$0.250 in · $2.00 out /1M
GPT-5 Nano$0.050 in · $0.400 out /1M
GPT-5 Pro$15.00 in · $120.00 out /1M
Range: $0.050–$15.00 input · $0.400–$120.00 output per 1M
GPT-4.1
1M-token context specialist.
Best for: Whole-repo code assistants, book-length summarisation, large-document RAG.
GPT-4.1$2.00 in · $8.00 out /1M
GPT-4.1 Mini$0.400 in · $1.60 out /1M
GPT-4.1 Nano$0.100 in · $0.400 out /1M
Range: $0.100–$2.00 input · $0.400–$8.00 output per 1M
GPT-4o
Legacy multimodal — kept alive for backwards compatibility.
Best for: Existing apps not yet ready to migrate. Prefer GPT-5.4 for new work.
GPT-4o$2.50 in · $10.00 out /1M
GPT-4o Mini$0.150 in · $0.600 out /1M
Range: $0.150–$2.50 input · $0.600–$10.00 output per 1M
o-series (reasoning)
Chain-of-thought reasoners for hard math, code and planning.
Best for: Multi-step agents, competitive programming, formal proofs, complex analysis.
o1$15.00 in · $60.00 out /1M
o1 Pro$150.00 in · $600.00 out /1M
o3$2.00 in · $8.00 out /1M
o3 Pro$20.00 in · $80.00 out /1M
o3 Mini$1.10 in · $4.40 out /1M
o4 Mini$1.10 in · $4.40 out /1M
Range: $1.10–$150.00 input · $4.40–$600.00 output per 1M
OpenAI service tiers
Standard vs Batch vs Flex vs Fast (Priority)
OpenAI meters the same models at four different rates depending on how much latency you can tolerate. Batch and Flex are 50% off, Fast (formerly Priority) is 2× the standard rate for a guaranteed low-latency queue.
Standard:Default pay-as-you-go rate for every model.
Latency: Real-timeAvailability: All models
Model
Input
Cached input
Output
Input (long ctx)
Cached (long)
Output (long)
GPT-5.6 Sol
GPT-5.6
$5.00
$0.500
$30.00
$10.00
$1.00
$45.00
GPT-5.6 Terra
GPT-5.6
$2.00
$0.200
$12.00
$4.00
$0.400
$18.00
GPT-5.6 Luna
GPT-5.6
$0.200
$0.020
$1.20
$0.400
$0.040
$1.80
GPT-5.5
GPT-5.5
$5.00
$0.500
$30.00
$10.00
$1.00
$45.00
GPT-5.5 Pro
GPT-5.5
$30.00
—
$180.00
—
—
—
GPT-5.4
GPT-5.4
$2.50
$0.250
$15.00
$5.00
$0.500
$22.50
GPT-5.4 Mini
GPT-5.4
$0.750
$0.075
$4.50
—
—
—
GPT-5.4 Nano
GPT-5.4
$0.200
$0.020
$1.25
—
—
—
GPT-5.4 Pro
GPT-5.4
$30.00
—
$180.00
$60.00
—
$270.00
GPT-5.2
GPT-5.2
$1.75
$0.175
$14.00
—
—
—
GPT-5.2 Pro
GPT-5.2
$21.00
—
$168.00
—
—
—
GPT-5.1
GPT-5.1
$1.25
$0.125
$10.00
—
—
—
GPT-5
GPT-5
$1.25
$0.125
$10.00
—
—
—
GPT-5 Mini
GPT-5
$0.250
$0.025
$2.00
—
—
—
GPT-5 Nano
GPT-5
$0.050
$0.005
$0.400
—
—
—
GPT-5 Pro
GPT-5
$15.00
—
$120.00
—
—
—
GPT-4.1
GPT-4.1
$2.00
$0.500
$8.00
—
—
—
GPT-4.1 Mini
GPT-4.1
$0.400
$0.100
$1.60
—
—
—
GPT-4.1 Nano
GPT-4.1
$0.100
$0.025
$0.400
—
—
—
GPT-4o
GPT-4o
$2.50
$1.25
$10.00
—
—
—
GPT-4o Mini
GPT-4o
$0.150
$0.075
$0.600
—
—
—
o1
o-series
$15.00
$7.50
$60.00
—
—
—
o1 Pro
o-series
$150.00
—
$600.00
—
—
—
o3
o-series
$2.00
$0.500
$8.00
—
—
—
o3 Pro
o-series
$20.00
—
$80.00
—
—
—
o3 Mini
o-series
$1.10
$0.550
$4.40
—
—
—
o4 Mini
o-series
$1.10
$0.275
$4.40
—
—
—
All rates in USD per 1,000,000 tokens. Long-context rates apply on GPT-5.5, GPT-5.4 and GPT-5.4 Pro when request length exceeds 272K tokens. “Cached input” is billed when a prompt prefix hits OpenAI’s automatic 5-minute prompt cache.
Beyond text
OpenAI image, audio, video, embedding and tool pricing
Text tokens are only part of the bill. Realtime voice, image generation with gpt-image-2, Sora 2 video, transcription with gpt-transcribe, embeddings and every hosted tool (web search, code interpreter, file search) has its own rate card.
Image generation
gpt-image-2 & gpt-image-1 pricing
Priced per 1M tokens for both text and image content. Image output tokens are the dominant cost for anything above a thumbnail.
Model
Text input
Cached text
Image input
Cached image
Image output
gpt-image-2
Newest image generation & editing model.
$5.00
$1.25
$8.00
$2.00
$30.00
gpt-image-1.5
Higher-fidelity 1.5 tier with text-in-image awareness.
$5.00
$1.25
$8.00
$2.00
$32.00
gpt-image-1 mini
Cheapest gpt-image tier — great for thumbnails and bulk work.
$2.00
$0.20
$2.50
$0.25
$8.00
gpt-image-1
Original gpt-image release — usually superseded by gpt-image-2.
$5.00
$1.25
$10.00
$2.50
$40.00
chatgpt-image-latest
Same model powering image generation inside ChatGPT.
$5.00
$1.25
$8.00
$2.00
$32.00
Video generation
Sora 2 & Sora 2 Pro pricing
Billed per second of finished video, not per token. Higher resolutions cost more; portrait and landscape are the same rate.
Model
Resolution
Sizes
Price / second
Sora 2
Standard 720p Sora 2 video generation.
720p
720×1280 / 1280×720
$0.10
Sora 2 Pro (720p)
Pro tier at 720p — faster iteration.
720p
720×1280 / 1280×720
$0.30
Sora 2 Pro (1024p)
1024p
1024×1792 / 1792×1024
$0.50
Sora 2 Pro (1080p)
1080p
1080×1920 / 1920×1080
$0.70
Realtime & voice
Realtime, voice and text-to-speech pricing
Realtime models bill separately for audio, text and image tokens on the same request. TTS models like tts-1 are billed per 1M characters.
Model
Audio in
Audio cached
Audio out
Text in
Text out
gpt-realtime-2.1
Newest realtime voice + text + vision model.
$32.00
$0.40
$64.00
$4.00
$24.00
gpt-realtime-2.1 mini
Cheapest realtime tier — 3× cheaper than the flagship.
$10.00
$0.30
$20.00
$0.60
$2.40
gpt-realtime-2
$32.00
$0.40
$64.00
$4.00
$24.00
gpt-realtime
$32.00
$0.40
$64.00
$4.00
$16.00
gpt-audio-1.5
$32.00
—
$64.00
$2.50
$10.00
gpt-audio-mini
$10.00
—
$20.00
$0.60
$2.40
gpt-4o-mini-tts
—
—
$12.00
$0.60
—
Model
Unit
Price
tts-1
$15.00 per 1M characters (input).
per 1M characters (input)
$15.00
tts-1-hd
$30.00 per 1M characters (input). Higher-fidelity voices.
Modern OpenAI ASR is per-minute; the older gpt-4o family also has per-token billing so long transcripts and diarised runs are cheaper to reason with downstream.
Model
Input tokens /1M
Output tokens /1M
Per minute
gpt-transcribe
Latest & cheapest ASR — $0.0045 per minute.
—
—
$0.0045
gpt-4o-transcribe
Token-priced + estimated $0.006/minute.
$2.50
$10.00
$0.0060
gpt-4o-mini-transcribe
$1.25
$5.00
$0.0030
gpt-4o-transcribe-diarize
Adds speaker diarization on top of gpt-4o-transcribe.
$2.50
$10.00
$0.0060
Whisper
The original OpenAI ASR — kept for legacy pipelines.
—
—
$0.0060
Embeddings
text-embedding-3 pricing
Embeddings are still the cheapest part of the OpenAI stack. Use text-embedding-3-small for most RAG; upgrade to text-embedding-3-large only when recall matters.
Model
Input /1M tokens
text-embedding-3-small
Cheap general-purpose embeddings — $0.02 per 1M tokens.
$0.020
text-embedding-3-large
Higher-quality embeddings for RAG at scale.
$0.130
text-embedding-ada-002
Legacy — prefer text-embedding-3-small unless downstream code depends on ada.
$0.100
Hosted tools
Web search, code interpreter, file search & Agent Kit
Hosted tools carry a per-call or storage cost on top of the model tokens they consume.
Tool
Pricing
Web search (all models)
Applies to all GPT-5.x, o-series and GPT-4.1 models.
$10 / 1k calls + search content tokens billed at model rate
Image web search
$10 / 1k calls + search content tokens billed at model rate
Web search preview (reasoning models)
$10 / 1k calls + search content tokens billed at model rate
Web search preview (non-reasoning)
$25 / 1k calls (search content tokens free)
Containers (Shell & Code Interpreter)
5-minute minimum per session; billed by the minute after.
$0.10 / GB-day (first 1 GB free per account per month)
Cost optimisation
How to cut your OpenAI API bill by 50–90%
Eight compounding levers we’ve seen shave OpenAI bills in half — in rough order of impact per hour of engineering work.
Use the Batch API for anything that isn’t user-facing
Up to 50% off
OpenAI’s Batch API returns completions within 24 hours at 50% off list price. Nightly summarisation, evals, document enrichment, dataset labelling, backfills, embeddings recomputes — anything that doesn’t need to render in a chat UI belongs on Batch.
Design prompts so the prefix is cache-hittable
~80% off input tokens
OpenAI automatically caches prompt prefixes ≥1024 tokens for ~5 minutes and replays them at 10–20% of the standard input rate (0.5% on GPT-4o, 10% on GPT-5.x). Put the stable system message, retrieved context and tool schemas at the start; put user turns at the end. RAG workloads see 40–70% input-cost reductions with no code changes.
Right-size the tier: Terra > Sol, Nano > Mini > Full
3–25× reduction
The GPT-5.6 line has a 25× spread from Luna to Sol; GPT-5.4 has a 12× spread from Nano to Pro. Most production traffic (classification, extraction, summarisation) works fine on a Mini or Nano tier. Reserve Sol / Pro for the 5–10% of requests that actually need frontier quality — route those explicitly.
Ask for structured JSON, not prose
2–10× on outputs
Output tokens cost 4–6× input tokens on every OpenAI model. A conversational answer might be 400 tokens; the same answer as { "verdict": "yes", "reasons": [...] } is 40. Combine this with response_format: { type: "json_schema" } and you get both cheaper and more parseable outputs.
Only reach for o-series when you actually need reasoning
5–10× vs default o1
The o-series (o1, o3, o3-pro) spends its money on invisible reasoning tokens — o1 emits 5–20K reasoning tokens on a single hard problem at $60/M output. For anything that isn’t multi-step math, code, or planning, GPT-5.6 Terra or GPT-5.4 will match quality at ~10% of the cost.
Skip Fast (Priority) unless latency is a hard requirement
Avoid 2× premium
Fast mode doubles the standard rate for a priority queue. If your P95 latency requirement is above ~2s (most agents, most chat apps), the Standard tier already meets it. Only pay the 2× premium when you have a documented, user-visible latency SLO.
Cap max_tokens on every request
Prevents 10–100× incidents
Bugs in prompt templates or agent loops can produce runaway completions that burn thousands of dollars in output tokens overnight. Set max_tokens to the largest sensible value for that endpoint (200 for a classifier, 800 for a chat reply, 4000 for a report) and return an error rather than a truncated reply — that error is your alarm bell.
Enable spend limits + usage alerts on your OpenAI project
Bounded worst case
Set a hard monthly limit and a soft alert at 50% and 80% inside the OpenAI dashboard. Combine with per-key tags (production, staging, evals) so a runaway staging job can’t take down your production budget.
Methodology & sources
How this OpenAI calculator stays honest
Every rate on this page is copied from OpenAI’s official pricing documentation and verified on a rolling monthly cadence. Last verified: .
Calculation formula
The number in the “Total cost” column is a deterministic function of four inputs — model, input tokens per call, output tokens per call, and calls per day:
Standard (real-time) rates, short-context (< 272K tokens) — the tier most requests fall into.
A 30-day billing month is used for the monthly total.
Prompt caching, Batch (‑50%), Flex (‑50%) and Fast (2×) tiers are shown in the tier table above and are not applied in the calculator by default.
For a model with tiered pricing (GPT-5.5, GPT-5.4, GPT-5.4 Pro), the short-context input/output rate is used.
Caveats
Reasoning tokens on the o-series are billed as output tokens and can be 5–20× the visible response length.
Vision / image inputs on GPT-5.x are billed at the same rate as text tokens after conversion — a 1024×1024 image is roughly 1,105 tokens.
Regional processing (data residency) endpoints carry a 10% uplift for models released after March 5, 2026.
Fine-tuned inference is billed at 1.5–3× the base model rate; see OpenAI’s fine-tuning table for exact rates.
Update history
Live verification
Cross-checked every rate against the official developers.openai.com/api/docs/pricing page.
Added GPT-5.6 Sol/Terra/Luna and refreshed short/long context rates on GPT-5.5, GPT-5.4 and GPT-5.4 Pro.
Added Sora 2 & Sora 2 Pro resolution-based per-second pricing.
Fast (Priority) rate refresh
Fast mode was renamed from Priority processing on July 30, 2026; both service_tier strings still work.
Applied the 2× multiplier to the new GPT-5.6 tier.
Batch API expansion
Batch discount now applies to the full GPT-5.6 line at 50% off.
Flex tier remains at 50% off but is queued behind Batch.
Primary source
All prices on this page are transcribed directly from OpenAI’s official developer documentation. Where OpenAI publishes multiple service tiers (Standard, Batch, Flex, Fast) or context-length tiers (short vs long), we mirror both.