Compare every major AI model
AI Models Compared
Compare the best AI models in one place. Explore benchmark scores, API pricing, speed, context windows, reasoning ability, multimodal capabilities, and real-world performance across OpenAI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, Alibaba, Midjourney, Runway, ElevenLabs, and more.
What is an AI model?
An AI model is a software system trained on large amounts of data so it can generate predictions, text, images, audio, or video. In product language, “AI model” usually means a foundation model — a large neural network trained once at enormous cost, then adapted to many tasks through prompting, fine-tuning, or tool use. Large language models (LLMs) such as GPT-5.5, Claude 4, and Gemini 2.5 take text in and return text out. Multimodal models also accept or produce images, audio, or video. Open-weights models such as Llama and DeepSeek can be downloaded and run on your own hardware; proprietary models are typically reached through a vendor API.
GPT-5.5 currently leads our intelligence rankings, while DeepSeek V3 offers one of the lowest API costs among strong open-weights options. Claude 4 Opus remains a top choice for writing and long-context code, and Gemini 2.5 Pro excels with long-context workflows thanks to its 2M-token window.
When people ask which model is “best,” they are usually asking several questions at once: which reasons most accurately, which answers fastest, which costs the least at production volume, which can read the longest documents, and which handles code, writing, or vision most reliably. No single score answers all of those. That is why this page compares models across a small set of transparent axes instead of crowning one universal winner — or open a GPT-5.5 vs Claude 4 Opus matchup for a head-to-head view.
How are AI models compared?
We compare models on public, reproducible metrics whenever possible. For language models the core columns are intelligence (a composite of academic benchmarks), generation speed in tokens per second, time-to-first-token latency, context window size, input and output API prices, and a blended price that reflects a typical traffic mix. Image, video, and speech models use modality-appropriate pricing — per image, per second of video, or per minute of audio — and appear on their image, video, and speech category pages with columns that match how those products are sold.
Side-by-side matchups let you pit any two models against each other on the same axes, with notes on strengths, weaknesses, and who each option is best for. Use the rankings below to scan the field, then open individual model pages or the comparator when you are ready to decide.
What is the Intelligence Index?
The Intelligence Index is our 0–100 summary of how strong a language model is on widely cited public benchmarks, including MMLU, MMLU-Pro, GPQA, MATH, and HumanEval. Higher is better — today GPT-5.5 leads, with Claude 4 Opus and Gemini 2.5 Pro close behind. The index is a rough ordering for scanning the catalog — not a substitute for reading the underlying scores. A model can lead on coding while trailing on knowledge-heavy exams and still land in the middle of the pack.
Reasoning models that spend more compute at inference time — such as OpenAI o1 and DeepSeek R1 — often score higher on hard math and planning tasks; chat-oriented models may look closer on everyday writing. Always treat the index as a starting point, then check task-specific benchmarks and, when stakes are high, run your own evaluation set.
How speed is measured
Speed on this site means median output tokens per second during steady-state generation across major API providers — not the time until the first token appears. Latency (time to first token, or TTFT) is listed separately because interactive apps feel sluggish when the first token is slow even if later tokens stream quickly.
Reasoning models that pause to “think” before answering will look slower on tokens-per-second; that is expected and does not mean the API is broken. Fast tiers like GPT-4o mini, Claude 3.5 Haiku, and Gemini 2.0 Flash usually top the speed board. Hosted open-source endpoints can also vary by provider, so treat speed as a directional signal rather than a guaranteed SLA.
Pricing methodology
Pricing for token-based models is shown as dollars per million input tokens and dollars per million output tokens, using list API rates from each vendor. Blended price mixes those rates at a 3:1 input-to-output ratio, which approximates many production workloads where prompts and retrieved context outweigh generated replies. That single number makes it easier to compare models whose input and output prices diverge sharply — for example Llama 3.1 8B at the budget end versus frontier rates on GPT-5.5.
Image, video, and audio models are priced in their native units. Promotional credits, enterprise discounts, and batch or cached-input rates can change effective cost — our tables reflect public list pricing so comparisons stay apples-to-apples.
Why benchmark scores differ
Benchmark scores differ across leaderboards for several good reasons. Contaminated or overlapping evaluation sets reward models that memorized similar questions. Prompt formatting, temperature, tool use, and whether chain-of-thought is allowed can swing results by large margins. Some vendors report self-reported scores; others are third-party replications. Multimodal and agentic tasks are still inconsistently measured.
When two sources disagree, prefer transparent methodology, look at the full score suite rather than a single headline number, and validate on your own prompts before locking in a production choice. See how we rank for the full methodology behind this catalog.
How often rankings update
Rankings and prices on What’s The Big Data are refreshed monthly. Each model entry carries a last-updated stamp; we re-pull API prices and public benchmark figures in the first week of each month. New model releases are added as soon as enough reliable data exists to place them fairly. If you spot a stale price or missing release, tell us — keeping the catalog current is the point of a living comparison page.
Top picks across each axis
Six quick leaderboards so you can shortlist without scrolling.
How we rank AI models
Transparent methodology for the intelligence index, speed figures, and API prices on this site — so you can trust the rankings and reproduce the logic. For a worked example, compare GPT-5.5 vs Claude 4 Opus using the same axes described below.
Intelligence score calculation
The Intelligence Index is a 0–100 composite that summarises how strong a language model is on widely cited public exams. We normalise each underlying benchmark to a 0–100 scale, then combine those scores into a single index so models can be scanned on one axis.
GPT-5.5 currently leads our intelligence rankings, while DeepSeek V3 offers one of the lowest API costs among strong open-weights options. Claude 4 Opus remains a top choice for writing tasks, and Gemini 2.5 Pro excels with long-context workflows. The index is a rough ordering for discovery — not a claim that one model is best at every task. Always open the model page for the full score breakdown.
Benchmark weighting
Language-model scores currently draw on MMLU, MMLU-Pro, GPQA, MATH, and HumanEval. Knowledge and reasoning suites (MMLU / MMLU-Pro / GPQA) carry the largest combined weight; MATH and HumanEval ensure hard quantitative and coding ability still move the index.
When a vendor or third-party report is missing a benchmark, we do not invent a perfect score — that entry is omitted from the composite or marked as partial data completeness so incomplete rows sort below fully measured models.
Speed testing
Speed is reported as median output tokens per second during steady-state generation across major API providers — not the time until the first token appears. Time to first token (TTFT) is listed separately because chat UIs feel slow when the first token lags, even if later tokens stream quickly.
Reasoning models that pause to “think” before answering will look slower on tokens-per-second; that is expected. Hosted open-weights endpoints can also vary by provider, so treat speed as a directional signal rather than a guaranteed SLA.
API price normalization
Token-priced models show list USD rates per million input tokens and per million output tokens from each vendor’s public API pricing. Blended price mixes those rates at a 3:1 input-to-output ratio — a common production traffic mix — so models with very different input vs output prices can be compared on one number.
Image, video, and speech models are priced in their native units (per image, per second of video, or per 1k characters). Enterprise discounts, batch APIs, and cached-input rates are excluded so comparisons stay apples-to-apples on public list pricing.
Update frequency
Rankings, prices, and benchmark figures are refreshed monthly. We re-pull public rates and scores in the first week of each month, and every model row carries a last-updated stamp. New releases are added once enough reliable data exists to place them fairly alongside peers.
Data sources
Primary inputs are vendor model cards and pricing pages, plus widely reported public benchmark results (including lab reports and independent leaderboards). Where sources disagree, we prefer transparent methodology and third-party replications over single self-reported headlines.
This catalog is editorially maintained — not an affiliate-sorted marketplace. If you spot a stale price, missing release, or incorrect score, tell us and we will correct it in the next refresh cycle.
AI Model Rankings Table
Click any column to sort. Filter by provider, license, capabilities, context, price, or release year.
| GPT-5.5 | OpenAI | 82 | 95 tok/s | 0.42s | 400k | $5.00 | $15.00 | $7.50 |
| Claude 4 Opus | Anthropic | 81 | 50 tok/s | 1.4s | 500k | $8.00 | $40.00 | $16.00 |
| Gemini 2.5 Pro | 78 | 110 tok/s | 0.7s | 2M | $1.25 | $5.00 | $2.19 | |
| OpenAI o1 | OpenAI | 76 | 32 tok/s | 12s | 200k | $15.00 | $60.00 | $26.25 |
| Claude 4 Sonnet | Anthropic | 75 | 95 tok/s | 0.85s | 500k | $3.00 | $15.00 | $6.00 |
| Grok 3 | xAI | 74 | 75 tok/s | 0.6s | 1M | $3.00 | $15.00 | $6.00 |
| DeepSeek R1 | DeepSeek | 73 | 60 tok/s | 1.5s | 128k | $0.55 | $2.19 | $0.96 |
| GPT-4o | OpenAI | 72 | 110 tok/s | 0.4s | 128k | $2.50 | $10.00 | $4.38 |
| Claude 3.5 Sonnet | Anthropic | 71 | 85 tok/s | 0.9s | 200k | $3.00 | $15.00 | $6.00 |
| OpenAI o3-mini | OpenAI | 70 | 60 tok/s | 6.5s | 200k | $1.10 | $4.40 | $1.93 |
| GPT-5.5 mini | OpenAI | 68 | 180 tok/s | 0.28s | 400k | $0.25 | $1.00 | $0.44 |
| DeepSeek V3 | DeepSeek | 67 | 90 tok/s | 0.45s | 64k | $0.27 | $1.10 | $0.48 |
| Gemini 1.5 Pro | 67 | 60 tok/s | 0.9s | 2M | $1.25 | $5.00 | $2.19 | |
| Llama 3.3 70B* | Meta | 66 | 200 tok/s | 0.4s | 128k | $0.23 | $0.40 | $0.27 |
| Gemini 2.0 Flash | 64 | 220 tok/s | 0.3s | 1M | $0.10 | $0.40 | $0.18 | |
| Llama 3.1 405B* | Meta | 64 | 32 tok/s | 0.7s | 128k | $2.70 | $2.70 | $2.70 |
| Qwen 2.5 72B* | Alibaba | 62 | 55 tok/s | 0.5s | 128k | $0.40 | $0.40 | $0.40 |
| Mistral Large 2 | Mistral | 60 | 70 tok/s | 0.6s | 128k | $2.00 | $6.00 | $3.00 |
| Claude 3.5 Haiku | Anthropic | 58 | 130 tok/s | 0.5s | 200k | $0.80 | $4.00 | $1.60 |
| GPT-4o mini | OpenAI | 56 | 145 tok/s | 0.32s | 128k | $0.15 | $0.60 | $0.26 |
| Mistral Small 3 | Mistral | 52 | 150 tok/s | 0.3s | 32k | $0.20 | $0.60 | $0.30 |
| Llama 3.1 8B* | Meta | 44 | 750 tok/s | 0.2s | 128k | $0.06 | $0.06 | $0.06 |
Showing 22 of 22 models. Click any column header to sort. Prices are USD per 1M tokens unless noted otherwise. Estimates marked with *.
AI Model Timeline
A short history of landmark releases — how today’s frontier and open-weights models got here.
- 2023
The modern frontier era begins — GPT-4 sets a new bar; open weights go mainstream.
- GPT-4OpenAI
- Claude 2Anthropic
- Llama 2Meta
- 2024
Multimodal defaults, long context, and open-weights challengers arrive at scale.
- 2025
Reasoning models and next-gen flagships tighten the race across labs.
- 2026
Current frontier — higher reasoning ceilings, longer context, and sharper cost tiers.
AI Model Charts
Visual snapshots of intelligence, price, speed, context, and provider coverage — click any model to open its full page.
Intelligence Leaderboard
Top models by intelligence index (0–100)
API Price Comparison
Blended $/1M tokens vs intelligence index
Speed vs Cost
Bubble size = intelligence index
Context Window Comparison
Largest context windows (tokens)
Provider Market Share
Share of models in this catalog by provider
- OpenAI 11 · 31%
- Anthropic 4 · 11%
- Google 4 · 11%
- Meta 3 · 8%
- DeepSeek 2 · 6%
- Mistral 2 · 6%
- xAI 1 · 3%
- Alibaba 1 · 3%
- Midjourney 1 · 3%
- Black Forest Labs 1 · 3%
- Stability AI 1 · 3%
- Runway 1 · 3%
- Kuaishou 1 · 3%
- ElevenLabs 1 · 3%
- Cartesia 1 · 3%
- Voyage AI 1 · 3%
Best AI Models for Coding
Picking a coding model is less about a single leaderboard number and more about the work you do every day: multi-file refactors, greenfield features, test generation, or hard algorithmic puzzles. On public coding exams such as HumanEval, Claude 4 Opus currently sits at the front of the pack, with GPT-5.5 close behind when the workflow includes tool use and agent loops. Those two are the defaults most engineering teams should evaluate first for production code assistance.
For repository-scale edits and long pull requests, Claude 4 Opus and Claude 4 Sonnet benefit from large context windows and strong instruction following. Sonnet is usually the better everyday workhorse — near-Opus quality at a friendlier price — while Opus is worth reserving for the toughest migrations. GPT-5.5 shines when the model must call linters, browsers, or internal APIs as part of an agent. OpenAI o1 and DeepSeek R1 remain the strongest picks when the task is pure reasoning: contest math, tricky algorithms, or proofs buried in a codebase.
Budget and volume matter too. GPT-5.5 mini, Gemini 2.0 Flash, and Claude 3.5 Haiku handle autocomplete-style prompts and cheap draft generations without burning the monthly bill. On the open-weights side, DeepSeek V3 and Llama 3.3 70B are credible self-hosted coding partners when data cannot leave your VPC. Browse our coding models hub, then open the GPT-5.5 vs Claude 4 Opus comparison before you standardize on a stack — and keep a small eval suite of real PRs so marketing benchmarks do not override what your engineers actually feel.
Related: Compare models · Rankings table · How we rank
Best AI Models for Writing
Long-form writing rewards coherence, taste, and the ability to hold a brief across thousands of tokens — not just raw benchmark scores. Claude 4 Opus and Claude 4 Sonnet are widely preferred for essays, docs, marketing drafts, and narrative work because they produce less repetitive prose and stay on voice longer than many peers. GPT-5.5 is often stronger when writing is tightly coupled to research, structured outlines, or multi-step briefs that mix analysis with drafting.
For editorial teams, Sonnet is usually the daily driver: fast enough for iteration, strong enough for publishable first drafts, and cheaper than Opus for high volume. Reach for Opus when the piece is flagship — a long report, a sensitive rewrite, or a document that must absorb a large source pack in one pass. Gemini 2.5 Pro earns a seat when the brief includes huge source libraries or mixed media context, thanks to its class-leading context window.
Lightweight models still have a role. GPT-4o mini and Gemini 2.0 Flash are fine for outlines, social variants, and SEO snippets where latency and cost matter more than literary polish. Open-weights options such as Llama 3.3 70B and Qwen 2.5 72B can power private writing assistants on your own hardware. Whatever you choose, keep a human editor in the loop for factual claims and brand voice — models draft; they do not own the byline. Explore individual model pages for strengths, weaknesses, and best-fit use cases before you lock a writing stack for the quarter.
Related: Compare models · Rankings table · How we rank
Best AI Models for Startups
Startups rarely need the absolute top intelligence index on day one. They need a model that ships features, keeps unit economics sane, and can be swapped later without rewriting the product. That usually means a mid-tier workhorse plus a premium model for the hard 10% of traffic. GPT-5.5 mini, Claude 4 Sonnet, Gemini 2.0 Flash, and DeepSeek V3 are the common shortlist: strong enough for chat, support, and light agents, cheap enough to iterate through product-market fit.
When quality becomes the bottleneck — complex agents, deep coding, or customer-facing research — graduate hot paths to GPT-5.5 or Claude 4 Opus rather than upgrading every call. A simple router (cheap model first, escalate on failure or low confidence) often cuts spend in half versus sending everything to the frontier. OpenAI o1 and DeepSeek R1 are worth a dedicated lane for planning and hard reasoning jobs that would otherwise burn scarce engineer time in a five-person team.
Self-hosting is attractive once volume is high and privacy requirements are strict. Llama 3.3 70B and DeepSeek V3 can run behind your own auth with predictable GPU cost, while proprietary APIs stay available as overflow. Track blended price, speed, and failure rates from week one so you know when to switch. Use the price and capability filters on the rankings table, then open head-to-head comparisons before you commit an annual contract or bake a single vendor into your architecture for the next eighteen months.
Related: Compare models · Rankings table · How we rank
Best Open Source AI Models
Open-weights models matter when you need to self-host, fine-tune, air-gap, or escape per-token API bills at scale. In 2026 the quality gap to proprietary frontier models has narrowed sharply on many tasks. DeepSeek R1 leads open reasoning, DeepSeek V3 is a strong generalist, and Llama 3.3 70B remains the most widely deployed community default. Llama 3.1 405B still shows up where raw parameter count and research baselines matter, while Qwen 2.5 72B and Mistral Small 3 cover multilingual and efficient serving niches.
“Open source” is not one license. Llama’s community license restricts very large commercial users; others are closer to Apache-style terms. Always read the license before you ship. Operationally, smaller weights (8B–70B) fit single-node or modest multi-GPU setups; 405B-class models need serious clusters or a hosted open-weights provider such as Together, Fireworks, or Groq. For image generation, Stable Diffusion 3.5 remains the open ecosystem to beat for fine-tunes and local creative tools.
When should you stay proprietary? If you need best-in-class tool use, the absolute top writing quality, or a vendor SLA with zero ops load, GPT-5.5, Claude 4 Opus, and Gemini 2.5 Pro still win many bake-offs. A hybrid pattern is common: open weights for bulk inference, closed APIs for the hardest prompts. Start from our open-weights and open-source hubs, sort the rankings table by intelligence and blended price, and validate on your own evaluation set before you migrate production traffic away from a managed API.
Related: Compare models · Rankings table · How we rank
Best AI Models by Price
API price is easy to misread if you only look at input rates. Real apps mix long prompts with shorter completions, which is why we publish a blended $/1M figure at a 3:1 input-to-output ratio. On that axis, Llama 3.1 8B is among the cheapest full entries in the catalog, while Gemini 2.0 Flash, GPT-4o mini, and Llama 3.3 70B deliver much stronger quality still under roughly thirty cents per million blended tokens. Those models are the right starting point for high-volume classification, extraction, and light chat.
Mid-tier options such as Claude 4 Sonnet, GPT-5.5 mini, and DeepSeek V3 occupy the value band where most products should live: good enough quality that users stay, cheap enough that margins survive. Frontier models — GPT-5.5, Claude 4 Opus, Gemini 2.5 Pro — cost several dollars per million tokens and should be reserved for prompts that actually need them. Reasoning models add another twist: they may look expensive per token and slower on tok/s because they spend hidden compute thinking, yet they can still be cheaper than a junior engineer hour on hard tasks.
Self-hosting flips the spreadsheet. Open-weights models have near-zero marginal API cost once GPUs are paid for, but you inherit utilization, reliability, and MLOps work. Compare list prices in our rankings table, use the price filters to shortlist under $0.50 or $2 blended, then open model pages for input versus output breakdowns. Pair any cost decision with quality checks — the cheapest model that fails your eval is the most expensive mistake. For a visual view, jump to the API price and speed-versus-cost charts above.
Related: Compare models · Rankings table · How we rank
Browse AI Models by category
Drill into a slice of the catalog — reasoning, coding, image, video, speech, embeddings, agents, open weights, and more.
By Use case
Frontier
10The most capable models from each major lab.
Reasoning Models
7Long chain-of-thought models built for hard math, code, and planning.
Coding Models
14Models specialised for software engineering and code generation.
Agents
13Models strong at tool use, multi-step workflows, and agentic systems.
Fast & Cheap
6Production-grade workhorses with the best speed and cost.
Multimodal
11Models that natively understand text, images, and beyond.
By Modality
Image Models
4Generate and edit still images from text or reference images.
Video Models
6Generate short video clips from text or image prompts.
Audio Models
7Models that understand or generate audio beyond plain speech TTS.
Speech Models
3Text-to-speech and voice models for natural spoken audio.
Embeddings
4Embedding models for search, RAG, clustering, and semantic similarity.
By License
By Audience
By Provider
OpenAI
11All models from OpenAI — GPT, o-series and beyond.
Anthropic
4Claude family of models from Anthropic.
Gemini family of models from Google DeepMind.
Meta
3Open-weights Llama family from Meta AI.
Mistral
2Models from Mistral AI — EU-hosted, multilingual.
DeepSeek
2Frontier-class open-weights models from DeepSeek.
xAI
1Grok models from xAI.
Alibaba
1Qwen family of open-weights models from Alibaba.
Voyage AI
1Embedding models from Voyage AI.
Midjourney
1Midjourney text-to-image models.
Black Forest Labs
1FLUX image-generation models from Black Forest Labs.
Stability AI
1Stable Diffusion image models from Stability AI.
Runway
1Gen-series text-to-video models from Runway.
Kuaishou
1Kling text-to-video models from Kuaishou.
ElevenLabs
1Text-to-speech voice models from ElevenLabs.
Cartesia
1Sonic low-latency text-to-speech models from Cartesia.
Recently Updated
Models with the freshest benchmark and pricing data.
Trending This Week
Models people are comparing and reading about right now.
Popular AI model comparisons
One-click comparisons for the matchups people search the most.
Frequently asked questions
Pick a model with confidence
Open any model below for a deep-dive — or pit two models head-to-head with the side-by-side comparator.