Compare every major AI model

AI Models Compared

Compare the best AI models in one place. Explore benchmark scores, API pricing, speed, context windows, reasoning ability, multimodal capabilities, and real-world performance across OpenAI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, Alibaba, Midjourney, Runway, ElevenLabs, and more.

Models
36
Providers
16
Categories
35
Updated
2026-06

What is an AI model?

An AI model is a software system trained on large amounts of data so it can generate predictions, text, images, audio, or video. In product language, “AI model” usually means a foundation model — a large neural network trained once at enormous cost, then adapted to many tasks through prompting, fine-tuning, or tool use. Large language models (LLMs) such as GPT-5.5, Claude 4, and Gemini 2.5 take text in and return text out. Multimodal models also accept or produce images, audio, or video. Open-weights models such as Llama and DeepSeek can be downloaded and run on your own hardware; proprietary models are typically reached through a vendor API.

GPT-5.5 currently leads our intelligence rankings, while DeepSeek V3 offers one of the lowest API costs among strong open-weights options. Claude 4 Opus remains a top choice for writing and long-context code, and Gemini 2.5 Pro excels with long-context workflows thanks to its 2M-token window.

When people ask which model is “best,” they are usually asking several questions at once: which reasons most accurately, which answers fastest, which costs the least at production volume, which can read the longest documents, and which handles code, writing, or vision most reliably. No single score answers all of those. That is why this page compares models across a small set of transparent axes instead of crowning one universal winner — or open a GPT-5.5 vs Claude 4 Opus matchup for a head-to-head view.

How are AI models compared?

We compare models on public, reproducible metrics whenever possible. For language models the core columns are intelligence (a composite of academic benchmarks), generation speed in tokens per second, time-to-first-token latency, context window size, input and output API prices, and a blended price that reflects a typical traffic mix. Image, video, and speech models use modality-appropriate pricing — per image, per second of video, or per minute of audio — and appear on their image, video, and speech category pages with columns that match how those products are sold.

Side-by-side matchups let you pit any two models against each other on the same axes, with notes on strengths, weaknesses, and who each option is best for. Use the rankings below to scan the field, then open individual model pages or the comparator when you are ready to decide.

What is the Intelligence Index?

The Intelligence Index is our 0–100 summary of how strong a language model is on widely cited public benchmarks, including MMLU, MMLU-Pro, GPQA, MATH, and HumanEval. Higher is better — today GPT-5.5 leads, with Claude 4 Opus and Gemini 2.5 Pro close behind. The index is a rough ordering for scanning the catalog — not a substitute for reading the underlying scores. A model can lead on coding while trailing on knowledge-heavy exams and still land in the middle of the pack.

Reasoning models that spend more compute at inference time — such as OpenAI o1 and DeepSeek R1 — often score higher on hard math and planning tasks; chat-oriented models may look closer on everyday writing. Always treat the index as a starting point, then check task-specific benchmarks and, when stakes are high, run your own evaluation set.

How speed is measured

Speed on this site means median output tokens per second during steady-state generation across major API providers — not the time until the first token appears. Latency (time to first token, or TTFT) is listed separately because interactive apps feel sluggish when the first token is slow even if later tokens stream quickly.

Reasoning models that pause to “think” before answering will look slower on tokens-per-second; that is expected and does not mean the API is broken. Fast tiers like GPT-4o mini, Claude 3.5 Haiku, and Gemini 2.0 Flash usually top the speed board. Hosted open-source endpoints can also vary by provider, so treat speed as a directional signal rather than a guaranteed SLA.

Pricing methodology

Pricing for token-based models is shown as dollars per million input tokens and dollars per million output tokens, using list API rates from each vendor. Blended price mixes those rates at a 3:1 input-to-output ratio, which approximates many production workloads where prompts and retrieved context outweigh generated replies. That single number makes it easier to compare models whose input and output prices diverge sharply — for example Llama 3.1 8B at the budget end versus frontier rates on GPT-5.5.

Image, video, and audio models are priced in their native units. Promotional credits, enterprise discounts, and batch or cached-input rates can change effective cost — our tables reflect public list pricing so comparisons stay apples-to-apples.

Why benchmark scores differ

Benchmark scores differ across leaderboards for several good reasons. Contaminated or overlapping evaluation sets reward models that memorized similar questions. Prompt formatting, temperature, tool use, and whether chain-of-thought is allowed can swing results by large margins. Some vendors report self-reported scores; others are third-party replications. Multimodal and agentic tasks are still inconsistently measured.

When two sources disagree, prefer transparent methodology, look at the full score suite rather than a single headline number, and validate on your own prompts before locking in a production choice. See how we rank for the full methodology behind this catalog.

How often rankings update

Rankings and prices on What’s The Big Data are refreshed monthly. Each model entry carries a last-updated stamp; we re-pull API prices and public benchmark figures in the first week of each month. New model releases are added as soon as enough reliable data exists to place them fairly. If you spot a stale price or missing release, tell us — keeping the catalog current is the point of a living comparison page.

How we rank AI models

Transparent methodology for the intelligence index, speed figures, and API prices on this site — so you can trust the rankings and reproduce the logic. For a worked example, compare GPT-5.5 vs Claude 4 Opus using the same axes described below.

Intelligence score calculation

The Intelligence Index is a 0–100 composite that summarises how strong a language model is on widely cited public exams. We normalise each underlying benchmark to a 0–100 scale, then combine those scores into a single index so models can be scanned on one axis.

GPT-5.5 currently leads our intelligence rankings, while DeepSeek V3 offers one of the lowest API costs among strong open-weights options. Claude 4 Opus remains a top choice for writing tasks, and Gemini 2.5 Pro excels with long-context workflows. The index is a rough ordering for discovery — not a claim that one model is best at every task. Always open the model page for the full score breakdown.

Benchmark weighting

Language-model scores currently draw on MMLU, MMLU-Pro, GPQA, MATH, and HumanEval. Knowledge and reasoning suites (MMLU / MMLU-Pro / GPQA) carry the largest combined weight; MATH and HumanEval ensure hard quantitative and coding ability still move the index.

When a vendor or third-party report is missing a benchmark, we do not invent a perfect score — that entry is omitted from the composite or marked as partial data completeness so incomplete rows sort below fully measured models.

Speed testing

Speed is reported as median output tokens per second during steady-state generation across major API providers — not the time until the first token appears. Time to first token (TTFT) is listed separately because chat UIs feel slow when the first token lags, even if later tokens stream quickly.

Reasoning models that pause to “think” before answering will look slower on tokens-per-second; that is expected. Hosted open-weights endpoints can also vary by provider, so treat speed as a directional signal rather than a guaranteed SLA.

API price normalization

Token-priced models show list USD rates per million input tokens and per million output tokens from each vendor’s public API pricing. Blended price mixes those rates at a 3:1 input-to-output ratio — a common production traffic mix — so models with very different input vs output prices can be compared on one number.

Image, video, and speech models are priced in their native units (per image, per second of video, or per 1k characters). Enterprise discounts, batch APIs, and cached-input rates are excluded so comparisons stay apples-to-apples on public list pricing.

Update frequency

Rankings, prices, and benchmark figures are refreshed monthly. We re-pull public rates and scores in the first week of each month, and every model row carries a last-updated stamp. New releases are added once enough reliable data exists to place them fairly alongside peers.

Data sources

Primary inputs are vendor model cards and pricing pages, plus widely reported public benchmark results (including lab reports and independent leaderboards). Where sources disagree, we prefer transparent methodology and third-party replications over single self-reported headlines.

This catalog is editorially maintained — not an affiliate-sorted marketplace. If you spot a stale price, missing release, or incorrect score, tell us and we will correct it in the next refresh cycle.

AI Model Rankings Table

Click any column to sort. Filter by provider, license, capabilities, context, price, or release year.

Capabilities
GPT-5.5OpenAI8295 tok/s0.42s400k$5.00$15.00$7.50
Claude 4 OpusAnthropic8150 tok/s1.4s500k$8.00$40.00$16.00
Gemini 2.5 ProGoogle78110 tok/s0.7s2M$1.25$5.00$2.19
OpenAI o1OpenAI7632 tok/s12s200k$15.00$60.00$26.25
Claude 4 SonnetAnthropic7595 tok/s0.85s500k$3.00$15.00$6.00
Grok 3xAI7475 tok/s0.6s1M$3.00$15.00$6.00
DeepSeek R1DeepSeek7360 tok/s1.5s128k$0.55$2.19$0.96
GPT-4oOpenAI72110 tok/s0.4s128k$2.50$10.00$4.38
Claude 3.5 SonnetAnthropic7185 tok/s0.9s200k$3.00$15.00$6.00
OpenAI o3-miniOpenAI7060 tok/s6.5s200k$1.10$4.40$1.93
GPT-5.5 miniOpenAI68180 tok/s0.28s400k$0.25$1.00$0.44
DeepSeek V3DeepSeek6790 tok/s0.45s64k$0.27$1.10$0.48
Gemini 1.5 ProGoogle6760 tok/s0.9s2M$1.25$5.00$2.19
Llama 3.3 70B*Meta66200 tok/s0.4s128k$0.23$0.40$0.27
Gemini 2.0 FlashGoogle64220 tok/s0.3s1M$0.10$0.40$0.18
Llama 3.1 405B*Meta6432 tok/s0.7s128k$2.70$2.70$2.70
Qwen 2.5 72B*Alibaba6255 tok/s0.5s128k$0.40$0.40$0.40
Mistral Large 2Mistral6070 tok/s0.6s128k$2.00$6.00$3.00
Claude 3.5 HaikuAnthropic58130 tok/s0.5s200k$0.80$4.00$1.60
GPT-4o miniOpenAI56145 tok/s0.32s128k$0.15$0.60$0.26
Mistral Small 3Mistral52150 tok/s0.3s32k$0.20$0.60$0.30
Llama 3.1 8B*Meta44750 tok/s0.2s128k$0.06$0.06$0.06

Showing 22 of 22 models. Click any column header to sort. Prices are USD per 1M tokens unless noted otherwise. Estimates marked with *.

AI Model Timeline

A short history of landmark releases — how today’s frontier and open-weights models got here.

  1. 2023

    The modern frontier era begins — GPT-4 sets a new bar; open weights go mainstream.

  2. 2024

    Multimodal defaults, long context, and open-weights challengers arrive at scale.

  3. 2025

    Reasoning models and next-gen flagships tighten the race across labs.

  4. 2026

    Current frontier — higher reasoning ceilings, longer context, and sharper cost tiers.

AI Model Charts

Visual snapshots of intelligence, price, speed, context, and provider coverage — click any model to open its full page.

API Price Comparison

Blended $/1M tokens vs intelligence index

API price vs intelligence scatter plot85.2840.480000000000004$0.00$28.4Blended $/1M →Intelligence ↑GPT-5.5: $7.5 · intel 82GPT-5.5 mini: $0.44 · intel 68GPT-4o: $4.4 · intel 72GPT-4o mini: $0.26 · intel 56OpenAI o1: $26.3 · intel 76OpenAI o3-mini: $1.9 · intel 70Claude 4 Opus: $16.0 · intel 81Claude 4 Sonnet: $6.0 · intel 75Claude 3.5 Sonnet: $6.0 · intel 71Claude 3.5 Haiku: $1.6 · intel 58Gemini 2.5 Pro: $2.2 · intel 78Gemini 2.0 Flash: $0.18 · intel 64Gemini 1.5 Pro: $2.2 · intel 67Llama 3.3 70B: $0.27 · intel 66Llama 3.1 405B: $2.7 · intel 64Llama 3.1 8B: $0.06 · intel 44DeepSeek R1: $0.96 · intel 73DeepSeek V3: $0.48 · intel 67Mistral Large 2: $3.0 · intel 60Mistral Small 3: $0.30 · intel 52Grok 3: $6.0 · intel 74Qwen 2.5 72B: $0.40 · intel 62

Speed vs Cost

Bubble size = intelligence index

Speed versus cost bubble chart8250$0.00$28.9Blended $/1M →Speed (tok/s) ↑GPT-5.5: $7.5 · 95 tok/s · intel 82GPT-5.5 mini: $0.44 · 180 tok/s · intel 68GPT-4o: $4.4 · 110 tok/s · intel 72GPT-4o mini: $0.26 · 145 tok/s · intel 56OpenAI o1: $26.3 · 32 tok/s · intel 76OpenAI o3-mini: $1.9 · 60 tok/s · intel 70Claude 4 Opus: $16.0 · 50 tok/s · intel 81Claude 4 Sonnet: $6.0 · 95 tok/s · intel 75Claude 3.5 Sonnet: $6.0 · 85 tok/s · intel 71Claude 3.5 Haiku: $1.6 · 130 tok/s · intel 58Gemini 2.5 Pro: $2.2 · 110 tok/s · intel 78Gemini 2.0 Flash: $0.18 · 220 tok/s · intel 64Gemini 1.5 Pro: $2.2 · 60 tok/s · intel 67Llama 3.3 70B: $0.27 · 200 tok/s · intel 66Llama 3.1 405B: $2.7 · 32 tok/s · intel 64Llama 3.1 8B: $0.06 · 750 tok/s · intel 44DeepSeek R1: $0.96 · 60 tok/s · intel 73DeepSeek V3: $0.48 · 90 tok/s · intel 67Mistral Large 2: $3.0 · 70 tok/s · intel 60Mistral Small 3: $0.30 · 150 tok/s · intel 52Grok 3: $6.0 · 75 tok/s · intel 74Qwen 2.5 72B: $0.40 · 55 tok/s · intel 62

Provider Market Share

Share of models in this catalog by provider

Provider market share pie chartOpenAI: 11 models (31%)Anthropic: 4 models (11%)Google: 4 models (11%)Meta: 3 models (8%)DeepSeek: 2 models (6%)Mistral: 2 models (6%)xAI: 1 models (3%)Alibaba: 1 models (3%)Midjourney: 1 models (3%)Black Forest Labs: 1 models (3%)Stability AI: 1 models (3%)Runway: 1 models (3%)Kuaishou: 1 models (3%)ElevenLabs: 1 models (3%)Cartesia: 1 models (3%)Voyage AI: 1 models (3%)36models
  • OpenAI 11 · 31%
  • Anthropic 4 · 11%
  • Google 4 · 11%
  • Meta 3 · 8%
  • DeepSeek 2 · 6%
  • Mistral 2 · 6%
  • xAI 1 · 3%
  • Alibaba 1 · 3%
  • Midjourney 1 · 3%
  • Black Forest Labs 1 · 3%
  • Stability AI 1 · 3%
  • Runway 1 · 3%
  • Kuaishou 1 · 3%
  • ElevenLabs 1 · 3%
  • Cartesia 1 · 3%
  • Voyage AI 1 · 3%

Best AI Models for Coding

Picking a coding model is less about a single leaderboard number and more about the work you do every day: multi-file refactors, greenfield features, test generation, or hard algorithmic puzzles. On public coding exams such as HumanEval, Claude 4 Opus currently sits at the front of the pack, with GPT-5.5 close behind when the workflow includes tool use and agent loops. Those two are the defaults most engineering teams should evaluate first for production code assistance.

For repository-scale edits and long pull requests, Claude 4 Opus and Claude 4 Sonnet benefit from large context windows and strong instruction following. Sonnet is usually the better everyday workhorse — near-Opus quality at a friendlier price — while Opus is worth reserving for the toughest migrations. GPT-5.5 shines when the model must call linters, browsers, or internal APIs as part of an agent. OpenAI o1 and DeepSeek R1 remain the strongest picks when the task is pure reasoning: contest math, tricky algorithms, or proofs buried in a codebase.

Budget and volume matter too. GPT-5.5 mini, Gemini 2.0 Flash, and Claude 3.5 Haiku handle autocomplete-style prompts and cheap draft generations without burning the monthly bill. On the open-weights side, DeepSeek V3 and Llama 3.3 70B are credible self-hosted coding partners when data cannot leave your VPC. Browse our coding models hub, then open the GPT-5.5 vs Claude 4 Opus comparison before you standardize on a stack — and keep a small eval suite of real PRs so marketing benchmarks do not override what your engineers actually feel.

Related: Compare models · Rankings table · How we rank

Best AI Models for Writing

Long-form writing rewards coherence, taste, and the ability to hold a brief across thousands of tokens — not just raw benchmark scores. Claude 4 Opus and Claude 4 Sonnet are widely preferred for essays, docs, marketing drafts, and narrative work because they produce less repetitive prose and stay on voice longer than many peers. GPT-5.5 is often stronger when writing is tightly coupled to research, structured outlines, or multi-step briefs that mix analysis with drafting.

For editorial teams, Sonnet is usually the daily driver: fast enough for iteration, strong enough for publishable first drafts, and cheaper than Opus for high volume. Reach for Opus when the piece is flagship — a long report, a sensitive rewrite, or a document that must absorb a large source pack in one pass. Gemini 2.5 Pro earns a seat when the brief includes huge source libraries or mixed media context, thanks to its class-leading context window.

Lightweight models still have a role. GPT-4o mini and Gemini 2.0 Flash are fine for outlines, social variants, and SEO snippets where latency and cost matter more than literary polish. Open-weights options such as Llama 3.3 70B and Qwen 2.5 72B can power private writing assistants on your own hardware. Whatever you choose, keep a human editor in the loop for factual claims and brand voice — models draft; they do not own the byline. Explore individual model pages for strengths, weaknesses, and best-fit use cases before you lock a writing stack for the quarter.

Related: Compare models · Rankings table · How we rank

Best AI Models for Startups

Startups rarely need the absolute top intelligence index on day one. They need a model that ships features, keeps unit economics sane, and can be swapped later without rewriting the product. That usually means a mid-tier workhorse plus a premium model for the hard 10% of traffic. GPT-5.5 mini, Claude 4 Sonnet, Gemini 2.0 Flash, and DeepSeek V3 are the common shortlist: strong enough for chat, support, and light agents, cheap enough to iterate through product-market fit.

When quality becomes the bottleneck — complex agents, deep coding, or customer-facing research — graduate hot paths to GPT-5.5 or Claude 4 Opus rather than upgrading every call. A simple router (cheap model first, escalate on failure or low confidence) often cuts spend in half versus sending everything to the frontier. OpenAI o1 and DeepSeek R1 are worth a dedicated lane for planning and hard reasoning jobs that would otherwise burn scarce engineer time in a five-person team.

Self-hosting is attractive once volume is high and privacy requirements are strict. Llama 3.3 70B and DeepSeek V3 can run behind your own auth with predictable GPU cost, while proprietary APIs stay available as overflow. Track blended price, speed, and failure rates from week one so you know when to switch. Use the price and capability filters on the rankings table, then open head-to-head comparisons before you commit an annual contract or bake a single vendor into your architecture for the next eighteen months.

Related: Compare models · Rankings table · How we rank

Best Open Source AI Models

Open-weights models matter when you need to self-host, fine-tune, air-gap, or escape per-token API bills at scale. In 2026 the quality gap to proprietary frontier models has narrowed sharply on many tasks. DeepSeek R1 leads open reasoning, DeepSeek V3 is a strong generalist, and Llama 3.3 70B remains the most widely deployed community default. Llama 3.1 405B still shows up where raw parameter count and research baselines matter, while Qwen 2.5 72B and Mistral Small 3 cover multilingual and efficient serving niches.

“Open source” is not one license. Llama’s community license restricts very large commercial users; others are closer to Apache-style terms. Always read the license before you ship. Operationally, smaller weights (8B–70B) fit single-node or modest multi-GPU setups; 405B-class models need serious clusters or a hosted open-weights provider such as Together, Fireworks, or Groq. For image generation, Stable Diffusion 3.5 remains the open ecosystem to beat for fine-tunes and local creative tools.

When should you stay proprietary? If you need best-in-class tool use, the absolute top writing quality, or a vendor SLA with zero ops load, GPT-5.5, Claude 4 Opus, and Gemini 2.5 Pro still win many bake-offs. A hybrid pattern is common: open weights for bulk inference, closed APIs for the hardest prompts. Start from our open-weights and open-source hubs, sort the rankings table by intelligence and blended price, and validate on your own evaluation set before you migrate production traffic away from a managed API.

Related: Compare models · Rankings table · How we rank

Best AI Models by Price

API price is easy to misread if you only look at input rates. Real apps mix long prompts with shorter completions, which is why we publish a blended $/1M figure at a 3:1 input-to-output ratio. On that axis, Llama 3.1 8B is among the cheapest full entries in the catalog, while Gemini 2.0 Flash, GPT-4o mini, and Llama 3.3 70B deliver much stronger quality still under roughly thirty cents per million blended tokens. Those models are the right starting point for high-volume classification, extraction, and light chat.

Mid-tier options such as Claude 4 Sonnet, GPT-5.5 mini, and DeepSeek V3 occupy the value band where most products should live: good enough quality that users stay, cheap enough that margins survive. Frontier models — GPT-5.5, Claude 4 Opus, Gemini 2.5 Pro — cost several dollars per million tokens and should be reserved for prompts that actually need them. Reasoning models add another twist: they may look expensive per token and slower on tok/s because they spend hidden compute thinking, yet they can still be cheaper than a junior engineer hour on hard tasks.

Self-hosting flips the spreadsheet. Open-weights models have near-zero marginal API cost once GPUs are paid for, but you inherit utilization, reliability, and MLOps work. Compare list prices in our rankings table, use the price filters to shortlist under $0.50 or $2 blended, then open model pages for input versus output breakdowns. Pair any cost decision with quality checks — the cheapest model that fails your eval is the most expensive mistake. For a visual view, jump to the API price and speed-versus-cost charts above.

Related: Compare models · Rankings table · How we rank

Browse AI Models by category

Drill into a slice of the catalog — reasoning, coding, image, video, speech, embeddings, agents, open weights, and more.

By Provider

Frequently asked questions

In 2026 the three highest-scoring models on our intelligence index are GPT-5.5 (82), Claude 4 Opus (81) and Gemini 2.5 Pro (78). The “best” depends on the task — GPT-5.5 leads on reasoning and tool use, Claude 4 Opus on long-context code, and Gemini 2.5 Pro on multimodal work with its 2M-token window.

Pick a model with confidence

Open any model below for a deep-dive — or pit two models head-to-head with the side-by-side comparator.