When was DeepSeek R1 released?

DeepSeek R1 was released by DeepSeek on 20 January 2025.

How much does DeepSeek R1 cost?

DeepSeek R1 costs $0.55 per 1 million input tokens and $2.19 per 1 million output tokens via DeepSeek’s API. At a typical 3:1 input/output ratio the blended cost works out to about $0.96 per 1M tokens.

How smart is DeepSeek R1?

DeepSeek R1 scores 73 out of 100 on our intelligence index — a strong production-tier composite of MMLU, MMLU Pro, GPQA, MATH and HumanEval benchmark scores. That places it #7 of 22 language models we track.

What is the context window of DeepSeek R1?

DeepSeek R1 has a 128k-token context window — roughly 96,000 words of text. That’s what you can fit into a single request (your prompt plus all uploaded documents and the model’s response combined).

What is DeepSeek R1 best for?

DeepSeek R1 is most useful for self-hosted reasoning, math & code and cost-sensitive agents. Its key strengths are reasoning at gpt-class scores, open weights and cheap.

How fast is DeepSeek R1?

DeepSeek R1 generates around 60 output tokens per second with a typical time-to-first-token of 1.5 seconds on the major API providers. Reasoning models that "think" before answering will appear slower on tokens-per-second since they spend time on internal chain-of-thought.

Is DeepSeek R1 better than DeepSeek V3?

DeepSeek R1 scores higher on our intelligence index (73 vs 67). On most general tasks DeepSeek R1 has the edge — but DeepSeek V3 can still be the better pick on cost or specific capabilities. See our full DeepSeek R1 vs DeepSeek V3 comparison at /ai-models/compare/deepseek-r1-vs-deepseek-v3/.

What is the best alternative to DeepSeek R1?

The closest alternatives to DeepSeek R1 are DeepSeek V3, GPT-5.5 and Claude 4 Opus. Each shares most of DeepSeek R1’s use-cases — pick by price, context window or specific capability rather than headline intelligence.

Is DeepSeek R1 open source?

Yes. DeepSeek R1 is open-source — the weights are publicly available, and you can self-host it or use a hosted provider (Together, Fireworks, Groq, Replicate). Some open-source licenses include usage caveats; check the model’s license file before deploying.

Can I use DeepSeek R1 for free?

Yes. DeepSeek R1 is open-source — you can download the weights and self-host for free. DeepSeek’s hosted chat at chat.deepseek.com also offers free access with rate limits; the API itself is paid but very cheap relative to comparable models.

DeepSeek Open sourceJan 2025

DeepSeek R1

Open-weights reasoning model that matches o1 at 1/25 the price.

Compare with…Open docs

Intelligence index

73/ 100

vs all models68th pctile

Composite of MMLU, GPQA, MATH & HumanEval

Speed

60tok/s

vs all models18th pctile

Median across providers, steady state

Blended price

$0.96/ 1M tokens

vs all models64th pctile

3:1 input:output blend

DeepSeek R1 Overview at a Glance

DeepSeek R1 is a large language model from DeepSeek, first released on 20 January 2025. It is open-source (open-weights) and sits in the reasoning, frontier, deepseek, open weights, and agents categories of our catalog. Open-weights reasoning model that matches o1 at 1/25 the price. This page covers DeepSeek R1 pricing, benchmarks, API limits, speed, modalities, best use cases, and how it compares with similar models — so you can decide whether it belongs in your stack in 2026.

As a language model, DeepSeek R1 is evaluated on reasoning quality, coding ability, latency, context window size, and dollars-per-million-tokens. The context window is 128k tokens (about 96k words), which determines how much prompt, document, and conversation history you can send in one request. At a typical 3:1 input-to-output mix, the blended API price is about $0.96 per 1M tokens. On our intelligence index it ranks #7 of 22 language models we track with a score of 73/100 (strong production-tier).

Teams usually shortlist DeepSeek R1 when they need a dependable DeepSeek option for production chat, agents, retrieval-augmented generation, or coding copilots. Common fits include self-hosted reasoning, math & code, and cost-sensitive agents. Reviewers consistently call out reasoning at gpt-class scores, open weights, and cheap as standout strengths. Trade-offs to weigh include slower than non-reasoning peers. The sections below break down pricing tables, benchmark charts, token limits, input/output modalities, and head-to-head comparisons so long-tail queries — from “DeepSeek R1 API pricing” to “DeepSeek R1 vs DeepSeek V3” — are answered on this page.

If you are migrating from an older DeepSeek model or switching labs entirely, treat this page as a decision brief: skim the overview stats, confirm API pricing fits your volume, check whether the context window covers your longest documents, then validate quality on a golden set of prompts. Benchmarks and charts help shortlist; your own evals decide. We refresh catalog numbers periodically (last update 2026-06) so figures stay useful through the year.

Context window: 128k tokens
Max output: 33k tokens
Input price: $0.55 / 1M tokens
Output price: $2.19 / 1M tokens
Time to first token: 1.5s
Input modalities: text
Output modalities: text
License: Open source
Provider: DeepSeek

Strengths

Reasoning at GPT-class scores
Open weights
Cheap

Weaknesses

Slower than non-reasoning peers

Best for

Self-hosted reasoning
Math & code
Cost-sensitive agents

DeepSeek R1 Pricing

DeepSeek R1 uses token-based API pricing from DeepSeek. You pay $0.55 per million input tokens and $2.19 per million output tokens. For planning budgets we quote a blended rate of $0.96 per 1M tokens at a 3:1 input-to-output ratio — the same convention used across our catalog so models are comparable. Output tokens usually dominate cost for chatty or agentic workloads, so watch generation length and system-prompt size.

When estimating production spend, multiply expected monthly tokens by the blended rate, then add a buffer for retries, tool-calling loops, and RAG context. Because DeepSeek R1 is open-weights, self-hosting can beat API pricing at high volume once GPU utilization is solid — hosted APIs remain cheaper to start.

Also compare DeepSeek R1 against cheaper siblings from DeepSeek for router patterns: send easy traffic to a mini/flash tier and reserve DeepSeek R1 for hard reasoning. That hybrid design often cuts billable tokens 30–70% without users noticing quality drops on simple turns.

Input price: $0.55 / 1M tokens
Output price: $2.19 / 1M tokens
Blended (3:1): $0.96 / 1M tokens

DeepSeek R1 Benchmarks

Public benchmark scores help compare DeepSeek R1 with other LLMs on knowledge, graduate-level science, competition math, and coding. Reported figures in our catalog include MMLU 87.1, MMLU Pro 75.9, GPQA 71.5, MATH 90.2, and HumanEval 91. These are not a substitute for evals on your own prompts, but they are useful for shortlisting.

Our intelligence index (73/100) normalizes those benchmarks into a single strong production-tier score so you can scan the leaderboard quickly. The performance chart below shows each benchmark against the current catalog leader.

MMLU

General knowledge across 57 subjects

87.1

leader: 91.8

MMLU Pro

Harder MMLU successor with more reasoning

75.9

leader: 80.0

GPQA

Graduate-level science Q&A

71.5

leader: 78.0

MATH

Competition mathematics

90.2

leader: 94.8

HumanEval

Python code generation pass@1

91.0

leader: 95.8

DeepSeek R1 API Pricing

API pricing for DeepSeek R1 is what you pay when calling DeepSeek’s developer endpoint (or a marketplace such as Azure, Bedrock, or Vertex when available). Unlike consumer chat apps with flat subscriptions, API bills scale with tokens processed. Cache prompt prefixes where the provider supports it, batch non-interactive jobs, and prefer smaller sibling models for classification or routing when full DeepSeek R1 quality is unnecessary.

To convert catalog numbers into a monthly forecast: estimate average input tokens per request (system prompt + user message + retrieved context), average output tokens, and request volume. Cost ≈ requests × ((inputTokens/1e6) × inputPrice + (outputTokens/1e6) × outputPrice). Our LLM pricing calculator can stress-test scenarios if you need a second opinion against peers.

Always verify live rates on the official docs — our figures are refreshed periodically (last catalog update: 2026-06) and providers change list prices. Official reference: https://api-docs.deepseek.com/.

DeepSeek R1 Context Window

DeepSeek R1 offers a 128k-token context window — roughly about 96k words of English text. Everything in a single API call counts against that budget: system instructions, chat history, retrieved documents, tool schemas, and the model’s reply. Exceeding the window truncates or errors depending on the provider.

Large windows help with long PDFs, multi-file code reviews, and multi-hour agent traces, but bigger contexts also cost more tokens and can add latency. Prefer retrieval that stuffs only relevant chunks, summarize old turns, and reserve headroom for up to 33k output tokens.

DeepSeek R1 Input / Output Modalities

DeepSeek R1 accepts text as input and produces text as output. Knowing the modality matrix matters when you design pipelines — for example, vision-capable language models can take screenshots or PDFs as images, while pure text models need an OCR or captioning step first.

If you need bidirectional voice, native video understanding, or tool-use with multimodal arguments, confirm support in DeepSeek’s API schema rather than assuming parity with the consumer chat app. Modality support also affects pricing: image or audio inputs may be tokenized differently than plain text.

Document which of DeepSeek R1’s listed modalities you will actually send in production. Turning on unused multimodal features can change tokenizers, rate limits, and safety filters unexpectedly.

Inputs: text
Outputs: text

DeepSeek R1 Token Limits

Token limits define how much DeepSeek R1 can read and write per request. Total context is capped at 128k tokens. Maximum completion length is 33k tokens — even if context remains, the model stops generating beyond that ceiling unless you continue in a follow-up call. Providers may also enforce organization-level rate limits (RPM/TPM) separate from these per-request caps.

Practical tip: set max_tokens intentionally. Leaving it unbounded wastes budget on verbose answers; setting it too low truncates JSON or code. For structured outputs, prefer schemas/tool calls and keep completions tight.

Context window: 128k tokens
Max output: 33k tokens

DeepSeek R1 Speed

Speed for DeepSeek R1 is measured two ways: time-to-first-token (how quickly streaming starts) and steady-state tokens per second. Catalog median throughput is about 60 tok/s. Typical TTFT is 1.5 s. Reasoning-heavy modes that think before answering will look slower on tok/s even when quality is higher.

Interactive chat wants low TTFT; batch extraction can tolerate higher latency for cheaper regions or providers. If DeepSeek R1 is too slow for your UX, evaluate a “mini/flash/haiku” sibling from the same lab before switching ecosystems.

Throughput: 60 tokens/sec
Time to first token: 1.5 s
Speed percentile: Faster than ~18% of tracked LLMs

DeepSeek R1 Performance Charts

The charts on this page visualize DeepSeek R1 against catalog peers. Benchmark bars show academic scores versus the current leader; the similar-models comparison table plots intelligence, speed, and blended price so you can see trade-offs at a glance. Use them to answer “is DeepSeek R1 fast enough?” and “is the quality jump worth the premium?” without opening a spreadsheet.

Benchmark performance vs catalog leaders

MMLU

General knowledge across 57 subjects

87.1

leader: 91.8

MMLU Pro

Harder MMLU successor with more reasoning

75.9

leader: 80.0

GPQA

Graduate-level science Q&A

71.5

leader: 78.0

MATH

Competition mathematics

90.2

leader: 94.8

HumanEval

Python code generation pass@1

91.0

leader: 95.8

Intelligence index vs similar models

DeepSeek R173

DeepSeek V367

GPT-5.582

Claude 4 Opus81

Comparison with Similar Models

Choosing an AI model is rarely absolute — it is relative to the next-best option. DeepSeek R1 is most often weighed against DeepSeek V3, GPT-5.5, and Claude 4 Opus. Compare intelligence (or generation quality), latency, price, license, and modality support. A slightly weaker but much cheaper model can win for high-volume workloads; a pricier frontier model wins when a single mistake is expensive.

Use the links and table below for structured DeepSeek R1 vs alternatives research. We also maintain dedicated head-to-head pages for popular matchups when available. If you are standardizing on DeepSeek, check sibling models from the same lab before leaving the ecosystem.

A practical bake-off: pick 20–50 real prompts, score accuracy/style, measure p50/p95 latency, and compute cost at projected volume for DeepSeek R1 and two peers. Ship the winner behind a feature flag so you can reverse the decision without a rewrite.

Model	Provider	Intelligence	Speed	Price
DeepSeek R1	DeepSeek	73	60 t/s	$0.96/1M
DeepSeek V3	DeepSeek	67	90 t/s	$0.48/1M
GPT-5.5	OpenAI	82	95 t/s	$7.50/1M
Claude 4 Opus	Anthropic	81	50 t/s	$16.00/1M

DeepSeek R1 vs popular alternatives

DeepSeek R1 Best Use Cases

Best use cases for DeepSeek R1 follow from its strengths, price point, and modality support. Match the model to the job: frontier reasoning for hard planning, fast/cheap tiers for classification, image/video/speech specialists for media pipelines.

Based on catalog notes, DeepSeek R1 is a particularly strong fit for self-hosted reasoning, math & code, and cost-sensitive agents. Validate with a short bake-off on your real prompts before a full cutover.

Anti-patterns: do not use a frontier-priced model like a generic classifier if a smaller model scores within a point on your eval; do not stuff entire corpora into context when retrieval would be cheaper; and do not skip structured outputs if you plan to parse DeepSeek R1 responses in code.

Self-hosted reasoning
Math & code
Cost-sensitive agents

DeepSeek R1 Pros & Cons

Every model trades quality, speed, cost, and openness. Here is a concise pros and cons list for DeepSeek R1 drawn from our catalog strengths and weaknesses — pair it with your own evals before committing.

Read pros as “reasons to shortlist” and cons as “risks to mitigate,” not as deal-breakers in isolation. A listed weakness (for example higher price or smaller context) may be irrelevant if your workload is bursty, short-context, or already standardized on DeepSeek.

After scanning this list, jump to the comparison table and FAQ for decision support, then lock a trial window with success metrics before replacing a production model with DeepSeek R1.

Pros

Reasoning at GPT-class scores
Open weights
Cheap

Cons

Slower than non-reasoning peers

DeepSeek R1 — frequently asked questions

DeepSeek R1 is a large language model from DeepSeek, released on 20 January 2025. Open-weights reasoning model that matches o1 at 1/25 the price.

Need help choosing between models?

Compare every option in one sortable table — intelligence, speed and price on a single page.

All AI models LLM pricing calculator