DeepSeek has launched its V4 Pro 0813, a large-scale mixture-of-experts AI model, offering competitive performance metrics and pricing.
DeepSeek: DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Providers
This model is hosted by one provider. OpenRouter forwards every request to it directly — no routing decisions to make.
| Provider | Input /M | Output /M | Cache read /M | Latency | Throughput | Uptime |
|---|
DeepSeek | $0.435 | $0.87 | $0.003625 | 1.96s | 45 tps | 100.00% |
Pricing
The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.
Weighted Average
Weighted Avg Input Price
$0.03856
/M tokens
Weighted Avg Output Price
$0.8696
/M tokens
DeepSeek V4 Pro 0813 — Price History
| Chart visibility | Provider | Effective in /M | Effective out /M | Listed in /M | Listed out /M | Cache hit rate | Token share 1d |
|---|
| DeepSeek | $0.03856 | $0.8696 | $0.435 | $0.87 | 91.9% | 100.0% |
Performance
Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).
Throughput
45tok/s
P50, best across providers
Latency
1.96s
P50, best provider
All locations
3 days
Throughput
P99
Avg100 tok/s
P95
Avg83 tok/s
P90
Avg77 tok/s
P75
Avg68 tok/s
P50
Avg56 tok/s
Latency
P50
Avg1.44 s
P75
Avg1.91 s
P90
Avg2.71 s
P95
Avg3.86 s
P99
Avg21.94 s
E2E Latency
P50
Avg10.33 s
P75
Avg27.14 s
P90
Avg59.35 s
P95
Avg93.77 s
P99
Avg237.31 s
Tool Call Error Rate
DeepSeek
Avg1.67 %
Cache Hit Rate
DeepSeek
Avg92.14 %
Uptime
Percent of requests that succeeded over the last 30 days. OpenRouter monitors every provider continuously and automatically retries on the next-best provider when one returns an error.
Avg. Provider Uptime (3d)
100.00%
averaged across all endpoints
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
Benchmarks
Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.
45.3\n\nArtificial Analysis\n\nIntelligence Index\n\nBetter than 70% of models compared 59.4\n\nArtificial Analysis\n\nCoding Index\n\nBetter than 73% of models compared 37.8\n\nArtificial Analysis\n\nAgentic Index\n\nBetter than 73% of models compared
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
88.8%
HLE
Humanity's Last Exam
37.5%
IFBench
Instruction-following benchmark
76.5%
τ²-Bench Telecom
Conversational AI agents in dual-control scenarios
96.2%
AA-LCR
Long context reasoning evaluation
70.0%
GDPval-AA
Economically valuable tasks
40.3%
CritPt
Research-level physics reasoning
12.9%
Coding
SciCode
Python programming for scientific computing
50.0%
Terminal-Bench Hard
Agentic coding & terminal use
46.2%
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
42.9%
AA-Omniscience Non-Hallucination Rate
Rate of avoiding hallucination among non-correct responses
5.9%
Metrics sourced from Artificial Analysis
Apps
Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Hermes Agent
Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.
2.67Btokens

omp
new
1.92Btokens

pi
There are many coding agents, but this one is yours.
1.56Btokens

Claude Code
Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and iterates on failures, all from natural language prompts.
866Mtokens

ISEKAI ZERO
AI adventures. Travel with your favorite characters
712Mtokens
Activity
Token volume and request traffic to this model over time.
Tokens
Prompt
6.04B
Reasoning
55.6M
Completion
24.8M
Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length.
Quick Start
Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.
1
Get your API key
Create an API key from your OpenRouter dashboard and set it as an environment variable:
Create API Key
Copy
shell
export OPENROUTER_API_KEY=sk-or-v1-...
2
Make your first request
Use deepseek/deepseek-v4-pro-0813 with the OpenRouter API:
OpenRouter supports reasoning-enabled models that can show their step-by-step thinking process. Use the reasoning parameter in your request to enable reasoning, and access the reasoning_details array in the response to see the model's internal reasoning before the final answer. When continuing a conversation, preserve the complete reasoning_details when passing messages back to the model so it can continue reasoning from where it left off. Learn more about reasoning tokens.
In the examples below, the OpenRouter-specific headers are optional. Setting them allows your app to appear on the OpenRouter leaderboards.
TypeScript SDKPythonTypeScript (fetch)cURLPython (OpenAI)TypeScript (OpenAI)
Copy
typescript
import { OpenRouter } from "@openrouter/sdk";
const openrouter = new OpenRouter({
apiKey: ""
});
// Stream the response to get reasoning tokens in usage
const stream = await openrouter.chat.send({
chatRequest: {
model: "deepseek/deepseek-v4-pro-0813",
messages: [
{
role: "user",
content: "How many r's are in the word 'strawberry'?"
}
],
stream: true
}
});
let response = "";
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) {
response += content;
process.stdout.write(content);
}
// Usage information comes in the final chunk
if (chunk.usage) {
console.log("\nReasoning tokens:", chunk.usage.completionTokensDetails?.reasoningTokens);
}
}
Using third-party SDKs
For information about using third-party SDKs and frameworks with OpenRouter, please see our frameworks documentation.
3
Enable streaming
Add "stream": true to your request body to receive responses as server-sent events:
Copy
shell
curl -N https://openrouter.ai/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-d
Endpoint
Sends a request for a model response for the given chat conversation. Supports both streaming and non-streaming modes.
POSThttps://openrouter.ai/api/v1/chat/completions
AuthorizationBearer $OPENROUTER_API_KEY
Content-Typeapplication/json
HTTP-Refereroptional — your site URL, for rankings
X-Titleoptional — your site name, for rankings
Modeldeepseek/deepseek-v4-pro-0813
Creates a streaming or non-streaming response using the OpenAI Responses API format.
Docs
POSThttps://openrouter.ai/api/v1/responses
AuthorizationBearer $OPENROUTER_API_KEY
Content-Typeapplication/json
HTTP-Refereroptional — your site URL, for rankings
X-Titleoptional — your site name, for rankings
Modeldeepseek/deepseek-v4-pro-0813
Creates a message using the Anthropic Messages API format. Supports text, images, PDFs, tools, and extended thinking.
Docs
POSThttps://openrouter.ai/api/v1/messages
AuthorizationBearer $OPENROUTER_API_KEY
Content-Typeapplication/json
HTTP-Refereroptional — your site URL, for rankings
X-Titleoptional — your site name, for rankings
Modeldeepseek/deepseek-v4-pro-0813
Parameters
| Name | Type | Default | Description |
|---|
reasoning | map | — | Controls reasoning behavior for models that support thinking tokens, including whether reasoning is enabled, the reasoning effort, maximum reasoning tokens, and whether reasoning is excluded from the response. |
max_tokens | integer | — | This sets the upper limit for the number of tokens the model can generate in response. |
temperature | float | 1 | This setting influences the variety in the model's responses. |
top_p | float | 1 | This setting limits the model's choices to a percentage of likely tokens: only the top tokens whose probabilities add up to P. |
stop | array | — | Stop generation immediately if the model encounter any token specified in the stop array. |
frequency_penalty | float | 0 | This setting aims to control the repetition of tokens based on how often they appear in the input. |
presence_penalty | float | 0 | Adjusts how often the model repeats specific tokens already used in the input. |
logprobs | boolean | — | Whether to return log probabilities of the output tokens or not. |
top_logprobs | integer | — | An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with an associated log probability. |
tools | array | — | Tool calling parameter, following OpenAI's tool calling request shape. |
tool_choice | string or object | — | Controls which (if any) tool is called by the model. |
response_format | map | — | Forces the model to produce specific output format. |
Frequently asked questions
What is DeepSeek V4 Pro 0813?
How much does DeepSeek V4 Pro 0813 cost?
What is the context length of DeepSeek V4 Pro 0813?
Does DeepSeek V4 Pro 0813 support tool calling and structured outputs?
What other text models does DeepSeek have?
When was DeepSeek V4 Pro 0813 released?
More models from DeepSeek
DeepSeek V4 Flash Latest\n\nThis model always redirects to the latest model in the DeepSeek V4 Flash family.
DeepSeek V4 Flash 0731\n\nDeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.
DeepSeek V4 Pro\n\nDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.\n\nBuilt on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical
DeepSeek V4 Flash 0423\n\nDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.\n\nThe model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.
DeepSeek V3.2 Speciale\n\nDeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on DeepSeek Sparse Attention (DSA) for efficient long-context processing, then scales post-training reinforcement learning to push capability beyond the base model. Reported evaluations place Speciale ahead of GPT-5 on difficult reasoning workloads, with proficiency comparable to Gemini-3.0-Pro, while retaining strong coding and tool-use reliability. Like V3.2, it benefits from a large-scale agentic task synthesis pipeline that improves compliance and generalization in interactive environments.
DeepSeek V3.2\n\nDeepSeek V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.\n\nUsers can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs
DeepSeek V3.2 Exp\n\nDeepSeek V3.2 Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios while maintaining output quality. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model was trained under conditions aligned with V3.1-Terminus to enable direct comparison. Benchmarking shows performance roughly on par with V3.1 across reasoning, coding, and agentic tool-use tasks, with minor tradeoffs and gains depending on the domain. This release focuses on validating architectural optimizations for extended context lengths rather than advancing raw task accuracy, making it primarily a research-oriented model for exploring efficient transformer designs.
DeepSeek V3.1 Terminus\n\nDeepSeek V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's performance in coding and search agents. It is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.
DeepSeek V3.1\n\nDeepSeek V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.\n\nIt succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.
DeepSeek V3.1 Base\n\nThis is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).\n\nDeepSeek V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.
R1 Distill Qwen 7B\n\nDeepSeek R1 Distill Qwen 7B is a 7 billion parameter dense language model distilled from DeepSeek R1, leveraging reinforcement learning-enhanced reasoning data generated by DeepSeek's larger models. The distillation process transfers advanced reasoning, math, and code capabilities into a smaller, more efficient model architecture based on Qwen2.5-Math-7B. This model demonstrates strong performance across mathematical benchmarks (92.8% pass@1 on MATH-500), coding tasks (Codeforces rating 1189), and general reasoning (49.1% pass@1 on GPQA Diamond), achieving competitive accuracy relative to larger models while maintaining smaller inference costs.
DeepSeek R1 0528 Qwen3 8B\n\nDeepSeek R1 0528 is a lightly upgraded release of DeepSeek R1 that taps more compute and smarter post-training tricks, pushing its reasoning and inference to the brink of flagship models like O3 and Gemini 2.5 Pro.\nIt now tops math, programming, and logic leaderboards, showcasing a step-change in depth-of-thought.\nThe distilled variant, DeepSeek R1 0528 Qwen3 8B, transfers this chain-of-thought into an 8 B-parameter form, beating standard Qwen3 8B by +10 pp and tying the 235 B “thinking” giant on AIME 2024.
R1 0528\n\nMay 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.\n\nFully open-source model.
DeepSeek Prover V2\n\nDeepSeek Prover V2 is a 671B parameter model, speculated to be geared towards logic and mathematics. Likely an upgrade from DeepSeek-Prover-V1.5 Not much is known about the model yet, as DeepSeek released it on Hugging Face without an announcement or description.
DeepSeek V3 Base\n\nNote that this is a base model mostly meant for testing, you need to provide detailed prompts for the model to return useful responses.\n\nDeepSeek V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.\n\nDeepSeek V3 Base is the pre-trained model behind DeepSeek V3
DeepSeek V3 0324\n\nDeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.\n\nIt succeeds the DeepSeek V3 model and performs really well on a variety of tasks.
DeepSeek R1 Zero\n\nDeepSeek R1 Zero is a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step. It's 671B parameters in size, with 37B active in an inference pass.\n\nIt demonstrates remarkable performance on reasoning. With RL, DeepSeek R1 Zero naturally emerged with numerous powerful and interesting reasoning behaviors.\n\nDeepSeek R1 Zero encounters challenges such as endless repetition, poor readability, and language mixing. See DeepSeek R1 for the SFT model.
R1 Distill Llama 8B\n\nDeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.\n\nHugging Face:
R1 Distill Qwen 1.5B\n\nDeepSeek R1 Distill Qwen 1.5B is a distilled large language model based on Qwen 2.5 Math 1.5B, using outputs from DeepSeek R1. It's a very small and efficient model which outperforms GPT 4o 0513 on Math Benchmarks.\n\nOther benchmark results include:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
R1 Distill Qwen 32B\n\nDeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\n- AIME 2024 pass@1: 72.6\n- MATH-500 pass@1: 94.3\n- CodeForces Rating: 1691\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
R1 Distill Qwen 14B\n\nDeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
R1 Distill Llama 70B\n\nDeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
R1\n\nDeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.\n\nFully open-source model & technical report.\n\nMIT licensed: Distill & commercialize freely!
DeepSeek V2.5\n\nDeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct. The new model integrates the general and coding abilities of the two previous versions. For model details, please visit DeepSeek-V2 page for more information.
Previous slideNext slide
Compare Playground
Verify your email
We could not find a primary email address on your account. Please contact support.
Cancel
Close
Quick Start
Drop-in code to call this model with OpenRouter's OpenAI-compatible API.
Get Code
OpenRouter SDKOpenAI SDKAnthropic SDKRaw
TypeScriptPythonGo
typescript
import { OpenRouter } from "@openrouter/sdk";
const openrouter = new OpenRouter({
apiKey: ""
});
const response = await openrouter.chat.send({
model: "deepseek/deepseek-v4-pro-0813",
messages: [
{
"role": "user",
"content": "What is the meaning of life?"
}
]
});
console.log(response.choices[0].message.content);
Close
Get Code
OpenRouter SDKOpenAI SDKAnthropic SDKRaw
TypeScriptPythonGo
typescript
import { OpenRouter } from "@openrouter/sdk";
const openrouter = new OpenRouter({
apiKey: ""
});
const response = await openrouter.chat.send({
model: "deepseek/deepseek-v4-pro-0813",
messages: [
{
"role": "user",
"content": "What is the meaning of life?"
}
]
});
console.log(response.choices[0].message.content);
Close
0 / 1
Latency percentiles on OpenRouter All locations
Close
DeepSeek V4 Pro 0813 — Price History
Close
Throughput percentiles on OpenRouter All locations
Close
End-to-End Latency percentiles on OpenRouter All locations
Close
Cache Hit Rate by Provider
Close
Tool Call Error Rate by Provider
Close
PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.