SATURDAY, OCTOBER 10, 2026|No. 18173
Artificial Intelligence · Models

DeepSeek Releases V4 Pro 0813 Large Language Model

DeepSeek has launched its V4 Pro 0813, a large-scale mixture-of-experts AI model, offering competitive performance metrics and pricing.

The DeepSeek V4 Pro 0813 model is a new large-scale AI offering.
The DeepSeek V4 Pro 0813 model is a new large-scale AI offering. · Photo by Growtika on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

DeepSeek: DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Providers

This model is hosted by one provider. OpenRouter forwards every request to it directly — no routing decisions to make.

ProviderInput /MOutput /MCache read /MLatencyThroughputUptime
Favicon for DeepSeekDeepSeek$0.435$0.87$0.0036251.96s45 tps100.00%

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Weighted Average

Weighted Avg Input Price

$0.03856

/M tokens

Weighted Avg Output Price

$0.8696

/M tokens

DeepSeek V4 Pro 0813 — Price History

Chart visibilityProviderEffective in /MEffective out /MListed in /MListed out /MCache hit rateToken share 1d
Favicon for DeepSeekDeepSeek$0.03856$0.8696$0.435$0.8791.9%100.0%

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Throughput

45tok/s

P50, best across providers

Latency

1.96s

P50, best provider

All locations

3 days

Throughput

P99

Avg100 tok/s

P95

Avg83 tok/s

P90

Avg77 tok/s

P75

Avg68 tok/s

P50

Avg56 tok/s

Latency

P50

Avg1.44 s

P75

Avg1.91 s

P90

Avg2.71 s

P95

Avg3.86 s

P99

Avg21.94 s

E2E Latency

P50

Avg10.33 s

P75

Avg27.14 s

P90

Avg59.35 s

P95

Avg93.77 s

P99

Avg237.31 s

Tool Call Error Rate

DeepSeek

Avg1.67 %

Cache Hit Rate

DeepSeek

Avg92.14 %

Uptime

Percent of requests that succeeded over the last 30 days. OpenRouter monitors every provider continuously and automatically retries on the next-best provider when one returns an error.

Avg. Provider Uptime (3d)

100.00%

averaged across all endpoints

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

45.3\n\nArtificial Analysis\n\nIntelligence Index\n\nBetter than 70% of models compared 59.4\n\nArtificial Analysis\n\nCoding Index\n\nBetter than 73% of models compared 37.8\n\nArtificial Analysis\n\nAgentic Index\n\nBetter than 73% of models compared

Reasoning

GPQA Diamond

Graduate-level scientific reasoning

88.8%

HLE

Humanity's Last Exam

37.5%

IFBench

Instruction-following benchmark

76.5%

τ²-Bench Telecom

Conversational AI agents in dual-control scenarios

96.2%

AA-LCR

Long context reasoning evaluation

70.0%

GDPval-AA

Economically valuable tasks

40.3%

CritPt

Research-level physics reasoning

12.9%

Coding

SciCode

Python programming for scientific computing

50.0%

Terminal-Bench Hard

Agentic coding & terminal use

46.2%

Knowledge

AA-Omniscience Accuracy

Proportion of correctly answered questions

42.9%

AA-Omniscience Non-Hallucination Rate

Rate of avoiding hallucination among non-correct responses

5.9%

Metrics sourced from Artificial Analysis

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Favicon for https://nousresearch.com

Hermes Agent

Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.

2.67Btokens

Favicon for https://omp.sh/

omp

new

1.92Btokens

Favicon for https://pi.dev/

pi

There are many coding agents, but this one is yours.

1.56Btokens

Favicon for https://claude.ai/apple-touch-icon.png

Claude Code

Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and iterates on failures, all from natural language prompts.

866Mtokens

Favicon for https://www.isekai.world/

ISEKAI ZERO

AI adventures. Travel with your favorite characters

712Mtokens

Activity

Token volume and request traffic to this model over time.

Tokens

Prompt

6.04B

Reasoning

55.6M

Completion

24.8M

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

1

Get your API key

Create an API key from your OpenRouter dashboard and set it as an environment variable:

Create API Key

Copy

shell

export OPENROUTER_API_KEY=sk-or-v1-...

2

Make your first request

Use deepseek/deepseek-v4-pro-0813 with the OpenRouter API:

OpenRouter supports reasoning-enabled models that can show their step-by-step thinking process. Use the reasoning parameter in your request to enable reasoning, and access the reasoning_details array in the response to see the model's internal reasoning before the final answer. When continuing a conversation, preserve the complete reasoning_details when passing messages back to the model so it can continue reasoning from where it left off. Learn more about reasoning tokens.

In the examples below, the OpenRouter-specific headers are optional. Setting them allows your app to appear on the OpenRouter leaderboards.

TypeScript SDKPythonTypeScript (fetch)cURLPython (OpenAI)TypeScript (OpenAI)

Copy

typescript

import { OpenRouter } from "@openrouter/sdk";

const openrouter = new OpenRouter({
 apiKey: ""
});

// Stream the response to get reasoning tokens in usage
const stream = await openrouter.chat.send({
 chatRequest: {
 model: "deepseek/deepseek-v4-pro-0813",
 messages: [
 {
 role: "user",
 content: "How many r's are in the word 'strawberry'?"
 }
 ],
 stream: true
 }
});

let response = "";
for await (const chunk of stream) {
 const content = chunk.choices[0]?.delta?.content;
 if (content) {
 response += content;
 process.stdout.write(content);
 }

 // Usage information comes in the final chunk
 if (chunk.usage) {
 console.log("\nReasoning tokens:", chunk.usage.completionTokensDetails?.reasoningTokens);
 }
}

Using third-party SDKs

For information about using third-party SDKs and frameworks with OpenRouter, please see our frameworks documentation.

3

Enable streaming

Add "stream": true to your request body to receive responses as server-sent events:

Copy

shell

curl -N https://openrouter.ai/api/v1/chat/completions \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $OPENROUTER_API_KEY" \
 -d 

Endpoint

Sends a request for a model response for the given chat conversation. Supports both streaming and non-streaming modes.

POSThttps://openrouter.ai/api/v1/chat/completions

AuthorizationBearer $OPENROUTER_API_KEY

Content-Typeapplication/json

HTTP-Refereroptional — your site URL, for rankings

X-Titleoptional — your site name, for rankings

Modeldeepseek/deepseek-v4-pro-0813

Creates a streaming or non-streaming response using the OpenAI Responses API format.

Docs

POSThttps://openrouter.ai/api/v1/responses

AuthorizationBearer $OPENROUTER_API_KEY

Content-Typeapplication/json

HTTP-Refereroptional — your site URL, for rankings

X-Titleoptional — your site name, for rankings

Modeldeepseek/deepseek-v4-pro-0813

Creates a message using the Anthropic Messages API format. Supports text, images, PDFs, tools, and extended thinking.

Docs

POSThttps://openrouter.ai/api/v1/messages

AuthorizationBearer $OPENROUTER_API_KEY

Content-Typeapplication/json

HTTP-Refereroptional — your site URL, for rankings

X-Titleoptional — your site name, for rankings

Modeldeepseek/deepseek-v4-pro-0813

Parameters

NameTypeDefaultDescription
reasoningmap—Controls reasoning behavior for models that support thinking tokens, including whether reasoning is enabled, the reasoning effort, maximum reasoning tokens, and whether reasoning is excluded from the response.
max_tokensinteger—This sets the upper limit for the number of tokens the model can generate in response.
temperaturefloat1This setting influences the variety in the model's responses.
top_pfloat1This setting limits the model's choices to a percentage of likely tokens: only the top tokens whose probabilities add up to P.
stoparray—Stop generation immediately if the model encounter any token specified in the stop array.
frequency_penaltyfloat0This setting aims to control the repetition of tokens based on how often they appear in the input.
presence_penaltyfloat0Adjusts how often the model repeats specific tokens already used in the input.
logprobsboolean—Whether to return log probabilities of the output tokens or not.
top_logprobsinteger—An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with an associated log probability.
toolsarray—Tool calling parameter, following OpenAI's tool calling request shape.
tool_choicestring or object—Controls which (if any) tool is called by the model.
response_formatmap—Forces the model to produce specific output format.

Frequently asked questions

What is DeepSeek V4 Pro 0813?

How much does DeepSeek V4 Pro 0813 cost?

What is the context length of DeepSeek V4 Pro 0813?

Does DeepSeek V4 Pro 0813 support tool calling and structured outputs?

What other text models does DeepSeek have?

When was DeepSeek V4 Pro 0813 released?

More models from DeepSeek

DeepSeek V4 Flash Latest\n\nThis model always redirects to the latest model in the DeepSeek V4 Flash family.

DeepSeek V4 Flash 0731\n\nDeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

DeepSeek V4 Pro\n\nDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.\n\nBuilt on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

DeepSeek V4 Flash 0423\n\nDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.\n\nThe model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

DeepSeek V3.2 Speciale\n\nDeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on DeepSeek Sparse Attention (DSA) for efficient long-context processing, then scales post-training reinforcement learning to push capability beyond the base model. Reported evaluations place Speciale ahead of GPT-5 on difficult reasoning workloads, with proficiency comparable to Gemini-3.0-Pro, while retaining strong coding and tool-use reliability. Like V3.2, it benefits from a large-scale agentic task synthesis pipeline that improves compliance and generalization in interactive environments.

DeepSeek V3.2\n\nDeepSeek V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.\n\nUsers can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs

DeepSeek V3.2 Exp\n\nDeepSeek V3.2 Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios while maintaining output quality. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model was trained under conditions aligned with V3.1-Terminus to enable direct comparison. Benchmarking shows performance roughly on par with V3.1 across reasoning, coding, and agentic tool-use tasks, with minor tradeoffs and gains depending on the domain. This release focuses on validating architectural optimizations for extended context lengths rather than advancing raw task accuracy, making it primarily a research-oriented model for exploring efficient transformer designs.

DeepSeek V3.1 Terminus\n\nDeepSeek V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's performance in coding and search agents. It is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.

DeepSeek V3.1\n\nDeepSeek V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs\n\nThe model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.\n\nIt succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.

DeepSeek V3.1 Base\n\nThis is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).\n\nDeepSeek V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.

R1 Distill Qwen 7B\n\nDeepSeek R1 Distill Qwen 7B is a 7 billion parameter dense language model distilled from DeepSeek R1, leveraging reinforcement learning-enhanced reasoning data generated by DeepSeek's larger models. The distillation process transfers advanced reasoning, math, and code capabilities into a smaller, more efficient model architecture based on Qwen2.5-Math-7B. This model demonstrates strong performance across mathematical benchmarks (92.8% pass@1 on MATH-500), coding tasks (Codeforces rating 1189), and general reasoning (49.1% pass@1 on GPQA Diamond), achieving competitive accuracy relative to larger models while maintaining smaller inference costs.

DeepSeek R1 0528 Qwen3 8B\n\nDeepSeek R1 0528 is a lightly upgraded release of DeepSeek R1 that taps more compute and smarter post-training tricks, pushing its reasoning and inference to the brink of flagship models like O3 and Gemini 2.5 Pro.\nIt now tops math, programming, and logic leaderboards, showcasing a step-change in depth-of-thought.\nThe distilled variant, DeepSeek R1 0528 Qwen3 8B, transfers this chain-of-thought into an 8 B-parameter form, beating standard Qwen3 8B by +10 pp and tying the 235 B “thinking” giant on AIME 2024.

R1 0528\n\nMay 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.\n\nFully open-source model.

DeepSeek Prover V2\n\nDeepSeek Prover V2 is a 671B parameter model, speculated to be geared towards logic and mathematics. Likely an upgrade from DeepSeek-Prover-V1.5 Not much is known about the model yet, as DeepSeek released it on Hugging Face without an announcement or description.

DeepSeek V3 Base\n\nNote that this is a base model mostly meant for testing, you need to provide detailed prompts for the model to return useful responses.\n\nDeepSeek V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.\n\nDeepSeek V3 Base is the pre-trained model behind DeepSeek V3

DeepSeek V3 0324\n\nDeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.\n\nIt succeeds the DeepSeek V3 model and performs really well on a variety of tasks.

DeepSeek R1 Zero\n\nDeepSeek R1 Zero is a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step. It's 671B parameters in size, with 37B active in an inference pass.\n\nIt demonstrates remarkable performance on reasoning. With RL, DeepSeek R1 Zero naturally emerged with numerous powerful and interesting reasoning behaviors.\n\nDeepSeek R1 Zero encounters challenges such as endless repetition, poor readability, and language mixing. See DeepSeek R1 for the SFT model.

R1 Distill Llama 8B\n\nDeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.\n\nHugging Face:

R1 Distill Qwen 1.5B\n\nDeepSeek R1 Distill Qwen 1.5B is a distilled large language model based on Qwen 2.5 Math 1.5B, using outputs from DeepSeek R1. It's a very small and efficient model which outperforms GPT 4o 0513 on Math Benchmarks.\n\nOther benchmark results include:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

R1 Distill Qwen 32B\n\nDeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\n- AIME 2024 pass@1: 72.6\n- MATH-500 pass@1: 94.3\n- CodeForces Rating: 1691\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

R1 Distill Qwen 14B\n\nDeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

R1 Distill Llama 70B\n\nDeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

R1\n\nDeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.\n\nFully open-source model & technical report.\n\nMIT licensed: Distill & commercialize freely!

DeepSeek V2.5\n\nDeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct. The new model integrates the general and coding abilities of the two previous versions. For model details, please visit DeepSeek-V2 page for more information.

Previous slideNext slide

Compare Playground

Verify your email

We could not find a primary email address on your account. Please contact support.

Cancel

Close

Quick Start

Drop-in code to call this model with OpenRouter's OpenAI-compatible API.

Get Code

OpenRouter SDKOpenAI SDKAnthropic SDKRaw

TypeScriptPythonGo

typescript


 import { OpenRouter } from "@openrouter/sdk";

 const openrouter = new OpenRouter({
 apiKey: ""
 });

 const response = await openrouter.chat.send({
 model: "deepseek/deepseek-v4-pro-0813",
 messages: [
 {
 "role": "user",
 "content": "What is the meaning of life?"
 }
 ]
 });

 console.log(response.choices[0].message.content);

Close

Get Code

OpenRouter SDKOpenAI SDKAnthropic SDKRaw

TypeScriptPythonGo

typescript


 import { OpenRouter } from "@openrouter/sdk";

 const openrouter = new OpenRouter({
 apiKey: ""
 });

 const response = await openrouter.chat.send({
 model: "deepseek/deepseek-v4-pro-0813",
 messages: [
 {
 "role": "user",
 "content": "What is the meaning of life?"
 }
 ]
 });

 console.log(response.choices[0].message.content);

Close

0 / 1

Latency percentiles on OpenRouter All locations

Close

DeepSeek V4 Pro 0813 — Price History

Close

Throughput percentiles on OpenRouter All locations

Close

End-to-End Latency percentiles on OpenRouter All locations

Close

Cache Hit Rate by Provider

Close

Tool Call Error Rate by Provider

Close

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →

Earlier on PAN

More in Technology →