DeepSeek V4 Flash 0731 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis
Model summary
- Intelligence: #3 / 101 — 50 on the Artificial Analysis Intelligence Index (4 out of 4 units)
- Speed: N/A — output tokens per second unknown
- Price: #22 / 101 — Input $0.14 per 1M tokens; Output $0.28 per 1M tokens (1 out of 4 units)
- Cache Hit Price: #1 / 101 — $0.003 per 1M tokens (-98%; 1 out of 4 units)
- Verbosity: #36 / 101 — 210M output tokens on the Intelligence Index (4 out of 4 units)
Comparison Summary
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is amongst the leading models in intelligence and well priced when comparing to other open weight models of similar size. The model supports text input, outputs text, and has a 1M tokens context window.
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) scores 50 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 25). When evaluating the Intelligence Index, it generated 210M tokens, which is very verbose in comparison to the median of 100M.
Pricing for DeepSeek V4 Flash 0731 (Reasoning, Max Effort) is $0.14 per 1M input tokens (competitively priced, median: $0.43) and $0.28 per 1M output tokens (competitively priced, median: $1.20). In total, it cost $72.02 to evaluate DeepSeek V4 Flash 0731 (Reasoning, Max Effort) on the Intelligence Index.
Technical specifications
| Reasoning | Yes. This page shows the reasoning version of this model. A non-reasoning variant may also exist. | | Input modality | Supports: text | | Output modality | Supports: text | | Context window | 1M (~1500 A4 pages of size 12 Arial font) | | Total parameters | 284B | | Active parameters | 13B (number of parameters active per token during inference) | | License | MIT | | Model weights | Hugging Face |
101 models in this class
Metrics are compared against models of the same class:
- Non-reasoning models → compared only with other non-reasoning models
- Reasoning models → compared across both reasoning and non-reasoning
- Open weights models → compared only with other open weights models of the same size class (Tiny: ≤4B; Small: 4B–40B; Medium: 40B–150B; Large: >150B)
- Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio ($1 per 1M tokens)
Intelligence
The Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) scores 50 on the index, ranking #3 out of 101 models in its class.
Benchmarks
The following intelligence evaluations are measured independently by Artificial Analysis:
- GDPval-AA v2 — Agentic real-world work tasks (Elo-500)/2000
- 𝜏³-Banking — Agentic tool use
- Terminal-Bench v2.1 — Agentic coding & terminal use
- SciCode — Coding
- Humanity's Last Exam — Reasoning & knowledge
- GPQA Diamond — Scientific reasoning
- CritPt — Physics reasoning
- AA-Omniscience Accuracy — Knowledge
- AA-Omniscience Non-Hallucination Rate — 1 - hallucination rate
- AA-LCR — Long context reasoning
- AA-Briefcase — Agentic knowledge work, Elo
- AutomationBench-AA — Agentic SaaS workflows
- Harvey LAB-AA — Legal agentic work, criterion pass rate
- EnterpriseOps-Gym-AA — Agentic business operations
- IFBench — Instruction following
- APEX-Agents-AA — Long-horizon agentic tasks
- ITBench-AA — Kubernetes incident root-cause analysis
- MMMU-Pro — Visual reasoning
AA-Omniscience
The AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Openness Index
The Artificial Analysis Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open).
Token Use
The number of tokens required per Intelligence Index task is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the Intelligence Index, then dividing by task count (excluding repeats). DeepSeek V4 Flash 0731 generated 210M output tokens across the index, with answer and reasoning tokens tracked separately.
Cost
Weighted average cost per Intelligence Index task is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. The total cost to run all evaluations in the Intelligence Index for this model was $72.02.
Cache Hit, Input, and Output Pricing are shown in USD per million tokens. Cache hit refers to the price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price; cache write and cache storage are billed separately and vary by provider.
Context Window
DeepSeek V4 Flash 0731 has a context window of 1.0M tokens. Larger context windows are relevant to RAG (Retrieval Augmented Generation) workflows, which typically involve reasoning and information retrieval of large amounts of data.
Model Size (Open Weights Models Only)
The model has 284 billion total parameters (trainable weights and biases), with 13 billion active parameters used during inference. It is a Mixture of Experts (MoE) model.
Frequently Asked Questions
When was DeepSeek V4 Flash 0731 (Reasoning, Max Effort) released?
July 31, 2026.
Who created DeepSeek V4 Flash 0731 (Reasoning, Max Effort)?
DeepSeek.
How intelligent is DeepSeek V4 Flash 0731 (Reasoning, Max Effort)?
It scores 50 on the Artificial Analysis Intelligence Index, placing it well above average among other open weight models of similar size (median: 25).
How much does DeepSeek V4 Flash 0731 (Reasoning, Max Effort) cost?
$0.14 per 1M input tokens (very competitive, median: $0.58) and $0.28 per 1M output tokens (very competitive, median: $2.20), based on DeepSeek's API. For a blended rate (7:2:1 cache hit/input/output ratio), this is $0.06 per 1M tokens. Pricing may vary by provider.
How verbose is DeepSeek V4 Flash 0731 (Reasoning, Max Effort)?
When evaluated on the Intelligence Index, it generated 210M output tokens, which is at the higher end compared to other open weight models of similar size (median: 100M).
Is DeepSeek V4 Flash 0731 (Reasoning, Max Effort) a reasoning model?
Yes, it uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer.
What input and output modalities does it support?
Text input and text output. It does not support image input and is not multimodal.
What is the context window?
1.0M tokens.
Is it open source?
Yes, it is an open weights model. The model weights are publicly available and can be downloaded for self-hosting.
How many parameters does it have?
284 billion total parameters (13 billion active). It is a Mixture of Experts (MoE) model.
What is the license?
The MIT license, which allows commercial use. View license
How does it perform on benchmarks?
It achieves a score of 50 on the Artificial Analysis Intelligence Index, a composite benchmark evaluating models across reasoning, knowledge, mathematics, and coding.
Is it available via API?
Yes, through 1 provider.




