DeepSeek-V4-Pro-0813
| Tier | Input | Output | Cached input |
|---|---|---|---|
PriorityLearn more | 702 | 1404 | 54 |
FlexLearn more | 374.4 | 748.8 | 28.8 |
DeepSeek V4 Pro 0813 is the official release of V4 Pro, and the point at which the series flagship became a production agent rather than a preview. It keeps the architecture that defines the series — 1.6 trillion total parameters with 49 billion active per token, a one-million-token context window, and a hybrid attention stack built to make that window practical — and rebuilds the agentic behaviour on top of it. Long-horizon software engineering, terminal automation, security workflows, and tool orchestration all improve substantially over the preview, in several cases by a factor of several. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license.

DeepSeek V4 Pro 0813
Overview
DeepSeek V4 Pro 0813 is the official release of DeepSeek V4 Pro, superseding the preview checkpoint. It keeps the architecture that defines the V4 series — a Mixture-of-Experts model with 1.6 trillion total parameters of which 49B activate per token, and a one-million-token context window — and rebuilds what the model does when it has to act.
The gains are concentrated in production behaviour rather than raw knowledge. On long-horizon software engineering the resolve rate moves from 12.8 to 62.7. On security-oriented code tasks, from 52.7 to 83.3. On terminal automation, from 72.1 to 87.9. These are not tuning adjustments; they are the difference between a model that starts a long task and one that finishes it.
Two properties define how it is used: reasoning effort is a per-request control with three levels, and the checkpoint ships with a speculative decoding module attached, which serving stacks use to generate faster without changing the output distribution.
At a Glance
| Field | Value |
|---|---|
| Model Type | Mixture-of-Experts transformer |
| Total Parameters | 1.6T |
| Activated Parameters | 49B per token |
| Context Window | 1,048,576 tokens (1M) |
| Precision | FP4 + FP8 mixed |
| Modality | Text in → text out |
| Reasoning | Three effort levels, separate reasoning_content field |
| Tool Calling | Supported |
| Speculative Decoding | Module included in the checkpoint |
| License | MIT |
Architecture
| Component | Detail |
|---|---|
| Sparsity | MoE — 1.6T total, 49B activated per token |
| Attention | Hybrid stack: Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA) |
| Residual path | Manifold-Constrained Hyper-Connections (mHC) |
| Speculative decoding | Draft module attached to the same checkpoint |
| Precision | MoE expert weights in FP4; attention, normalization and router in FP8 |
Hybrid attention is what makes the context window usable rather than nominal. CSA and HCA together bring per-token inference compute down to about a quarter, and key-value cache to about a tenth, of the previous DeepSeek generation at the 1M-token setting.
mHC strengthens the conventional residual connection, improving stability of signal propagation across layers while preserving expressivity — a training-stability property, but part of why a model at this scale stays coherent across enormous inputs.
Speculative decoding is unusual here in that the draft weights live in the same checkpoint as the target model rather than in a separate smaller one. The practical effect is lower generation latency, which matters most on the very long outputs this model produces at high reasoning effort.
Relationship to the Preview Checkpoint
This release supersedes the V4 Pro preview. Same architecture, same context window, same activated parameter count — rebuilt agentic behaviour, a renamed reasoning-effort scale, and an attached speculative decoding module.
If you are currently calling the preview checkpoint, migration is a model ID change plus one
adjustment: the fastest reasoning level is now named low.
Reasoning Effort Levels
The most important operational decision when using this model.
| Level | Behaviour | Use it for |
|---|---|---|
low | Minimal deliberation, fast responses | Routine tasks, classification, extraction, formatting |
high | Explicit reasoning before answering | Complex problems, planning, code, analysis |
max | Reasoning pushed to its fullest extent | Long-horizon agent work, the hardest problems |
DeepSeek evaluates the agentic benchmarks below at the max level. If you are building an
agent that must complete a long task rather than answer a question, that is the level the
published results describe.
Selecting a level
DEVUP AI forwards the complete request body upstream without stripping unknown fields, so reasoning effort can be passed directly in your payload:
{
"model": "deepseek-ai/DeepSeek-V4-Pro-0813",
"messages": [{ "role": "user", "content": "Fix the failing test in this repository." }],
"reasoning_effort": "max",
"temperature": 1.0,
"top_p": 0.95
}Capabilities
| Capability | Value |
|---|---|
input_types | text |
output_types | text |
image_input | Not supported |
context_window | 1048576 |
reasoning | Native — low, high, max |
reasoning_field | reasoning_content — separate from content |
streaming | Supported |
tool_calling | Supported |
requires_prompt | Yes — text prompt required |
Recommended Use Cases
- The hardest long-horizon coding work — tasks spanning many turns, many files, and many failed attempts, where the lighter model in the series starts to lose the thread.
- Security and correctness-critical code analysis — one of the largest improvements in this release.
- Terminal and infrastructure automation — planning command sequences, reading output, and recovering from errors without a human in the loop.
- Research-grade question answering with tools — the combination of deep reasoning and tool use is where this checkpoint gained the most.
- Whole-corpus analysis at full window — a codebase, a contract set, or a long log archive held entirely in context, with the cross-document relationships that chunking destroys.
- The escalation tier in a routed system — see below.
Choosing Between This Model and the Lighter Release
Both official releases in the series share an API surface, a context window, and a reasoning-effort control, which makes routing between them a configuration change rather than an integration.
The gap is not uniform, and that is the point:
| Benchmark | This model | Lighter release | Gap |
|---|---|---|---|
| Agents' Last Exam | 25.7 | 25.2 | 0.5 |
| DSBench-FullStack | 71.1 | 68.7 | 2.4 |
| Toolathlon-Verified | 74.1 | 70.3 | 3.8 |
| Terminal-Bench 2.1 | 87.9 | 82.7 | 5.2 |
| AutomationBench | 31.8 | 25.1 | 6.7 |
| Cybergym | 83.3 | 76.7 | 6.6 |
| NL2Repo | 61.5 | 54.2 | 7.3 |
| DSBench-Hard | 67.2 | 59.6 | 7.6 |
| DeepSWE | 62.7 | 54.4 | 8.3 |
On general tool orchestration and routine full-stack work the two are close. The flagship separates itself on difficulty: hard coding-agent problems, security analysis, and long-horizon repository work. Route accordingly — paying flagship latency for a task the lighter model resolves at within two points is a poor trade.
Using DeepSeek V4 Pro 0813 on DEVUP AI
Base URL: https://api.devupai.com/v1 · Model ID: deepseek-ai/DeepSeek-V4-Pro-0813
Quick start — cURL
curl https://api.devupai.com/v1/chat/completions \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-ai/DeepSeek-V4-Pro-0813",
"messages": [
{
"role": "user",
"content": "Two services write to the same row without a transaction. Walk through the failure modes in order of likelihood and propose the smallest fix for each."
}
],
"reasoning_effort": "high",
"temperature": 1.0,
"top_p": 1.0,
"max_tokens": 32768
}'Node.js — DEVUP AI SDK
npm install devupaiimport DevupAI from "devupai";
const client = new DevupAI({
apiKey: process.env.DEVUP_API_KEY,
});
const response = await client.chat.completions.create({
model: "deepseek-ai/DeepSeek-V4-Pro-0813",
messages: [
{
role: "system",
content:
"You are a security reviewer. Report every input that reaches a query, a shell, or a " +
"filesystem path without validation. For each, give the reachable path and the " +
"smallest safe fix.",
},
{ role: "user", content: sourceBundle },
],
reasoning_effort: "max",
temperature: 1.0,
top_p: 0.95,
max_tokens: 32768,
});
console.log(response.choices[0].message.content);Agent loop — Python
The scenario this release targets. Note top_p: 0.95 and reasoning_effort: "max", which are
the settings DeepSeek uses for its own agent evaluations.
import os
import json
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEVUP_API_KEY"],
base_url="https://api.devupai.com/v1",
)
TOOLS = [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Return the contents of a file at a repository-relative path.",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"],
},
},
},
{
"type": "function",
"function": {
"name": "run_tests",
"description": "Run the test suite and return pass/fail counts with failure output.",
"parameters": {"type": "object", "properties": {}},
},
},
]
def read_file(path: str) -> dict:
"""Replace with your real, sandboxed file access."""
raise NotImplementedError
def run_tests() -> dict:
"""Replace with your real, sandboxed test runner."""
raise NotImplementedError
HANDLERS = {"read_file": read_file, "run_tests": run_tests}
messages = [{"role": "user", "content": "The invoice rounding test is failing. Find the cause and fix it."}]
for _ in range(40): # bounded loop — never let an agent iterate without a ceiling
response = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4-Pro-0813",
messages=messages,
tools=TOOLS,
reasoning_effort="max",
temperature=1.0,
top_p=0.95,
max_tokens=32768,
)
message = response.choices[0].message
messages.append(message)
if not message.tool_calls:
print(message.content)
break
for call in message.tool_calls:
handler = HANDLERS.get(call.function.name)
if handler is None:
result = {"error": "unknown_tool", "name": call.function.name}
else:
try:
result = handler(**json.loads(call.function.arguments or "{}"))
except Exception as exc: # surface the failure to the model, do not crash
result = {"error": type(exc).__name__, "detail": str(exc)}
messages.append(
{
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
}
)
else:
print("Agent loop exceeded its iteration limit.")Returning a structured error rather than raising is deliberate. This checkpoint was trained on trajectories where actions fail and the agent recovers, so a described failure is information it can use.
Reading the reasoning trace
Reasoning arrives in a separate field, not inline in the answer. Read it explicitly, and null-check it — not every model on the platform populates it.
message = response.choices[0].message
reasoning = getattr(message, "reasoning_content", None)
if reasoning:
# Log it, do not show it. Reasoning traces are intermediate, not conclusions.
logger.debug("trace length: %d chars", len(reasoning))
print(message.content)Never concatenate reasoning_content into content before parsing or display. Doing so
breaks JSON parsing on structured-output paths and shows users an unpolished draft of an
answer they never asked to see.
Streaming with usage
stream = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4-Pro-0813",
messages=[{"role": "user", "content": "Design a retry policy for a webhook delivery system."}],
reasoning_effort="high",
temperature=1.0,
top_p=1.0,
max_tokens=32768,
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print(f"\n\nTokens — in: {chunk.usage.prompt_tokens}, out: {chunk.usage.completion_tokens}")Setting stream_options.include_usage returns a final chunk carrying token counts. At high
and max effort the trace can dominate the output budget, and this is the only way to see it.
Delegating access with a scoped JWT
The combination of a million-token window, an unbounded reasoning budget, and an unattended agent loop is the most expensive mistake available on this platform. Issue a token restricted to this model with an expiry and a spending limit instead of sharing your API key:
curl -X POST "https://api.devupai.com/v1/scoped-jwt" \
-H "Authorization: Bearer $DEVUP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"api_key_name": "auto",
"models": ["deepseek-ai/DeepSeek-V4-Pro-0813"],
"expires_delta": 7200,
"spending_limit": 1000
}'The returned token is used exactly like an API key in the Authorization header. Requests for
any other model, or past the expiry or spending limit, are rejected — a hard ceiling on what a
runaway loop can consume.
Recommended Generation Parameters
| Parameter | Value |
|---|---|
temperature | 1.0 |
top_p | 0.95 for agentic scenarios, 1.0 otherwise |
max_tokens | Large — see below |
DeepSeek recommends allowing up to 384K output tokens at the high and max effort
levels. Truncating a reasoning model mid-trace yields an unfinished, unusable response rather
than a shorter one, so size max_tokens to the effort level you selected, not to the answer
you expect.
Benchmark Results
As reported by DeepSeek, evaluated at the max reasoning effort level with
temperature = 1.0, top_p = 0.95, using DeepSeek's own agent harness in minimal mode. The
comparison columns are the other checkpoints in the series.
| Benchmark | This release | Lighter release | Pro (preview) | Flash (preview) |
|---|---|---|---|---|
| HLE (without / with tools) | 42.7 / 60.0 | 37.8 / 51.5 | 37.7 / 48.2 | 34.8 / 45.1 |
| Terminal-Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 |
| AutomationBench (public) | 31.8 | 25.1 | 12.8 | 10.8 |
| DSBench-FullStack | 71.1 | 68.7 | 41.8 | 37.0 |
| DSBench-Hard | 67.2 | 59.6 | 31.1 | 25.8 |
DSBench-FullStack and DSBench-Hard are DeepSeek's internal test sets for full-stack development and difficult coding-agent problems respectively. DeepSeek also publishes comparisons against leading models from other developers, which are not reproduced here.
Two rows are worth reading twice. DeepSWE, 12.8 to 62.7 on the same architecture at the same size — long-horizon software engineering is what this release changed. And HLE with tools, 48.2 to 60.0 against 42.7 without them, which says the gain is in using tools well, not in knowing more.
Best Practices
- Set
reasoning_effortdeliberately, per request. It is the highest-impact parameter on this model, and no single value is right for every path in an application. - Use
maxfor agent work. The published agentic results describe that level; running an agent atlowis not a cheaper version of the same behaviour. - Set
top_pto 0.95 in agentic scenarios, 1.0 elsewhere. This is a documented split, not a preference. - Route by difficulty, not by default. On routine agentic work the lighter release lands within a few points; reserve this tier for the hard cases.
- Budget output tokens generously at
highandmax. A truncated reasoning model returns nothing useful. - Read
reasoning_contentas a separate field. Do not merge it intocontent, and null-check it — other models on the platform leave it empty. - Return tool errors as data. This checkpoint was trained to recover from failed actions, so a described failure is more useful than a raised exception.
- Bound every agent loop with an iteration ceiling and a scoped token.
- Use the context window instead of building retrieval where the corpus fits.
Limitations
- Text only. No image, audio, or document input.
loweffort is a different capability tier, not merely a faster one. Treat the levels as distinct configurations rather than a speed dial.- Reasoning traces are not conclusions. Content in
reasoning_contentmay be unpolished or contradict the final answer. - Agentic gains do not imply knowledge gains. This release rebuilt long-horizon action; factual recall is a separate axis and moved far less.
- Benchmark figures are vendor-reported and depend on the harness and settings used. They compare most reliably within the series.
- Not a safety layer. Apply your own moderation and validation before acting on model output in a production system.