Modelsdeepseek-aiDeepSeek-V4-Pro-0813
providerdeepseek-ai /

DeepSeek-V4-Pro-0813

455 DZD in 910 DZD out 35 DZD cached/ 1M tokens
Service tier pricing, in DZD per 1M tokens
TierInputOutputCached input
PriorityLearn more
702140454
374.4748.828.8
Prices in DZD per 1M tokens

DeepSeek V4 Pro 0813 is the official release of V4 Pro, and the point at which the series flagship became a production agent rather than a preview. It keeps the architecture that defines the series — 1.6 trillion total parameters with 49 billion active per token, a one-million-token context window, and a hybrid attention stack built to make that window practical — and rebuilds the agentic behaviour on top of it. Long-horizon software engineering, terminal automation, security workflows, and tool orchestration all improve substantially over the preview, in several cases by a factor of several. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license.

Publicfp8JSON
DeepSeek-V4-Pro-0813
Capabilities
ToolsReasoning (optional)Structured output
ArchitectureTransformer
Context Window1M

DeepSeek V4 Pro 0813

Overview

DeepSeek V4 Pro 0813 is the official release of DeepSeek V4 Pro, superseding the preview checkpoint. It keeps the architecture that defines the V4 series — a Mixture-of-Experts model with 1.6 trillion total parameters of which 49B activate per token, and a one-million-token context window — and rebuilds what the model does when it has to act.

The gains are concentrated in production behaviour rather than raw knowledge. On long-horizon software engineering the resolve rate moves from 12.8 to 62.7. On security-oriented code tasks, from 52.7 to 83.3. On terminal automation, from 72.1 to 87.9. These are not tuning adjustments; they are the difference between a model that starts a long task and one that finishes it.

Two properties define how it is used: reasoning effort is a per-request control with three levels, and the checkpoint ships with a speculative decoding module attached, which serving stacks use to generate faster without changing the output distribution.


At a Glance

FieldValue
Model TypeMixture-of-Experts transformer
Total Parameters1.6T
Activated Parameters49B per token
Context Window1,048,576 tokens (1M)
PrecisionFP4 + FP8 mixed
ModalityText in → text out
ReasoningThree effort levels, separate reasoning_content field
Tool CallingSupported
Speculative DecodingModule included in the checkpoint
LicenseMIT

Architecture

ComponentDetail
SparsityMoE — 1.6T total, 49B activated per token
AttentionHybrid stack: Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA)
Residual pathManifold-Constrained Hyper-Connections (mHC)
Speculative decodingDraft module attached to the same checkpoint
PrecisionMoE expert weights in FP4; attention, normalization and router in FP8

Hybrid attention is what makes the context window usable rather than nominal. CSA and HCA together bring per-token inference compute down to about a quarter, and key-value cache to about a tenth, of the previous DeepSeek generation at the 1M-token setting.

mHC strengthens the conventional residual connection, improving stability of signal propagation across layers while preserving expressivity — a training-stability property, but part of why a model at this scale stays coherent across enormous inputs.

Speculative decoding is unusual here in that the draft weights live in the same checkpoint as the target model rather than in a separate smaller one. The practical effect is lower generation latency, which matters most on the very long outputs this model produces at high reasoning effort.


Relationship to the Preview Checkpoint

This release supersedes the V4 Pro preview. Same architecture, same context window, same activated parameter count — rebuilt agentic behaviour, a renamed reasoning-effort scale, and an attached speculative decoding module.

If you are currently calling the preview checkpoint, migration is a model ID change plus one adjustment: the fastest reasoning level is now named low.


Reasoning Effort Levels

The most important operational decision when using this model.

LevelBehaviourUse it for
lowMinimal deliberation, fast responsesRoutine tasks, classification, extraction, formatting
highExplicit reasoning before answeringComplex problems, planning, code, analysis
maxReasoning pushed to its fullest extentLong-horizon agent work, the hardest problems

DeepSeek evaluates the agentic benchmarks below at the max level. If you are building an agent that must complete a long task rather than answer a question, that is the level the published results describe.

Selecting a level

DEVUP AI forwards the complete request body upstream without stripping unknown fields, so reasoning effort can be passed directly in your payload:

JSON
{
  "model": "deepseek-ai/DeepSeek-V4-Pro-0813",
  "messages": [{ "role": "user", "content": "Fix the failing test in this repository." }],
  "reasoning_effort": "max",
  "temperature": 1.0,
  "top_p": 0.95
}

Capabilities

CapabilityValue
input_typestext
output_typestext
image_inputNot supported
context_window1048576
reasoningNative — low, high, max
reasoning_fieldreasoning_content — separate from content
streamingSupported
tool_callingSupported
requires_promptYes — text prompt required

Recommended Use Cases

  • The hardest long-horizon coding work — tasks spanning many turns, many files, and many failed attempts, where the lighter model in the series starts to lose the thread.
  • Security and correctness-critical code analysis — one of the largest improvements in this release.
  • Terminal and infrastructure automation — planning command sequences, reading output, and recovering from errors without a human in the loop.
  • Research-grade question answering with tools — the combination of deep reasoning and tool use is where this checkpoint gained the most.
  • Whole-corpus analysis at full window — a codebase, a contract set, or a long log archive held entirely in context, with the cross-document relationships that chunking destroys.
  • The escalation tier in a routed system — see below.

Choosing Between This Model and the Lighter Release

Both official releases in the series share an API surface, a context window, and a reasoning-effort control, which makes routing between them a configuration change rather than an integration.

The gap is not uniform, and that is the point:

BenchmarkThis modelLighter releaseGap
Agents' Last Exam25.725.20.5
DSBench-FullStack71.168.72.4
Toolathlon-Verified74.170.33.8
Terminal-Bench 2.187.982.75.2
AutomationBench31.825.16.7
Cybergym83.376.76.6
NL2Repo61.554.27.3
DSBench-Hard67.259.67.6
DeepSWE62.754.48.3

On general tool orchestration and routine full-stack work the two are close. The flagship separates itself on difficulty: hard coding-agent problems, security analysis, and long-horizon repository work. Route accordingly — paying flagship latency for a task the lighter model resolves at within two points is a poor trade.


Using DeepSeek V4 Pro 0813 on DEVUP AI

Base URL: https://api.devupai.com/v1 · Model ID: deepseek-ai/DeepSeek-V4-Pro-0813

Quick start — cURL

BASH
curl https://api.devupai.com/v1/chat/completions \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-V4-Pro-0813",
    "messages": [
      {
        "role": "user",
        "content": "Two services write to the same row without a transaction. Walk through the failure modes in order of likelihood and propose the smallest fix for each."
      }
    ],
    "reasoning_effort": "high",
    "temperature": 1.0,
    "top_p": 1.0,
    "max_tokens": 32768
  }'

Node.js — DEVUP AI SDK

BASH
npm install devupai
JAVASCRIPT
import DevupAI from "devupai";

const client = new DevupAI({
  apiKey: process.env.DEVUP_API_KEY,
});

const response = await client.chat.completions.create({
  model: "deepseek-ai/DeepSeek-V4-Pro-0813",
  messages: [
    {
      role: "system",
      content:
        "You are a security reviewer. Report every input that reaches a query, a shell, or a " +
        "filesystem path without validation. For each, give the reachable path and the " +
        "smallest safe fix.",
    },
    { role: "user", content: sourceBundle },
  ],
  reasoning_effort: "max",
  temperature: 1.0,
  top_p: 0.95,
  max_tokens: 32768,
});

console.log(response.choices[0].message.content);

Agent loop — Python

The scenario this release targets. Note top_p: 0.95 and reasoning_effort: "max", which are the settings DeepSeek uses for its own agent evaluations.

PYTHON
import os
import json
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEVUP_API_KEY"],
    base_url="https://api.devupai.com/v1",
)

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "read_file",
            "description": "Return the contents of a file at a repository-relative path.",
            "parameters": {
                "type": "object",
                "properties": {"path": {"type": "string"}},
                "required": ["path"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "run_tests",
            "description": "Run the test suite and return pass/fail counts with failure output.",
            "parameters": {"type": "object", "properties": {}},
        },
    },
]


def read_file(path: str) -> dict:
    """Replace with your real, sandboxed file access."""
    raise NotImplementedError


def run_tests() -> dict:
    """Replace with your real, sandboxed test runner."""
    raise NotImplementedError


HANDLERS = {"read_file": read_file, "run_tests": run_tests}

messages = [{"role": "user", "content": "The invoice rounding test is failing. Find the cause and fix it."}]

for _ in range(40):  # bounded loop — never let an agent iterate without a ceiling
    response = client.chat.completions.create(
        model="deepseek-ai/DeepSeek-V4-Pro-0813",
        messages=messages,
        tools=TOOLS,
        reasoning_effort="max",
        temperature=1.0,
        top_p=0.95,
        max_tokens=32768,
    )

    message = response.choices[0].message
    messages.append(message)

    if not message.tool_calls:
        print(message.content)
        break

    for call in message.tool_calls:
        handler = HANDLERS.get(call.function.name)
        if handler is None:
            result = {"error": "unknown_tool", "name": call.function.name}
        else:
            try:
                result = handler(**json.loads(call.function.arguments or "{}"))
            except Exception as exc:  # surface the failure to the model, do not crash
                result = {"error": type(exc).__name__, "detail": str(exc)}

        messages.append(
            {
                "role": "tool",
                "tool_call_id": call.id,
                "content": json.dumps(result),
            }
        )
else:
    print("Agent loop exceeded its iteration limit.")

Returning a structured error rather than raising is deliberate. This checkpoint was trained on trajectories where actions fail and the agent recovers, so a described failure is information it can use.

Reading the reasoning trace

Reasoning arrives in a separate field, not inline in the answer. Read it explicitly, and null-check it — not every model on the platform populates it.

PYTHON
message = response.choices[0].message

reasoning = getattr(message, "reasoning_content", None)
if reasoning:
    # Log it, do not show it. Reasoning traces are intermediate, not conclusions.
    logger.debug("trace length: %d chars", len(reasoning))

print(message.content)

Never concatenate reasoning_content into content before parsing or display. Doing so breaks JSON parsing on structured-output paths and shows users an unpolished draft of an answer they never asked to see.

Streaming with usage

PYTHON
stream = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Pro-0813",
    messages=[{"role": "user", "content": "Design a retry policy for a webhook delivery system."}],
    reasoning_effort="high",
    temperature=1.0,
    top_p=1.0,
    max_tokens=32768,
    stream=True,
    stream_options={"include_usage": True},
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print(f"\n\nTokens — in: {chunk.usage.prompt_tokens}, out: {chunk.usage.completion_tokens}")

Setting stream_options.include_usage returns a final chunk carrying token counts. At high and max effort the trace can dominate the output budget, and this is the only way to see it.

Delegating access with a scoped JWT

The combination of a million-token window, an unbounded reasoning budget, and an unattended agent loop is the most expensive mistake available on this platform. Issue a token restricted to this model with an expiry and a spending limit instead of sharing your API key:

BASH
curl -X POST "https://api.devupai.com/v1/scoped-jwt" \
  -H "Authorization: Bearer $DEVUP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "api_key_name": "auto",
    "models": ["deepseek-ai/DeepSeek-V4-Pro-0813"],
    "expires_delta": 7200,
    "spending_limit": 1000
  }'

The returned token is used exactly like an API key in the Authorization header. Requests for any other model, or past the expiry or spending limit, are rejected — a hard ceiling on what a runaway loop can consume.


Recommended Generation Parameters

ParameterValue
temperature1.0
top_p0.95 for agentic scenarios, 1.0 otherwise
max_tokensLarge — see below

DeepSeek recommends allowing up to 384K output tokens at the high and max effort levels. Truncating a reasoning model mid-trace yields an unfinished, unusable response rather than a shorter one, so size max_tokens to the effort level you selected, not to the answer you expect.


Benchmark Results

As reported by DeepSeek, evaluated at the max reasoning effort level with temperature = 1.0, top_p = 0.95, using DeepSeek's own agent harness in minimal mode. The comparison columns are the other checkpoints in the series.

BenchmarkThis releaseLighter releasePro (preview)Flash (preview)
HLE (without / with tools)42.7 / 60.037.8 / 51.537.7 / 48.234.8 / 45.1
Terminal-Bench 2.187.982.772.161.8
NL2Repo61.554.238.539.4
Cybergym83.376.752.738.7
DeepSWE62.754.412.87.3
Toolathlon-Verified74.170.355.949.7
Agents' Last Exam25.725.216.515.8
AutomationBench (public)31.825.112.810.8
DSBench-FullStack71.168.741.837.0
DSBench-Hard67.259.631.125.8

DSBench-FullStack and DSBench-Hard are DeepSeek's internal test sets for full-stack development and difficult coding-agent problems respectively. DeepSeek also publishes comparisons against leading models from other developers, which are not reproduced here.

Two rows are worth reading twice. DeepSWE, 12.8 to 62.7 on the same architecture at the same size — long-horizon software engineering is what this release changed. And HLE with tools, 48.2 to 60.0 against 42.7 without them, which says the gain is in using tools well, not in knowing more.


Best Practices

  • Set reasoning_effort deliberately, per request. It is the highest-impact parameter on this model, and no single value is right for every path in an application.
  • Use max for agent work. The published agentic results describe that level; running an agent at low is not a cheaper version of the same behaviour.
  • Set top_p to 0.95 in agentic scenarios, 1.0 elsewhere. This is a documented split, not a preference.
  • Route by difficulty, not by default. On routine agentic work the lighter release lands within a few points; reserve this tier for the hard cases.
  • Budget output tokens generously at high and max. A truncated reasoning model returns nothing useful.
  • Read reasoning_content as a separate field. Do not merge it into content, and null-check it — other models on the platform leave it empty.
  • Return tool errors as data. This checkpoint was trained to recover from failed actions, so a described failure is more useful than a raised exception.
  • Bound every agent loop with an iteration ceiling and a scoped token.
  • Use the context window instead of building retrieval where the corpus fits.

Limitations

  • Text only. No image, audio, or document input.
  • low effort is a different capability tier, not merely a faster one. Treat the levels as distinct configurations rather than a speed dial.
  • Reasoning traces are not conclusions. Content in reasoning_content may be unpolished or contradict the final answer.
  • Agentic gains do not imply knowledge gains. This release rebuilt long-horizon action; factual recall is a separate axis and moved far less.
  • Benchmark figures are vendor-reported and depend on the harness and settings used. They compare most reliably within the series.
  • Not a safety layer. Apply your own moderation and validation before acting on model output in a production system.