DeepSeek V4 Pro 0813: Pricing, Benchmarks, and Open Weights

DeepSeek V4 Pro 0813 is the August 13, 2026 general-availability build of DeepSeek's largest V4 model. It combines 1.6 trillion total parameters with about 49 billion activated parameters per token, supports a one-million-token context window, and is available both through DeepSeek's API and as MIT-licensed weights. Those headline numbers are real, but they are easy to misuse. Activated parameters describe computation per token, not how much of the checkpoint must be stored.
This guide focuses on what developers can verify from the DeepSeek V4 technical report, the 0813 model card, and the current API documentation. Product and pricing details were checked on September 14, 2026.
DeepSeek V4 Pro 0813 at a glance
| Item | Verified detail |
|---|---|
| Release | GA on August 13, 2026; supersedes V4 Pro Preview |
| API model name | deepseek-v4-pro |
| Architecture size | 1.6T total parameters, approximately 49B activated |
| Context and output | 1M-token context; up to 384K output through the documented API |
| Modalities | Text; DeepSeek's pricing page says V4 Pro does not support vision |
| Interfaces | Chat Completions, Responses API, and Anthropic-compatible API |
| Open weights | deepseek-ai/DeepSeek-V4-Pro-0813, MIT license |
| Thinking controls | low, high, and max; thinking is enabled by default at high |
DeepSeek's September 10 change log also says the company will continue V4 Pro API service after September 14, with the billing method unchanged until further notice. The deepseek-v4-pro API name still maps to 0813. This differs from the Flash line: V4 Flash and V4 Flash Vision Experimental are retired, and their legacy names now route to V4.1 Flash. The recommended current Flash name is deepseek-flash. That is the current status, not a permanent availability guarantee.
What the 1.6T/49B MoE design means
DeepSeek V4 Pro is a mixture-of-experts model. Its 1.6T figure counts the complete parameter set, while routing activates roughly 49B parameters for a token's forward pass. Sparse activation reduces the computation required for each token compared with activating every expert. It does not turn the checkpoint into a 49B model for storage or deployment.
The technical report describes a hybrid of Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), together with Manifold-Constrained Hyper-Connections (mHC). DeepSeek says the V4 models were pretrained on more than 32 trillion tokens. At a one-million-token context, V4 Pro uses 27% of the single-token inference FLOPs and 10% of the KV cache of V3.2 in DeepSeek's tests. That comparison does not apply to every transformer.
The released checkpoint configuration provides more implementation detail: 61 hidden layers, 384 routed experts, six routed experts selected per token, and one shared expert. It sets max_position_embeddings to 1,048,576. These fields support the one-million-token claim, but they do not promise that every deployment can serve the maximum context at acceptable latency or concurrency.
The 0813 model card says a DSpark speculative-decoding module is attached and documents flags for vLLM and SGLang. Measure throughput on your engine, hardware, context, and batch configuration.
The official agent benchmarks and their conditions
The GA release reports a set of agent-oriented results. The most useful public-suite figures are below.
| Benchmark | DeepSeek V4 Pro 0813 | What to keep in mind |
|---|---|---|
| HLE | 42.7 without tools / 60.0 with tools | Tool access materially changes the setup |
| Terminal-Bench 2.1 | 87.9 | Agent framework and sampling settings matter |
| NL2Repo | 61.5 | Repository-generation benchmark |
| CyberGym | 83.3 | Security-agent benchmark |
| DeepSWE | 62.7 | Reported by DeepSeek for the 0813 release |
| Toolathlon-Verified | 74.1 | Verified tool-use task set |
DeepSeek ran public code-agent tasks with DeepSeek Harness in minimal mode, max effort, temperature = 1.0, and top_p = 0.95. Comparisons require the same harness, tool policy, retry budget, and sampling setup.
The release lists DSBench-FullStack at 71.1 and DSBench-Hard at 67.2, but labels both internal. They are not reproducible from the cited public materials. See the DeepSeek Harness evaluation guide for the public-suite runtime context.
The official sources reviewed here contain no 96.4% SWE-bench Verified result for V4 Pro 0813. A valid replacement would require a pinned harness and dataset revision, plus a disclosed attempt budget.
Open weights: storage is based on all parameters
The MIT-licensed Hugging Face repository showed approximately 893 GB of files when checked. Its config uses mixed formats, including FP4 experts and FP8 quantization metadata. The model card provides a vLLM recipe for one four-GPU GB300 node, an SGLang recipe, and local conversion demos.
The practical memory lesson is simple: all experts still have weights. A bare 4-bit calculation for 1.6T parameters is about 800 GB in decimal units:
1.6 trillion parameters × 4 bits ÷ 8 = 800 billion bytes
That is a sizing floor. Scales, non-4-bit tensors, runtime buffers, KV cache, and framework overhead add memory. The active 49B subset is only 24.5 GB at a theoretical four bits, but the remaining experts still need storage.
A 96 GB workstation cannot hold the complete Q4 checkpoint. CPU or disk offload changes where weights live; it does not remove them. Capacity planning must also cover context, batch size, concurrency, and the serving engine.
Use the official vLLM and SGLang recipes as version-sensitive starting points. They were not run for this article, and DeepSeek's four-GB300 example does not validate an eight-H100 or consumer-GPU setup.
DeepSeek API pricing as of September 14, 2026
DeepSeek bills V4 Pro input differently for cache hits and cache misses, then applies weekday peak and off-peak rates. Prices below are US dollars per one million tokens on DeepSeek's direct API.
| Billing item | Off-peak | Peak |
|---|---|---|
| Cached input | $0.022 | $0.044 |
| Uncached input | $0.66 | $1.32 |
| Output | $1.98 | $3.96 |
Peak periods are Monday through Friday, 01:00-04:00 UTC and 06:00-10:00 UTC. All other times are off-peak. DeepSeek notes that prices can change, so production calculators should store rates with an effective date and cache status.
A request with 100,000 uncached input tokens and 10,000 output tokens costs $0.0858 off-peak or $0.1716 at peak. That excludes retries and tool-loop turns. Use the LLM API pricing comparison guide for broader cost modeling.
For gateway access, check the exact model route and billing conditions in the Tokenhot model catalog. Its gateway prices are separate from the DeepSeek direct prices above. The DeepSeek access outside China guide covers account, API, and deployment checks.
A supported Responses API starting point
DeepSeek documents native Responses API support at https://api.deepseek.com. This minimal Python example uses an environment variable for the key and the official V4 Pro model name:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPSEEK_API_KEY"],
base_url="https://api.deepseek.com",
)
response = client.responses.create(
model="deepseek-v4-pro",
instructions="Review the code for correctness and cite the relevant functions.",
input="Explain the failure mode in this retry loop: ...",
)
print(response.output_text)
Install a current OpenAI Python SDK that exposes client.responses.create, set DEEPSEEK_API_KEY, and replace the sample input with your content. The snippet follows DeepSeek's documented request shape but was not run for this article because no credentialed API test was performed.
Thinking defaults to high. DeepSeek's thinking-mode documentation says temperature, presence_penalty, and frequency_penalty have no effect in that mode. Multi-turn tool calls must return reasoning_content as documented. Handle HTTP 429 responses; V4 Pro's listed account concurrency limit is 500, and higher capacity requires a request to DeepSeek.
API or self-hosting: a practical decision
Use the API for quick evaluation or variable demand. Measure task success, output use, cache hits, and retries; a low token price does not guarantee a low cost per successful task.
Self-host when offline operation, infrastructure control, or weight research justifies the checkpoint. Confirm V4 and DSpark support, then model the full memory budget. DeepSeek's four-GB300 recipe does not validate a smaller cluster. The OpenRouter alternatives guide covers gateway comparison criteria.
Frequently asked questions
Is DeepSeek V4 Pro 0813 really a 1.6T model with only 49B active parameters?
Yes. The 49B figure reduces per-token computation; it does not reduce the stored 1.6T checkpoint to 49B parameters.
Can the full model run in 96 GB with Q4 quantization?
No. Four bits across 1.6T parameters is about 800 GB before metadata and runtime memory. A 96 GB machine requires major offload or a smaller model.
Did DeepSeek V4 Pro score 96.4% on SWE-bench Verified?
No official 0813 source reviewed here reports that result. Use the published agent benchmark table and its harness conditions instead.
What does the one-million-token context window cover?
The API lists a 1M-token context window, and the checkpoint config sets 1,048,576 positions. Plan the input and output budget together rather than treating the headline window as an unrestricted prompt allowance. Cost and speed also depend on caching, concurrency, and hardware.
Does V4 Pro support images?
No, according to DeepSeek's current model table. The table marks vision support for deepseek-flash and marks it unsupported for deepseek-v4-pro.
Is the V4 Pro API being discontinued after September 14, 2026?
DeepSeek's September 10 update says API service will continue after September 14 with billing unchanged, with further notice to come if that status changes. Check the change log before building a long-term dependency around the model.
DeepSeek V4 Pro 0813 is an MIT-licensed, open-weight mixture-of-experts model with 1.6 trillion total parameters, about 49 billion activated parameters, and a one-million-token context window. This guide separates those architectural facts from deployment memory, explains the official benchmark conditions, documents DeepSeek's direct API pricing as checked on September 14, 2026, and provides a supported API starting point.


