DeepSeek V4: Specs, Pricing, Benchmarks, and How to Run It (2026 Guide)

DeepSeek V4 is an open-weight MoE family under MIT: V4-Pro-0813 (1.6T params, 49B active) and V4-Flash-0731 (284B, 13B active), both GA with a 1M-token context. Specs, official peak and off-peak pricing, independent benchmarks, InferenceX serving runs on B200 and B300, self-host requirements and the kernels that break on Ampere, cache-hit economics, and working Claude Code, Codex, and OpenCode configs.

August 21, 2026 · 2 min read
V4 Pro total params · 49B active
1.6T
V4 Pro total params · 49B active
V4 Pro single user on B300 (InferenceX, SGLang FP4)
216 tok/s
V4 Pro single user on B300 (InferenceX, SGLang FP4)
V4 Pro output, off-peak ($3.96 peak)
$1.98/M
V4 Pro output, off-peak ($3.96 peak)
8x B200 node to self-host V4 Pro (Lambda)
$53.52/hr
8x B200 node to self-host V4 Pro (Lambda)

DeepSeek V4 Quick Reference

DeepSeek V4 is an open-weight mixture-of-experts model family from DeepSeek, released under the MIT license. It ships in two sizes: V4 Pro, 1.6 trillion parameters with 49 billion active per token, and V4 Flash, 284 billion with 13 billion active. Both carry a 1-million-token context window and 384K max output. Off-peak API pricing is $0.66 and $0.22 per million input tokens.

DeepSeek V4 at a glance (verified September 7, 2026)
AttributeDeepSeek V4 ProDeepSeek V4 Flash
Current checkpointV4-Pro-0813V4-Flash-0731
GA dateAugust 13, 2026July 31, 2026
Total / active parameters1.6T / 49B284B / 13B
Context / max output1,000,000 / 384K1,000,000 / 384K
LicenseMITMIT
API input / output per 1M (off-peak)$0.66 / $1.98$0.22 / $0.66
API cache hit per 1M (off-peak)$0.022$0.007
ModalityText onlyText; images via deepseek-v4-flash-vision-exp
Reference self-host node8x B200 (SGLang, NVFP4) or 4x GB300 (vLLM)8x B200; 4x DGX Spark reported at 49-54 tok/s
Weightsdeepseek-ai/DeepSeek-V4-Pro-0813deepseek-ai/DeepSeek-V4-Flash-0731

Where to run it: the first-party API at api.deepseek.com (OpenAI ChatCompletions, OpenAI Responses, and Anthropic formats), OpenRouter, Cloudflare Workers AI, DeepInfra, Together, Lightning, NVIDIA build and NIM, LM Studio for local GGUF builds, and Morph for morph-dsv4flash at 16-bit activations. Full comparison in the providers section.

TL;DR

Last updated September 7, 2026 with the NVIDIA NVFP4 checkpoint and its accuracy deltas, the Ampere kernel blockers, 1M-context TTFT measurements, the reasoning_content tool-call rule, the working Codex config, current Artificial Analysis scores, and Cloudflare Workers AI availability.

DeepSeek V4 Pro is a 1.6T-parameter mixture-of-experts model with 49B active per token. DeepSeek V4 Flash is its 284B sibling with 13B active. Both are open weight under MIT, both default to a 1M-token context with 384K max output, and both are GA: V4-Flash-0731 since July 31, 2026 and V4-Pro-0813 since August 13. Off-peak API rates are $0.66/M input and $1.98/M output for Pro, $0.22/M and $0.66/M for Flash, double at weekday peak. Artificial Analysis scores Pro-0813 at 36 on its Intelligence Index (rank 7 of 112), 1 point above Flash-0731. Per output token off-peak, V4-Pro is 12.6x cheaper than Claude Opus 4.8 and 15.2x cheaper than GPT-5.5.

Two numbers change the Pro-versus-Flash decision more than the list price. Flash burns 240M output tokens to run the Artificial Analysis Intelligence Index against Pro's 160M, so Flash is 1.5x more verbose for 1 point less intelligence. And in a well-behaved agentic loop, above 99% of input tokens are cache hits at $0.022/M for Pro, which is why practitioners report per-request costs near a tenth of a cent. Both are covered below.

Serving is where Pro gets expensive. InferenceX measures 216 tok/s for one user on a B300 node and 40 tok/s per user at concurrency 64 on B200. An 8x B200 node rents for $53.52 an hour on Lambda, about $0.12 per million tokens at full utilization and $0.40 at 30 percent. Morph serves V4 Flash at $0.141953125/M input and $0.399625/M output with no node to keep busy.

What it is

Two MoE models with 1M-token context. Attention: token-wise compression + DeepSeek Sparse Attention (DSA). Open weights on Hugging Face (V4-Flash-0731, V4-Pro-0813), MIT license. API speaks both OpenAI ChatCompletions and Anthropic formats.

Why it matters

Frontier-adjacent agentic scores at $1.98/M output off-peak. Opus 4.8 output costs $25/M. A dollar buys 505K off-peak output tokens from V4-Pro, 40K from Opus 4.8, 33K from GPT-5.5.

What Is DeepSeek V4?

DeepSeek V4 is an open-weight mixture-of-experts (MoE) model family from DeepSeek under the MIT license, previewed April 24, 2026 and now GA. It ships in two variants: V4-Pro (1.6T total parameters, 49B active per token; current checkpoint V4-Pro-0813) and V4-Flash (284B total, 13B active; current checkpoint V4-Flash-0731). Both default to a 1M-token context window with 384K max output, and the weights are on Hugging Face. The API speaks both the OpenAI ChatCompletions and Anthropic formats, so it drops into tools like Claude Code and OpenCode without a proxy.

$0.141953125 / $0.399625
morph-dsv4flash input / output per 1M (16-bit bf16, no fp8 quant)

Most serverless hosts quantize V4 activations to fp8 to cut cost, which moves output away from the reference weights. Morph serves morph-dsv4flash (DeepSeek V4 Flash) at 16-bit bf16 with no fp8 activation quantization, so output matches the published weights, at $0.141953125/M input and $0.399625/M output. See Morph Open Source Models and pricing.

Release Date and Models

DeepSeek released V4 on April 24, 2026 as a preview, announced on its API docs news page. Two models shipped the same day, and each has since rolled to a GA checkpoint:

  • deepseek-v4-pro: 1.6T total parameters, 49B active per token. Serves DeepSeek-V4-Pro-0813 since August 13, 2026.
  • deepseek-v4-flash: 284B total parameters, 13B active per token. Serves DeepSeek-V4-Flash-0731 since July 31, 2026.

Both run with a 1M-token context window by default across all official DeepSeek services, support JSON output, tool calls, and thinking plus non-thinking modes with three reasoning effort levels (low, high, max), and expose 384K max output tokens. The legacy deepseek-chat and deepseek-reasoner endpoints retired on July 24, 2026 15:59 UTC. The API supports both the OpenAI ChatCompletions format and the Anthropic API format, which is what makes the Claude Code setup below a three-line config.

Weights are on Hugging Face (deepseek-ai/DeepSeek-V4-Pro-0813, deepseek-ai/DeepSeek-V4-Flash-0731), released under the MIT license per the model cards.

What Changed August 15 to September 7, 2026

Everything below is verified against DeepSeek's API docs, changelog, and the cited third-party pages on September 7, 2026.

DeepSeek V4 changes since the Pro GA
DateChangeWhat it means
Aug 13V4-Pro-0813 GA; Responses API and Codex support; DeepSeek Harness developer previewdeepseek-v4-pro serves the 0813 checkpoint. The API accepts OpenAI Responses format, so Codex works with a one-click config. Harness is an MIT-licensed, model-agnostic agent harness built on Cordis, still marked as breaking-change territory.
Aug 16, 16:00 UTCPeak / off-peak pricing livePeak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday only. Weekends bill off-peak all day. Off-peak is exactly half of peak on every line item.
Aug 21deepseek-v4-flash-vision-exp released; Files API launchedExperimental image-input variant of V4-Flash-0731 with the same 1M context and 384K output. Images bill as input tokens at Flash rates, up to 384 tokens per image. Files API lets you upload an image once and reference it by ID.
Aug 25 to 26InferenceX publishes V4-Pro-0813 runs on B200 and B300The first independent serving measurements of the GA checkpoint. Numbers in the serving section below.
Aug 14Cloudflare Workers AI adds both modelsdeepseek-v4-pro-0813 and deepseek-v4-flash-0731 are the first Workers AI models with the full 1,048,576-token context. Thinking mode and function calling are supported. Requires the Workers Paid plan or prepaid AI Gateway credits.
Aug 27NVIDIA publishes nvidia/DeepSeek-V4-Pro-0813-NVFP4An NVFP4 quantization of the GA checkpoint built with TensorRT Model Optimizer, launched on SGLang with --tp 8 on 8x B200. The card lists 1.65T total / 49B active and reports near-parity accuracy against the MXFP4 source. Numbers in the self-host section.
Sep 7 snapshotConcurrency limits, third-party rates, current Artificial Analysis scoresDeepSeek caps deepseek-v4-pro at 500 concurrent requests per account and Flash at 2,500. OpenRouter lists Pro from $0.87/M input and $1.74/M output across 17 providers, Flash from $0.0679/M and $0.168/M. Artificial Analysis now scores Pro-0813 at 36 (rank 7 of 112) and Flash-0731 at 35 (rank 8), with Pro streaming 72.4 output tok/s and Flash 125.8.

Note the moving target. Artificial Analysis showed 53 for Pro-0813 and 52 for Flash-0731 on September 1 under its previous index, and shows 36 and 35 on September 7 under Intelligence Index v4.3, which swapped in ten harder evaluations. The checkpoints did not change; the yardstick did. Treat an Intelligence Index score as valid only within one index version, and never mix a v4.3 score with one quoted from a post written under the old scale. What held across both readings is the gap: Pro sits exactly one point above Flash.

V4 Pro vs V4 Flash: Specs and Pricing

1.6T / 49B
V4-Pro total / active params
284B / 13B
V4-Flash total / active params
1M / 384K
Context / max output (both)
DeepSeek V4 Pro vs Flash (official API, August 2026, off-peak / peak)
SpecificationV4-ProV4-Flash
Total parameters1.6 trillion284 billion
Active parameters per token49B13B
Context window1,000,000 tokens1,000,000 tokens
Max output tokens384K384K
Input, cache miss / 1M tokens$0.66 / $1.32$0.22 / $0.44
Input, cache hit / 1M tokens$0.022 / $0.044$0.007 / $0.014
Output / 1M tokens$1.98 / $3.96$0.66 / $1.32
Concurrency limit per account500 requests2,500 requests
Current checkpointV4-Pro-0813 (Aug 13, 2026)V4-Flash-0731 (Jul 31, 2026)
WeightsHugging Face, MITHugging Face, MIT

Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday. Those two windows are 09:00-12:00 and 14:00-18:00 in Beijing, and the weekday boundary is Beijing time, not yours: OpenCode's pricing docs were amended in PR 44286 specifically to say peak applies on Beijing-time weekdays and that Saturday and Sunday are off-peak all day. If you schedule batch jobs from the Americas, a Friday-evening local run is already Saturday in Beijing and bills at half rate. The other 17 weekday hours also bill off-peak. The cache-hit input prices remain the standout line items. A cache hit on V4-Pro costs $0.022/M off-peak, 30x less than its cache-miss rate. Agentic coding loops that resend the same system prompt and file context on every turn see most of their input tokens land in cache, so effective per-session cost runs far below the list cache-miss rate.

Which to pick: Pro activates 3.8x more parameters per token. Flash costs 3x less on every line item, and Artificial Analysis scores Pro-0813 only 1 point above Flash-0731 (36 vs 35 on its Intelligence Index, September 7, 2026). Flash is the default for batch pipelines, evals, coding, and high-volume extraction; reach for Pro when a hard reasoning task measurably needs it. The verbosity section below narrows that 3x price gap to roughly 2x on realized spend.

DeepSeek V4 variant matrix: Pro vs Pro-Max vs Flash

Three names circulate, which causes confusion. V4-Pro is the served 1.6T MoE. V4-Pro-Max is the benchmark configuration of Pro (the entry that posts the 80.6% SWE-bench Verified score); DeepSeek does not list a separate Pro-Max API price, so the output rate below is Pro's. V4-Flash is the smaller, cheaper variant.

DeepSeek V4 variant matrix (August 2026)
VariantTotal / active paramsContextSWE-bench VerifiedDeepSWEOutput / 1M (off-peak)Current checkpoint
V4-Pro1.6T / 49B1Msee Pro-Max8% preview / 62.7 self-reported (0813)$1.980813 (Aug 13)
V4-Pro-Max1.6T / 49B1M80.6% (preview era)8% (preview, reported)$1.98 (Pro rate)0813 (Aug 13)
V4-Flash284B / 13B1Mno entry54.4 self-reported (0731)$0.660731 (Jul 31)

Read the DeepSWE column carefully. The 8% pass@1 figure came from the independent DeepSWE tracker running the April preview. The 62.7 (Pro-0813) and 54.4 (Flash-0731) figures are DeepSeek's own, from the GA model cards, with no independent reproduction yet. The reconciliation section below covers the preview-era spread.

The Verbosity Tax: Flash Is 3x Cheaper Per Token, Not Per Task

Flash's list price is a third of Pro's on every line item. Flash's realized cost is not, because it writes more tokens to reach the same answer. Three independent measurements say the same thing.

Output volume for the same work (Artificial Analysis, September 7, 2026)
MetricV4-Pro-0813V4-Flash-0731
Output tokens to run the Intelligence Index160M240M (1.5x Pro)
Intelligence Index score36 (rank 7/112)35 (rank 8/112)
Median output tokens per second72.4125.8
Median time to first token1.66 s0.92 s
Blended $/1M at a 7:2:1 cache-hit / input / output mix$0.69$0.23

Artificial Analysis runs the same fixed benchmark suite against every model and publishes the output-token count. Flash needed 240M output tokens to Pro's 160M, and finished one index point lower. On a 3x list-price gap, 1.5x more output closes roughly half of it before you have measured anything about your own workload.

Practitioners report the same shape at higher magnitudes. On the Hacker News thread for the Pro-0813 launch, pixelesque compared the two on code reviews run through OpenRouter and found Flash's output volume often 5x Pro's, with the first pass frequently wrong ("there's a bug" when there is none, or "the code won't compile" when it does) followed by "Wait, let me re-check" and a second pass that lands correctly. An independent review by Thomas Wiegold measured the same tax, putting V4's output at 4 to 5x the median token count and concluding it "reduces the real-world cost advantage."

The practical rule: benchmark on cost per completed task, not cost per million tokens. If your harness caps output or your tasks are short, the list-price gap is real. If your tasks are long-horizon agentic loops where a wrong first pass triggers a re-check, Pro at $1.98/M output can land close to Flash at $0.66/M, and it gets there in fewer turns.

Cache Hits Decide Your DeepSeek V4 Bill

V4-Pro costs $0.66/M for a cache miss and $0.022/M for a cache hit off-peak, a 30x spread. In an agentic loop that resends the same system prompt, file context, and conversation history every turn, nearly every input token should be a hit. The list cache-miss price is close to irrelevant if your harness is behaving.

The concrete numbers come from the same launch thread. Using OpenCode's published token split for DeepSeek V4 Pro sessions, 750 uncached input, 290 output, and 82,000 cached input per request, commenter xynelius computed $0.000875 per request against roughly $0.052 for Claude Opus on the same split, ignoring cache-write cost. That is where the "60x cheaper" figure in the launch discussion comes from, and it is a very different ratio from the 12.6x you get comparing raw output prices.

Three ways teams destroy the cache discount
  • Rewriting history to save context. In the same thread, p1necone reported a standard agentic loop should sit above 99% cache hits, and that a 50% rate means the harness is mutating earlier turns. Stripping old thinking tokens or compacting tool results feels like a cost saving and is the opposite: it invalidates the prefix and reprices the whole conversation at the cache-miss rate.
  • Rotating providers mid-session. Cache is per-provider. Most third-party hosts price cached tokens near 10x DeepSeek's first-party rate, so unpinned failover routing quietly moves you onto the expensive schedule. Pin a primary provider and treat fallbacks as an availability measure, not a routing default.
  • Non-deterministic system prompts. A timestamp, a session ID, or a shuffled tool list at the head of the prompt changes the prefix on every request and drops the hit rate to zero. Put volatile content at the end.

Morph's cache-read rate for morph-dsv4flash is $0.03594/M against $0.14195/M uncached, so the same discipline applies. See pricing for the full table.

DeepSeek V4 Pro Serving Benchmarks

Benchmark scores tell you what V4 Pro can do. Serving numbers tell you what it costs to get that out of a GPU. V4 Pro's FP4 weights alone are about 0.8 TB (1.6T parameters at 4 bits), so every public run uses an 8-GPU Blackwell node or a multi-node rack. The independent data comes from SemiAnalysis InferenceX, which publishes the exact scripts and GitHub Actions runs behind each number.

InferenceX runs of V4-Pro-0813 (August 25, 2026)

DeepSeek V4 Pro on SGLang, FP4, 8K input / 1K output
GPUWorkloadSpeed per userThroughput per GPUTTFTFramework / precisionSource
B3008K input, 1K output216 tok/s216 tok/s350 msSGLang, FP4InferenceX GitHub run
B2008K input, 1K output at concurrency 6440 tok/s2,560 tok/s4,250 msSGLang, FP4InferenceX GitHub run

Read the two rows as the ends of one curve. One user on a B300 node gets 216 tok/s with a 350 ms first token because the whole node works on one request. At concurrency 64 on B200, each user gets 40 tok/s and waits 4.25 s for the first token because 64 prefills of 8K tokens each queue for the same GPUs. Aggregate throughput is what you pay for; per-user speed is what your agent feels.

The B200 vs B300 frontier at three interactivity targets

InferenceX's comparison page sweeps concurrency and reports the best measured throughput per chip at fixed per-user speeds for V4-Pro-0813 in FP4 on the 8K/1K lane. The cost column is InferenceX's own, from its hardware cost model, not a cloud list price.

DeepSeek V4 Pro 0813, FP4, 8K/1K (InferenceX compare page, September 1, 2026)
Speed per userB200 tok/s per chipB300 tok/s per chipB200 $/M (InferenceX)B300 $/M (InferenceX)B300 advantage
59 tok/s15,41018,735$0.031$0.03422% more throughput
102 tok/s6,49710,187$0.074$0.06257% more throughput, 16% cheaper
145 tok/s3,3735,530$0.142$0.11464% more throughput, 20% cheaper

The interactivity tax is steep: going from 59 to 145 tok/s per user cuts B200 throughput per chip by 4.6x. B300's extra HBM matters most at the fast end, where the decode batch is small and the run is memory-bound.

Rack scale and other silicon

  • GB300 NVL72, disaggregated prefill/decode (Dynamo + SGLang): about 11,200 tok/s per GPU at roughly 50 tok/s per user with MTP speculative decoding in June 2026, up from about 2,200 tok/s per GPU on the day-0 April stack, per the SGLang team's PyTorch blog post on the public InferenceX GB300 lane.
  • AMD MI355X, single 8-GPU node, TP8 (SGLang): 2,256 tok/s per GPU on the 8K/1K lane by May 21, 2026, up 110.5x from the FP8-only day-0 number, per InferenceX.
  • Public API providers (Artificial Analysis medians, September 7, 2026): Pro-0813 streams 72.4 output tok/s with a 1.66 s first token; Flash-0731 streams 125.8 tok/s with a 0.92 s first token. Flash is 1.7x faster to a user for the same reason it is 3x cheaper: 13B active parameters versus 49B.

Filter the attributed runs below, or see every model on the dedicated inference benchmarks page. For how these numbers translate into GPU counts, see LLM inference explained.

Interactive explorer

Compare speed and inference cost

Filter attributed measurements by model, GPU, and minimum generation speed. A missing cost means the source run did not report enough information to calculate it.

DeepSeek V4 Pro on B300

8K input, 1K output

independent
User speed
216 tok/sec
GPU throughput
216 tok/sec
Estimated compute cost
Not reported
Serving setup
SGLang, FP4
Open exact source

DeepSeek V4 Pro on B200

8K input, 1K output at concurrency 64

independent
User speed
40 tok/sec
GPU throughput
2,560 tok/sec
Estimated compute cost
Not reported
Serving setup
SGLang, FP4
Open exact source

Reference data is directional. Model version, workload, context length, concurrency, cache state, precision, framework, and topology must match before a result can size a production endpoint.

Run DeepSeek V4 Yourself: Hardware, Precision, and What Breaks

The released checkpoint is already mixed-precision, which changes the usual self-host calculus. NVIDIA's DeepSeek-V4-Pro-0813-NVFP4 model card (published August 27, 2026) describes the layout: routed experts are a lossless bit-cast from MXFP4 to NVFP4, attention projections stay FP8, norms and embeddings stay BF16, and DeepSeek's DSpark speculative-decoding heads are carried through unquantized. So the experts ship at 4 bits from DeepSeek. There is no bf16 reference build to degrade away from on the MoE weights.

That is why the NVFP4 conversion costs almost nothing in quality. NVIDIA's card reports it against the MXFP4 source:

NVFP4 vs the MXFP4 source checkpoint (NVIDIA Model Optimizer, August 27, 2026)
BenchmarkNVFP4MXFP4 source
GPQA Diamond88.4288.51
AA-LCR69.3368.67
Tau-squared Bench Telecom98.2596.49

NVIDIA launches it on SGLang with python3 -m sglang.launch_server --model-path nvidia/DeepSeek-V4-Pro-0813-NVFP4 --tp 8 on 8x B200. DeepSeek's own Pro model card points at vLLM on a single 4x GB300 node. Both are the realistic floor for Pro: 1.6T parameters is roughly 0.8 TB of expert weights at 4 bits before attention, KV cache, and activations.

Ampere does not run V4 yet

This is the gotcha that costs people a weekend. Neither V4 checkpoint runs on SM8x (A100, A800, RTX 30xx) on vLLM main. The tracking issue, vLLM #50576, documents the blocker chain empirically: Triton on SM80 cannot emit fp8e4nv and FP8 KV cache is mandatory for V4; flash_mla_sparse_fwd is SM90a and SM100f only; DeepGEMM fp8_einsum and MegaMoE are Hopper and newer. The issue author's working prototype on 8x A800 80GB reaches 40.6 tok/s single-stream decode with CUDA graphs, about 115 tok/s aggregate at batch 8, boots at max_model_len=1048576, and retrieves a needle from a 200K-token context. Usable, but it is a prototype on a fork, not a supported path.

What people actually get on small hardware

  • Single DGX Spark / GX10, 128 GB unified memory: 6 to 7 tok/s on a DeepSeek-V4-Flash IQ2_XXS GGUF (86.7 GB) under llama.cpp, per a practitioner writeup on the NVIDIA developer forums thread. Stock llama.cpp crashes past an 8K context on a graph-buffer overflow; the thread carries the patch.
  • Four DGX Sparks, TP=4 over RDMA with MTP: 49 to 54 tok/s single-stream on V4 Flash using a community vLLM fork, from the same thread.
  • Consumer GPUs: commenters on the launch thread run 3-bit quants of Flash at roughly 15 tok/s. Pro at 1.6T is not a single-workstation model.

Long context is a time-to-first-token problem, not a memory problem

The 1M window is the headline feature and the hardest thing to serve. In SGLang's DeepSeek V4 perf tracking issue, a contributor with a reserved 8x B200 SXM node posted a same-node topology comparison at 1M context on V4-Flash-0731: shared-KV context parallelism served a p50 time to first token of 61.8 s, plain TP8 77.4 s, and DP8 attention 533.5 s. Serving topology moves 1M-context TTFT by roughly 9x on identical hardware. If you plan to actually use the million tokens, benchmark the topology before you buy the node.

Two self-host behaviors that differ from the API

  • SGLang leaks raw DSML markup on tool_choice: "none". SGLang #35736 reproduces 7 out of 7 times: with a non-empty tools array and tool_choice: "none", the response returns HTTP 200 with finish_reason: "stop", no structured tool calls, and the model's raw DSML tool-call markup sitting inside message.content. The first-party API does not reproduce it, because it strips the tool definitions from the prompt when tools are disabled. SGLang copies them into the system message regardless.
  • No Jinja chat template ships with the model card. Thomas Wiegold's review flags this: budget for supplying your own template in the tokenization pipeline rather than assuming the repo has one.

Self-Host DeepSeek V4 Pro vs Call V4 Flash on Morph

The self-host question has a fixed cost and a utilization problem. Lambda lists an 8x B200 SXM6 node at $6.69 per GPU-hour on demand, so one node costs $53.52 an hour, or about $39070 a month if it never sleeps. Divide that by the tokens InferenceX measured on B200 and you get the cost per million tokens at each interactivity target:

8x B200 on Lambda list price, DeepSeek V4 Pro FP4, 8K/1K, InferenceX throughput
Speed per usertok/s per chipConcurrency$/M at 100% utilization$/M at 30% utilization
59 tok/s15,410~29$0.12$0.40
102 tok/s6,497~9$0.29$0.95
145 tok/s3,373~5$0.55$1.84

The math: $6.69 per GPU-hour divided by (tok/s per chip times 3,600 seconds), per million tokens. InferenceX counts the tokens on its 8K/1K lane the same way for its own $/M column; its lower figures come from a hardware cost model rather than a cloud list price. Thirty percent utilization is generous for a single-team deployment. Coding agents are bursty and a node bought for weekday peaks idles overnight.

Blended cost per million tokens on an 8K input / 1K output request (no cache hits)
OptionInput $/MOutput $/MBlended $/M (8:1)Fixed cost
Self-host V4 Pro, 8x B200, 59 tok/s per user, 100% busyn/an/a$0.12$39070/month
Self-host V4 Pro, 8x B200, 59 tok/s per user, 30% busyn/an/a$0.40$39070/month
DeepSeek API, V4 Pro, off-peak$0.66$1.98$0.81none
DeepSeek API, V4 Pro, weekday peak$1.32$3.96$1.61none
DeepSeek API, V4 Flash, off-peak$0.22$0.66$0.27none
Morph API, morph-dsv4flash, all hours$0.14195$0.400$0.171none

At full utilization a rented B200 node serves V4 Pro for about the same per-token price Morph charges for V4 Flash. At 30% utilization the node costs 3.4x more per token, and 15x more if your agents need 145 tok/s per user. You also still owe the engineering time that took SGLang from 2,200 to 11,200 tok/s per GPU between April and June. Morph's cache-read rate is $0.036/M, which is what agent loops that resend the same context actually pay on most input tokens. Pro's 1-point Intelligence Index lead over Flash is the whole quality delta you would be buying.

If you need Pro specifically, or a private endpoint, the honest path is a dedicated deployment sized from measured throughput rather than list prices. Use the dedicated inference calculator to size it, or read how Morph dedicated inference prices GPU-hours against the tokens they actually produce.

DeepSeek V4 Architecture

Per DeepSeek's release notes, V4's attention combines token-wise compression with DSA (DeepSeek Sparse Attention). DeepSeek's own documentation is sparse on internals; the detail below comes from the model card and third-party technical analyses of the weights, and is marked as such.

Attention: token-wise compression + DSA

The official description: each layer compresses the KV cache token-wise, then applies DeepSeek Sparse Attention over the compressed representation. This is what makes a 1M-token default context economical to serve at $0.66/M off-peak input.

Third-party analyses (Lambda's launch breakdown and model-card readers) describe the mechanism as Compressed Sparse Attention: KV caches compressed 4x along the sequence dimension, with a lightning indexer selecting the top 1,024 compressed KV entries per query. Treat the 4x and top-1,024 figures as secondary-source; DeepSeek's news page does not publish them.

CSA and HCA: the hybrid that makes 1M context cheap

The Hugging Face DeepSeek V4 blog names the attention stack a hybrid of two mechanisms interleaved across layers, not a single "DSA" block (DSA is the V3.2-era term). The two are Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA).

  • CSA compresses KV entries 4x along the sequence dimension using softmax-gated pooling with a learned positional bias. A lightning indexer (run in FP4, a ReLU-scored multi-head dot product) selects the top-k compressed blocks per query, and a sliding-window branch handles the most recent uncompressed tokens.
  • HCA compresses KV entries 128x and drops sparse selection entirely: every query attends densely to every compressed block.

Per the HF blog, V4-Pro is a 61-layer stack where layers 0 and 1 are HCA and layers 2 through 60 alternate CSA and HCA. Both paths store most KV entries in FP8 and keep BF16 only for the RoPE dimensions. The compound effect: the KV cache lands at roughly 2% the size of 8-head grouped-query attention in bfloat16 (V4-Pro at about 10% of V3.2's KV memory, V4-Flash at about 7%), and single-token inference runs at 27% of V3.2's FLOPs for Pro and 10% for Flash. That is what makes a 1M-token default context economical even at $0.66/M off-peak input.

DeepSeek V4 Pro architecture

V4-Pro is the 1.6T-total-parameter MoE with 49B active per token. Versus V3 (671B total, 37B active), that is 2.4x the total parameter count and 1.3x the active compute per token, with the context window expanded 8x from 128K to 1M. The sparse-attention stack is what keeps the larger model servable: DeepSeek prices Pro at $1.98/M output off-peak, 3x the output price of V4-Flash but a fraction of closed-model rates.

DeepSeek V4 Flash architecture

V4-Flash shares the attention design but at 284B total parameters with 13B active per token. The smaller expert pool is why DeepSeek can price Flash at $0.22/M off-peak input, a third of Pro on every line item. It also inherits the same 1M context and 384K max output, so the variant choice is about quality per token, not context capability.

What DeepSeek has and has not published

Confirmed by DeepSeek's own release notes: MoE parameter counts (1.6T/49B and 284B/13B), token-wise compression + DSA attention, 1M default context, 384K max output, thinking and non-thinking modes, OpenAI and Anthropic API compatibility. Not published first-party as of June 18, 2026: the CSA compression ratios, indexer internals, optimizer details, and training token counts that circulate in third-party writeups. We cite those as secondary-source where used.

Benchmarks: SWE-bench Verified

The independently tracked number is DeepSeek-V4-Pro-Max at 80.6% on SWE-bench Verified (llm-stats tracker, June 2026). That is the highest open-weights entry, tied with Gemini 3.1 Pro, 0.1 points ahead of MiniMax M3 and 0.2 ahead of Qwen3.7 Max. Closed frontier models score higher: Claude Fable 5 leads at 95.0% (currently suspended, see note).

SWE-bench Verified, top 10 (llm-stats tracker, June 2026)
ModelScoreOutput price / 1M tokens
Claude Fable 595.0%$50.00
Claude Mythos Preview93.9%restricted access
Claude Opus 4.888.6%$25.00
Claude Opus 4.787.6%$25.00
Claude Opus 4.580.9%$25.00
Claude Opus 4.680.8%$25.00
DeepSeek-V4-Pro-Max80.6%$1.98 (off-peak)
Gemini 3.1 Pro80.6%$12.00
MiniMax M380.5%$1.20
Qwen3.7 Max80.4%$2.40 (Qwen3.5-Plus rate)

Read the price column against the score column. The 14.4-point gap between Fable 5 (95.0%) and V4-Pro-Max (80.6%) costs 25x more per output token at V4-Pro's off-peak rate. The 8-point gap to Opus 4.8 costs 12.6x more. Whether that trade is worth it depends entirely on whether your tasks live in the band those extra points unlock.

DeepSeek-reported launch numbers

At launch, DeepSeek reported V4-Pro-Max scoring 93.5 on LiveCodeBench Pass@1 and a 3206 Codeforces rating. These are vendor-run numbers from the April 2026 release coverage, not independent leaderboard entries; vendor scaffolds routinely score above standardized harnesses (see our SWE-bench Pro breakdown for how large that gap runs).

GA checkpoint benchmarks (self-reported, August 2026)

The GA checkpoints post-train heavily for agentic work, and the model-card numbers move accordingly. Flash-0731 vs the April preview: Terminal-Bench 2.1 61.8 to 82.7, DeepSWE 7.3 to 54.4, Toolathlon-Verified 49.7 to 70.3. Pro-0813 vs the April Pro preview: Terminal-Bench 2.1 72.1 to 87.9, DeepSWE 12.8 to 62.7, NL2Repo 38.5 to 61.5, CyberGym 52.7 to 83.3. Every cell is DeepSeek-run; no independent reproduction of a GA-specific figure exists yet. The independent signal so far is Artificial Analysis: Pro-0813 scores 36 on Intelligence Index v4.3 (checked September 7, 2026), 1 point above Flash-0731's 35.

Does DeepSeek V4 Really Score 80.6%? Self-Reported vs Independent

The 80.6% SWE-bench Verified number is real, but it sits at the top of a harness with a loose verifier. On DeepSWE, a written-from-scratch, contamination-free benchmark, the April V4-Pro preview scored 8% pass@1 versus GPT-5.5 at 70% and Opus 4.7 at 54%. The same ordering held on SWE-bench Pro. Read all three together before trusting any one. Note the checkpoint scoping: everything in this section measured the April preview. DeepSeek self-reports 62.7 on DeepSWE for Pro-0813, but on its own harness; the independent trackers have not re-run the GA build.

DeepSeek V4-Pro across three harnesses vs frontier (June 2026)
HarnessV4-Pro scoreGPT-5.5Opus 4.7What it measures
SWE-bench Verified80.6% (Pro-Max)not on tracker87.6%Patch passes held-out tests; loose verifier
SWE-bench Pro76.2%82.6%82.0%Harder repos; verifier ~24% false negatives
DeepSWE8% (reported)70.0%54.0%Written-from-scratch, 91 repos, 5 langs, no contamination

How to reconcile the spread. The yage.ai DeepSWE audit sampled tasks across both benchmarks and an external LLM judge, and found SWE-bench Pro's verifier rejects about 24% of functionally correct solutions and accepts about 8.5% of incorrect ones, while DeepSWE's verifier runs at 1.1% false negatives and 0.3% false positives. A tighter verifier and contamination-free tasks make DeepSWE harder to inflate.

That said, V4-Pro's 8% DeepSWE result is anomalously low relative to peers and to its own 76.2% on SWE-bench Pro. The audit's read is that the gap reflects a real long-horizon-agent capability difference rather than a verifier artifact, because V4-Pro ranks lowest on both independent harnesses. The SWE-bench Pro figures above are the ones reported in the yage.ai audit, not Scale's SEAL board (which still has no V4 entry); treat the Opus 4.7 and GPT-5.5 numbers as reported by the DeepSWE and audit trackers, not independently re-run by us.

DeepSeek V4 Flash and SWE-bench: What Exists

As of July 2026, Flash has aggregated and self-reported SWE-bench numbers but still no contamination-free independent re-run: llm-stats lists DeepSeek-V4-Flash-Max at 79.0% SWE-bench Verified, DeepSeek's own technical report says 73.7%, and our DeepSeek V4 Flash deep-dive covers the gap between those two figures. The state of the trackers:

  • Scale's SEAL SWE-bench Pro leaderboards (public and private sets) list no DeepSeek V4 entry of any variant. The top open-weights entry there is qwen3-coder-480b-a35b at 38.7% on the public set.
  • The llm-stats SWE-bench Verified tracker lists DeepSeek-V4-Pro-Max (80.6%) but no Flash entry.

If you see a Flash SWE-bench number quoted, check whether it is a vendor-run scaffold result or a community harness; neither currently appears on the two trackers above. For agentic coding where benchmark evidence exists, the published data points at Pro, not Flash.

Harness Choice Moves the Score More Than the Checkpoint Does

The cleanest independent check on DeepSeek's GA claims is not a leaderboard. It is someone re-running the vendor's own benchmark in a different agent harness and publishing the contract.

In OpenCode issue 42553, a maintainer ran Terminal-Bench 2.1 against deepseek/deepseek-v4-flash under OpenCode 1.18.7 with Harbor 0.20.0 in Docker, one attempt per task, no retries, on a pinned dataset digest. Result: 60 of 89 tasks passed, 67.42%. DeepSeek reports 82.7% for V4-Flash-0731 on the same benchmark. That is a 15.28-point gap on an identical dataset.

The gap is not proof of an inflated vendor number. DeepSeek ran its own harness in minimal mode at max reasoning effort; the OpenCode run set no variant and used the API default of high. But the issue also lists four open OpenCode bugs that would each depress the score on their own, including --variant being silently ignored when the value is not in the model's reasoning options, and reasoning_content being dropped from tool-call messages. Read every V4 agentic number as a model-plus-harness result.

How large is the harness effect in general? On Artificial Analysis's coding-agent harness comparison, swapping harnesses moved Claude Opus 4.7 by about 10 points on the coding agent index, larger than the gap between two adjacent frontier checkpoints. For V4 specifically, practitioners in the launch thread report the same sensitivity: Flash-0731 rated "much better in OpenCode than in Pi" on identical tasks. If your V4 results look far off the published numbers, change the harness before you change the model.

Cost Math vs Opus 4.8, GPT-5.5, Gemini 3.1 Pro

Reddit and X discussion of V4 settled on a "17x cheaper" shorthand at launch. The August 16, 2026 repricing shrank the ratios: increases ran 50% to 1,100% across line items. The exact ratios now depend on which token type you compare and on the hour. At off-peak list prices:

V4-Pro list-price ratios vs frontier models (August 2026, off-peak)
ModelInput / 1MOutput / 1MOutput tokens per $1Output cost vs V4-Pro
DeepSeek V4-Flash (off-peak)$0.22$0.661.52M0.33x
DeepSeek V4-Pro (off-peak)$0.66$1.98505K1x
MiniMax M3 (≤512K)$0.30$1.20833K0.61x
Gemini 3.1 Pro (≤200K)$2.00$12.0083K6.1x
GPT-5.4$2.50$15.0067K7.6x
Claude Sonnet 4.6$3.00$15.0067K7.6x
Claude Opus 4.8$5.00$25.0040K12.6x
GPT-5.5$5.00$30.0033K15.2x
Claude Fable 5$10.00$50.0020K25.3x

The launch-era "17x cheaper" shorthand no longer holds. Against GPT-5.4 and Sonnet 4.6 output, V4-Pro off-peak is now 7.6x cheaper; at peak, 3.8x. Against Opus 4.8 it is 12.6x off-peak on output. Against Fable 5 it is 25.3x. Note MiniMax M3 output ($1.20/M) now undercuts V4-Pro at every hour. Cached input still widens the frontier gaps: a V4-Pro cache hit costs $0.022/M off-peak vs $0.50/M for an Opus 4.8 cache hit, 23x apart.

A concrete daily workload, 20 requests of 50K input + 10K output (1M input, 200K output per day), assuming zero cache hits and off-peak hours:

  • V4-Flash: $0.35/day, about $11/month
  • V4-Pro: $1.06/day, about $32/month
  • Opus 4.8: $10.00/day, about $300/month
  • GPT-5.5: $11.00/day, about $330/month
  • Claude Fable 5: $20.00/day, about $600/month

Full Anthropic-side rates, including cache and batch multipliers, are in our Anthropic API pricing guide.

Where to Run It: First-Party API vs OpenRouter

The first-party API (api.deepseek.com) is the reference deployment: $0.66/$1.98 off-peak for Pro and $0.22/$0.66 for Flash (double at peak), with the cache-hit discounts above. On September 1, 2026 OpenRouter listed deepseek/deepseek-v4-pro from $0.87/M input and $1.74/M output with 1M context across 17 providers (DigitalOcean cheapest, CoreWeave fastest at 50 tok/s), and deepseek/deepseek-v4-flash from $0.0679/M input and $0.168/M output. Per-host rates vary by routing tier, so check the live OpenRouter provider list before committing volume. The image-input variant deepseek-v4-flash-vision-exp (released August 21, 2026) is first-party only and bills at Flash rates.

Cloudflare added both models to Workers AI on August 14, 2026, and they are the first Workers AI models carrying the full 1,048,576-token context, with thinking mode and function calling supported. They are reachable through the Workers AI binding, the REST API, the OpenAI-compatible endpoint, or AI Gateway, and require the Workers Paid plan or prepaid AI Gateway credits. That is the shortest path if your agent already runs at the edge, and it is why a Workers AI changelog post sits on page one for this query.

Reasons to pick first-party: documented cache-hit pricing at $0.022/M off-peak (Pro) and the Anthropic-format endpoint. Reasons to pick OpenRouter: one key across models, provider failover, no peak-hour surcharge, and easy A/B against MiniMax M3 or Qwen3.5 at the prices in the table above. Because the weights are open, you can also self-host; Flash at 284B total is the realistic target, Pro at 1.6T is multi-node territory.

Output fidelity is where serverless hosts diverge. Most serverless providers quantize activations to fp8 to cut serving cost, which moves output away from the reference weights. Morph Open Source Models serves DeepSeek with 16-bit (bf16) activations and does not quantize activations to fp8, so output matches the published weights. For coding agents specifically, Morph adds codegen-tuned speculative decoding plus custom low-level inference kernels built for code generation, which makes it the fastest and highest-quality option for codegen. morph-dsv4flash (DeepSeek V4 Flash) runs at $0.141953125/M input and $0.399625/M output; see pricing for the full list.

Thinking Mode and the reasoning_content Rule

Both V4 models expose thinking mode through two request parameters. Send "thinking": {"type": "enabled"} or {"type": "disabled"} to switch it, and reasoning_effort with low, high, or max to set the budget. DeepSeek documents low for simple tasks, high for daily agent workflows, and max for complex ones. high is the default, which is why benchmark runs that forget to request max land below DeepSeek's published numbers.

The single most common DeepSeek V4 agent bug

The reasoning text comes back in reasoning_content, alongside content. DeepSeek's thinking-mode guide states that when the request carries a tools parameter, reasoning_content from all previous turns must be passed back and will be concatenated into context, including turns where the model made no tool call. Without tools it is ignored and can be dropped.

Harnesses that strip it break in two different ways depending on the provider. OpenCode issue 35689 documents both: OpenRouter returns HTTP 400 with "The reasoning_content in the thinking mode must be passed back to the API," while an OpenAI-compatible endpoint swallows it and returns finish_reason: "stop", so the agent silently exits the loop mid-task. If your DeepSeek agent stops after a few steps with no error, check whether your message conversion preserves reasoning_content on assistant messages that also carry tool_calls.

Use DeepSeek V4 in Claude Code, Codex, and OpenCode

Claude Code

DeepSeek's API speaks the Anthropic format natively, so Claude Code needs only environment variables, no proxy:

export ANTHROPIC_BASE_URL="https://api.deepseek.com/anthropic"
export ANTHROPIC_AUTH_TOKEN="sk-your-deepseek-key"
export ANTHROPIC_MODEL="deepseek-v4-pro"
export ANTHROPIC_SMALL_FAST_MODEL="deepseek-v4-flash"
claude

Pointing the small-fast model at deepseek-v4-flash keeps background tasks on the $0.66/M-off-peak-output tier. For routing multiple providers behind one endpoint instead, see Claude Code with LiteLLM.

One failure mode shows up mid-session rather than at startup. Claude Code's adaptive thinking sends a thinking.type value DeepSeek's Anthropic-compatible endpoint rejects, so a run dies after the first thinking block with API Error: 400 'type' must be in ["enabled", "disabled", "auto"]. Compacting does not clear it. The workaround users converged on in cc-switch issue 5897 is to pin thinking to a fixed mode in ~/.claude/settings.json instead of letting Claude Code negotiate it:

{
  "env": {
    "ANTHROPIC_BASE_URL": "https://api.deepseek.com/anthropic",
    "ANTHROPIC_AUTH_TOKEN": "sk-your-deepseek-key",
    "ANTHROPIC_MODEL": "deepseek-v4-pro",
    "CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING": "1"
  }
}

Disabling adaptive thinking keeps thinking mode available at a fixed effort. Setting CLAUDE_CODE_DISABLE_THINKING or MAX_THINKING_TOKENS=0 also clears the error but gives up reasoning entirely, which is the wrong trade on a model whose agentic scores depend on it.

OpenCode

OpenCode ships a DeepSeek provider. Run opencode auth login, select DeepSeek, paste your API key, then set the model in opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "model": "deepseek/deepseek-v4-pro"
}

Codex CLI

This changed with the Pro GA and older guides are now wrong, including an earlier version of this page. DeepSeek's August 13, 2026 release notes state "Native OpenAI Responses API support, optimized for Codex with one-click setup," so Codex talks to api.deepseek.com directly with wire_api = "responses" and no proxy. In ~/.codex/config.toml:

model = "deepseek-v4-pro"
model_provider = "deepseek"
preferred_auth_method = "apikey"
forced_login_method = "api"
model_reasoning_effort = "high"
model_catalog_json = "~/.codex/models.json"

[model_providers.deepseek]
name = "deepseek"
base_url = "https://api.deepseek.com/"
wire_api = "responses"
experimental_bearer_token = "<your DeepSeek API key>"

Two things bite here. The models.json catalog is required, not optional, and a stale one omits deepseek-v4-pro entirely; DeepSeek ships a setup script that writes both files. And a leftover wire_api = "chat" from a pre-GA guide surfaces as a 404, 405, or a Responses API error rather than a clear message. The same config shape works for any provider; see our Codex provider configuration guide.

A calibrating data point from the launch thread: one commenter ran the same new feature through Codex CLI on both models via OpenRouter. DeepSeek V4 Pro took 12 minutes 2 seconds and $0.12 and shipped a bug; Grok 4.6 took 3 minutes 18 seconds and $1.41 and did not. V4 is 12x cheaper per task and slower, and on that task it was also wrong. Single runs prove nothing on their own, but the shape of the trade is consistent with the verbosity numbers above.

Vision: deepseek-v4-flash-vision-exp and Its 800x800 Ceiling

Base V4 is text only. The experimental image-input variant deepseek-v4-flash-vision-exp shipped August 21, 2026, first-party only, at Flash token rates with the same 1M context.

The constraint that matters is resolution. DeepSeek resizes every image so the total pixel count lands near 800x800, scaling up anything below roughly 384x384, which caps an image at 384 input tokens. A 2000x2000 image and a 5000x5000 image therefore cost the same and carry the same detail. On the Hacker News thread for the release, that ceiling is the main complaint: it is too coarse for OCR of a full A4 or Letter page, and too coarse for schematics where symbols carry spatial meaning.

The workaround that thread converged on is a crop tool rather than a bigger upload. Because DeepSeek only downscales images that exceed the budget, a small cropped subimage arrives unresized. Give the model a zoom tool that returns a crop of the original at native resolution, let it read the 800x800 overview first, then request the regions it needs.

The variant also fixes a real annoyance. Multiple users on that thread report that text-only Flash-0731 assumes it can see, pulls screenshots unprompted, and then invents pixel-analysis workarounds when it cannot, breaking the session. If your agent takes screenshots, either route them to the vision variant or tell the text model explicitly that it has no vision.

What Changed from DeepSeek V3

DeepSeek V3 vs V4-Pro
DimensionDeepSeek V3DeepSeek V4-Pro
Total parameters671B1.6T (2.4x larger)
Active parameters37B per token49B per token
Context window128K tokens1M tokens (8x larger)
AttentionMulti-head Latent Attention (MLA)Token-wise compression + DSA
Max output8K tokens384K tokens
API formatsOpenAI-compatibleOpenAI + Anthropic formats
Endpointsdeepseek-chat / deepseek-reasonerdeepseek-v4-pro / deepseek-v4-flash

The structural shift is the attention stack: replacing MLA with token-wise compression plus DSA is what moves the default context from 128K to 1M without 1M-context pricing. The legacy V3-era endpoints survived only as aliases onto deepseek-v4-flash and retired on July 24, 2026.

Where DeepSeek R1 sits: R1 (671B total, 37B active, 128K context, MIT) launched January 20, 2025 as the model behind deepseek-reasoner. Per DeepSeek's changelog that endpoint moved to R1-0528 on May 28, 2025, then to the thinking modes of V3.1 (August 21, 2025), V3.2-Exp, and V3.2, then aliased to V4-Flash thinking mode on April 24, 2026 before retiring. V4 has no separate reasoning model. Both V4-Pro and V4-Flash expose thinking mode with low, high, and max reasoning effort, and R1's 49.2% on SWE-bench Verified sits 31 points under V4-Pro-Max's 80.6%. R1 remains on Hugging Face for self-hosting and its license explicitly permits distillation.

Limitations

Known limitations
  • Pricing volatility: the August 16, 2026 repricing raised line items 50% to 1,100% and introduced peak/off-peak billing. Budgets and blog posts written against the launch rates are stale; re-check the pricing page before committing volume. Direction of travel is unclear rather than one-way: at launch DeepSeek indicated Pro pricing could fall once Huawei Ascend 950 supernodes deploy at scale in the second half of 2026. Treat that as an unconfirmed forward-looking statement, not a rate you can plan against.
  • GA benchmarks are self-reported: the large agentic gains on the 0731 and 0813 checkpoints (Terminal-Bench 2.1, DeepSWE, NL2Repo) all come from DeepSeek's own harness. No independent tracker has re-run the GA builds yet.
  • 8-point gap to the frontier on SWE-bench Verified: V4-Pro-Max scores 80.6% vs Opus 4.8's 88.6% and Fable 5's 95.0%. For tasks where those points matter, the cheap model retries its way into costing you time instead of money.
  • No independent SWE-bench Pro entry: Scale's SEAL leaderboard has no DeepSeek V4 result, so agentic performance under a standardized harness is unverified.
  • Sparse first-party architecture docs: compression ratios and indexer internals circulate only in third-party analyses.
  • It does not signal uncertainty. Two independent readings put V4's hallucination rate in the same band on the AA-Omniscience evaluation: Thomas Wiegold's review reports 94% when uncertain, and the AINews launch digest records 94 to 96% across the family despite its ranking gains. The model answers confidently rather than declining. In a codebase audit it over-flagged non-issues while missing real bugs a frontier model caught. Pair it with a test suite or a reviewing model rather than trusting a single confident pass.
  • Ampere and older GPUs are unsupported: vLLM main cannot serve either checkpoint on SM8x. Blackwell or Hopper, or a community fork.
  • Verbosity eats the price gap: Flash burns 1.5x Pro's output tokens on the Artificial Analysis suite, and 4 to 5x median in independent reviews. Budget on cost per completed task.
  • Self-hosting Pro is heavy: 1.6T total parameters means an 8-GPU Blackwell node at minimum, about $39,070 a month at Lambda list price (see the self-host section above). Flash at 284B is the practical self-host target.

Frequently Asked Questions

When was DeepSeek V4 released?

April 24, 2026 as a preview, with V4-Pro and V4-Flash shipping the same day. GA came in two waves: V4-Flash-0731 on July 31, 2026 and V4-Pro-0813 on August 13, 2026. The legacy deepseek-chat and deepseek-reasoner endpoints retired on July 24, 2026 15:59 UTC.

What is the DeepSeek V4 architecture?

A mixture-of-experts transformer with token-wise compression plus DeepSeek Sparse Attention (DSA). Pro: 1.6T total / 49B active. Flash: 284B total / 13B active. Both: 1M context, 384K max output. Third-party analyses add 4x KV compression and a top-1,024 lightning indexer, which DeepSeek has not published first-party.

What does DeepSeek V4 cost?

Since August 16, 2026 the official API bills peak and off-peak. Off-peak: Flash $0.22/M input (miss), $0.007/M (cache hit), $0.66/M output; Pro $0.66/M, $0.022/M, $1.98/M. Peak hours (01:00-04:00 and 06:00-10:00 UTC, Monday through Friday) bill double; weekends are off-peak all day. On September 1, 2026 OpenRouter listed Pro from $0.87/$1.74 and Flash from $0.0679/$0.168.

How fast does DeepSeek V4 Pro run on B200 and B300?

InferenceX runs of V4-Pro-0813 on SGLang in FP4 (8K input, 1K output) measure 216 tok/s for a single user on a B300 node with a 350 ms first token, and 40 tok/s per user at concurrency 64 on B200 with a 4.25 s first token. At 59 tok/s per user the InferenceX frontier is 15,410 tok/s per chip on B200 and 18,735 on B300. Public API providers stream Pro at a median 72.4 tok/s and Flash at 125.8 (Artificial Analysis, September 7, 2026).

What does it cost to self-host DeepSeek V4 Pro?

An 8x B200 node on Lambda is $53.52 an hour, about $39070 a month. At InferenceX's 15,410 tok/s per chip that is $0.12 per million tokens at 100% utilization and $0.40 at 30%. Calling V4 Flash on Morph is $0.171/M blended on the same 8:1 mix with no fixed cost. Full table in the self-host section above.

Where does DeepSeek R1 sit versus V4?

R1 was the January 2025 model behind deepseek-reasoner. That endpoint walked through R1-0528, V3.1, V3.2, and a V4-Flash alias before retiring on July 24, 2026. V4 folds reasoning into both Pro and Flash as an effort parameter (low, high, max). R1 scored 49.2% on SWE-bench Verified; V4-Pro-Max scores 80.6%.

What is deepseek-v4-flash's SWE-bench score?

llm-stats lists Flash-Max at 79.0% SWE-bench Verified; DeepSeek's technical report says 73.7%. Neither is an independent contamination-free re-run, and Scale SEAL still has no Flash entry. Full breakdown on the DeepSeek V4 Flash page.

Does DeepSeek V4 really score 80.6% on SWE-bench?

On SWE-bench Verified, yes: V4-Pro-Max posts 80.6% (llm-stats), but that harness has a loose verifier. On DeepSWE, a contamination-free written-from-scratch benchmark, the April V4-Pro preview scored 8% pass@1 versus GPT-5.5 at 70% and Opus 4.7 at 54%. DeepSeek self-reports 62.7 on DeepSWE for the Pro-0813 GA build, on its own harness, with no independent re-run yet.

How does V4-Pro-Max compare to Claude on SWE-bench Verified?

V4-Pro-Max 80.6% vs Opus 4.6 80.8%, Opus 4.8 88.6%, Fable 5 95.0% (llm-stats, June 2026). Per output token V4-Pro costs $1.98 off-peak ($3.96 peak) vs $25 for Opus 4.8 and $50 for Fable 5.

Is V4 open source?

Open weights on Hugging Face under MIT. The current GA checkpoints are deepseek-ai/DeepSeek-V4-Flash-0731 and deepseek-ai/DeepSeek-V4-Pro-0813. Download, run, fine-tune.

Can I use V4 in Claude Code?

Yes. Set ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic, ANTHROPIC_AUTH_TOKEN to your DeepSeek key, and ANTHROPIC_MODEL=deepseek-v4-pro. If a session dies with 400 'type' must be in ["enabled", "disabled", "auto"] after the first thinking block, add CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1. The full snippet is in the setup section above.

How do I run DeepSeek V4 Pro locally, and how much VRAM does it need?

Pro needs a multi-GPU Blackwell node. NVIDIA serves its NVFP4 build on SGLang with 8x B200 (--tp 8); DeepSeek's model card points at vLLM on a single 4x GB300 node. Expert weights alone are roughly 0.8 TB at 4 bits. Ampere (A100, A800, RTX 30xx) does not work on vLLM main. Flash at 284B is the practical self-host target: 6 to 7 tok/s on one DGX Spark with an IQ2_XXS GGUF, 49 to 54 tok/s across four with TP=4 over RDMA.

How many parameters does DeepSeek V4 Pro have?

About 1.6 trillion total with 49 billion active per token. NVIDIA's quantized card lists 1.65T total; the Hugging Face tensor index reads higher because the checkpoint also ships the DSpark speculative-decoding module. Flash is 284B total, 13B active.

How do I enable thinking mode in the DeepSeek V4 API?

Send "thinking": {"type": "enabled"} and set reasoning_effort to low, high, or max. High is the default. The reasoning text returns in reasoning_content, and when your request carries tools you must pass it back on every subsequent turn or the agent loop breaks.

Is DeepSeek V4 multimodal?

Base V4 is text only. deepseek-v4-flash-vision-exp (August 21, 2026, first-party only) accepts images at Flash rates, resized to about 800x800 total pixels and capped at 384 tokens each. Use a crop tool for fine detail; small subimages are not downscaled.

How do I use DeepSeek V4 in VS Code?

Any VS Code extension that accepts a custom OpenAI-compatible endpoint works: set the base URL to https://api.deepseek.com, paste a DeepSeek API key, and pick deepseek-v4-pro or deepseek-v4-flash. The Codex extension reads the same ~/.codex/config.toml as the CLI, so the Responses-API config above covers it.

Related Articles

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 100 billion tokens per day run this way.

Talk to us about a private deployment

Use WarpGrep with DeepSeek V4 for Better Code Search Context

WarpGrep is an agentic code search tool that works as an MCP server. Connect it to any DeepSeek-powered agent for high-precision codebase context, so V4's 1M-token window gets filled with the right code, not noise. $0.80 per 100K tokens.

Sources