Where deepseek-v4-pro and deepseek-v4-flash
+sit against the field, what the 0813 checkpoint changed, and the economics
+argument the Chinese community calls the 斩杀线, the
+kill line. The numbers below are DeepSeek's own launch-day figures unless
+marked otherwise; read the caveats first.
+
+
+read this before quoting a number
+
+
Vendor numbers, one harness. The table is DeepSeek's own
+agent-benchmark chart, published 2026-08-12 with the GA. A score is a
+model-and-harness result, not a model-only one, and no independent same-harness
+run of 0813 exists yet. Treat these as the claim, not the verdict.
+
No GPT-5.6 column. DeepSeek's chart compares against Kimi
+K3, GLM-5.2, Claude Opus 4.8 and Fable 5 only. GPT figures elsewhere on this
+page are drawn from those vendors' own releases and are cross-vendor, so they
+are looser still.
+
Kimi and GLM are single-sourced. Those two columns appear
+only on the extended variant of the chart; the shared columns are identical
+across every copy, so the numbers are consistent, but the Kimi and GLM rows
+rest on one source.
+
Independent history says be careful. The one held-out
+check on the V4-Pro preview, from NIST/CAISI, put it closer to GPT-5
+and roughly eight months behind the frontier, below its self-reported
+position. No equivalent 0813 evaluation exists yet.
+
+
+
+
The launch table
+
Higher is better. HLE is shown as without-tools / with-tools. Every figure
+is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not
+report it.
+
+
+
+
Benchmark
+
V4-Pro 0813
+
V4-Flash 0731
+
Kimi K3
+
GLM-5.2
+
Opus 4.8
+
Fable 5
+
+
+
Terminal-Bench 2.1
87.9
82.7
88.3
81.6
85.0
88.0
+
DeepSWE
62.7
54.4
67.5
46.2
58.0
70.0
+
Toolathlon-Verified
74.1
70.3
76.5
59.9
76.2
77.9
+
CyberGym
83.3
76.7
80.0
–
78.3
83.1
+
NL2Repo
61.5
54.2
–
48.9
69.7
–
+
AutomationBench
31.8
25.1
30.8
12.9
27.2
29.1
+
DSBench-FullStack
71.1
68.7
63.0
51.8
71.6
77.2
+
DSBench-Hard
67.2
59.6
73.7
54.5
71.7
68.3
+
Agents' Last Exam
25.7
25.2
24.5
23.9
25.7
–
+
HLE (no / tools)
42.7/60.0
37.8/51.5
43.5/56.0
40.5/54.7
49.8/57.9
53.3/63.0
+
+
+
+
The honest read of the row-by-row: V4-Pro is within a point of Fable
+5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on
+Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi
+K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in
+the cluster, not clear of it. On knowledge without tools (HLE) the closed
+models still lead. Flash trails Pro across the board but stays remarkably close
+for a fifth of the price, which is the whole point of the next two sections.
+
+
What GA changed
+
The model ID stayed deepseek-v4-pro and the
+rate card did not move. What moved is the
+checkpoint. Against the April preview, DeepSeek's own chart shows the gains
+landing almost entirely in the agentic and SWE suites:
+
+
+
Benchmark
Preview
GA 0813
Delta
+
+
DeepSWE
12.8
62.7
+49.9
+
DSBench-Hard
31.1
67.2
+36.1
+
CyberGym
52.7
83.3
+30.6
+
DSBench-FullStack
41.8
71.1
+29.3
+
NL2Repo
38.5
61.5
+23.0
+
Toolathlon-Verified
55.9
74.1
+18.2
+
Terminal-Bench 2.1
72.1
87.9
+15.8
+
+
+
+
A near five-fold jump on DeepSWE is not a new base model. Gains shaped like
+this come from agent post-training: better tool-error recovery, better context
+policy, reinforcement on the harness the benchmark runs in. That is real and it
+is useful, and it is also exactly the kind of gain that can be harness-specific,
+which is why the caveats above matter and why the preview numbers are the floor,
+not these.
+
+
The kill line (斩杀线)
+
The term comes from Chinese gaming: the 斩杀线 is the
+health threshold below which a target can be executed outright. Applied to
+models, the idea is that DeepSeek's price-to-capability ratio draws a line, and
+any model that is both weaker and more expensive falls below it and has
+no reason to be chosen. The lever is price, and the gap is not small:
+
+
+
Per 1M tokens
v4-flash
v4-pro
GPT-5.6 Sol
pro is cheaper by
+
+
input, cache miss
$0.14
$0.435
$5.00
~11.5x
+
input, cache hit
$0.0028
$0.003625
$0.50
~138x
+
output
$0.28
$0.87
$30.00
~34.5x
+
+
+
+
DeepSeek prices are the published USD rate card of
+2026-08-02, unchanged at GA; they are a conversion of the RMB card
+(¥3 / ¥0.025 / ¥6 per 1M for pro) at one consistent rate. GPT-5.6
+Sol prices are from OpenAI's own listing. A broad
+DeepSeek repricing is announced but has no date, so it is not applied here.
+
The sober version matters as much as the slogan. The kill line is real for
+the middle of the market: a model that costs more than V4-Pro and
+scores below it on the table above is hard to justify, and that is most of the
+field. It is not real for the frontier. On the hardest
+multi-step work the strongest closed models still finish in fewer turns and
+need less steering, and per-attempt reliability is a thing you can measure in
+wall-clock and interventions, not just in dollars. DeepSeek does not have to win
+per attempt to win per dollar, and it does not have to win per dollar to lose
+the one task where getting it right the first time is the whole job.
+
+
What it means in practice
+
The economics only pay off if the workflow is built for them. Three moves,
+each of which this CLI is shaped to support:
+
+
Route by role. Use deepseek-v4-pro for
+planning, ambiguous changes, security review and recovery; let
+deepseek-v4-flash do bounded implementations and parallel work.
+ds chat -m deepseek-v4-pro and the
+Anthropic remap make the switch one flag.
+
Structure prompts for the cache. A cached input token
+costs about 1/50th of an uncached one. Keep the system prompt, tool schemas,
+repository map and durable instructions in an identical prefix and put the
+volatile part last; ds usage reports what the cache saved so you
+can see whether it is working. This is where the 138x cache-hit number turns
+from a table cell into a bill.
+
Measure successful-task cost, not per-token price. A
+cheaper model that retries five times can cost more than a dearer one that lands
+first. The ledger stores exact token counts
+per call, so you can price a whole task under any rate card, including the one
+that has not been announced yet.
+
+
+
Sources
+
The launch table and the preview deltas are transcribed from DeepSeek's
+official agent-benchmark chart, published on the
+Models & Pricing page and
+circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2
+columns was the widest copy available. Rate-card figures are DeepSeek's own and
+OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review
+of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts;
+cross-check a live leaderboard before betting on a single cell.
+
+
+
+
+
+
+
diff --git a/site/build.py b/site/build.py
index 2fd1055..ef7ddc9 100644
--- a/site/build.py
+++ b/site/build.py
@@ -66,6 +66,7 @@ def _release() -> str:
("commands/", "commands"),
("formats/", "formats"),
("cost/", "cost"),
+ ("bench/", "bench"),
("news/", "news"),
("agents/", "agents"),
("playground/", "playground"),
@@ -1184,6 +1185,180 @@ def jstr(s):
""",
))
+PAGES.append(dict(
+ slug="bench/",
+ crumb="bench",
+ title="DeepSeek V4-Pro benchmarks vs GPT, Claude, Kimi and GLM, and the kill line",
+ description="How deepseek-v4-pro (GA 0813) and v4-flash score against Kimi K3, GLM-5.2, Claude Opus 4.8 and Fable 5 on the agent suites, what the GA checkpoint changed, and the economics idea behind the DeepSeek kill line.",
+ keywords="deepseek v4 pro benchmarks, deepseek v4 pro vs claude, deepseek vs gpt-5.6, deepseek vs kimi k3, deepseek vs glm, deepseek 斩杀线, deepseek kill line, deepseek v4 pro 0813, deepseek agent benchmarks, terminal bench deepseek, deepswe, toolathlon",
+ jsonld=faq([
+ ("How does DeepSeek-V4-Pro score against other models on agent benchmarks?",
+ "On DeepSeek's own launch-day chart (2026-08-12), V4-Pro-0813 scores Terminal-Bench 2.1 87.9, DeepSWE 62.7, Toolathlon-Verified 74.1, CyberGym 83.3, HLE-with-tools 60.0 and AutomationBench 31.8. It sits in the same cluster as Kimi K3, Claude Fable 5 and Opus 4.8: within a point of Fable 5 on Terminal-Bench (88.0) and CyberGym (83.1), ahead of Opus 4.8 on several execution suites, but behind Kimi K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. These are vendor numbers from one harness and are not yet independently reproduced."),
+ ("What did the V4-Pro GA (0813) checkpoint change over the preview?",
+ "The model ID and the rate card did not change; the checkpoint did. Against the April V4-Pro preview, DeepSeek's chart shows large agentic gains: DeepSWE 12.8 to 62.7, DSBench-Hard 31.1 to 67.2, CyberGym 52.7 to 83.3, Terminal-Bench 2.1 72.1 to 87.9, Toolathlon 55.9 to 74.1. Jumps that size point to agent post-training and better tool-error handling rather than a new base model, and none of them are independently verified yet."),
+ ("What is the DeepSeek kill line (斩杀线)?",
+ "The kill line is a community idea that DeepSeek's price-to-capability ratio sets a threshold that removes the reason to exist for any model that is both weaker and more expensive. V4-Pro is roughly 11x cheaper than GPT-5.6 Sol on cache-miss input and 34x cheaper on output, and about 138x cheaper on cache-hit input. It does not kill the frontier: the strongest closed models still finish the hardest tasks in fewer turns. It kills the middle, where a model costs more and does less."),
+ ("Is DeepSeek-V4-Pro better than Claude or GPT for coding agents?",
+ "Per attempt, the strongest closed models remain more reliable on the hardest multi-step tasks and usually need less steering. Per dollar, V4-Pro changes the arithmetic: its cheap cached input makes repeated review, parallel workers and long tool loops affordable in a way per-token-stronger models are not. The practical answer is to route by role, run an internal bake-off, and measure successful-task cost, not per-token price."),
+ ]),
+ body="""
+
Benchmarks
+
Where deepseek-v4-pro and deepseek-v4-flash
+sit against the field, what the 0813 checkpoint changed, and the economics
+argument the Chinese community calls the 斩杀线, the
+kill line. The numbers below are DeepSeek's own launch-day figures unless
+marked otherwise; read the caveats first.
+
+
+read this before quoting a number
+
+
Vendor numbers, one harness. The table is DeepSeek's own
+agent-benchmark chart, published 2026-08-12 with the GA. A score is a
+model-and-harness result, not a model-only one, and no independent same-harness
+run of 0813 exists yet. Treat these as the claim, not the verdict.
+
No GPT-5.6 column. DeepSeek's chart compares against Kimi
+K3, GLM-5.2, Claude Opus 4.8 and Fable 5 only. GPT figures elsewhere on this
+page are drawn from those vendors' own releases and are cross-vendor, so they
+are looser still.
+
Kimi and GLM are single-sourced. Those two columns appear
+only on the extended variant of the chart; the shared columns are identical
+across every copy, so the numbers are consistent, but the Kimi and GLM rows
+rest on one source.
+
Independent history says be careful. The one held-out
+check on the V4-Pro preview, from NIST/CAISI, put it closer to GPT-5
+and roughly eight months behind the frontier, below its self-reported
+position. No equivalent 0813 evaluation exists yet.
+
+
+
+
The launch table
+
Higher is better. HLE is shown as without-tools / with-tools. Every figure
+is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not
+report it.
+
+
+
+
Benchmark
+
V4-Pro 0813
+
V4-Flash 0731
+
Kimi K3
+
GLM-5.2
+
Opus 4.8
+
Fable 5
+
+
+
Terminal-Bench 2.1
87.9
82.7
88.3
81.6
85.0
88.0
+
DeepSWE
62.7
54.4
67.5
46.2
58.0
70.0
+
Toolathlon-Verified
74.1
70.3
76.5
59.9
76.2
77.9
+
CyberGym
83.3
76.7
80.0
–
78.3
83.1
+
NL2Repo
61.5
54.2
–
48.9
69.7
–
+
AutomationBench
31.8
25.1
30.8
12.9
27.2
29.1
+
DSBench-FullStack
71.1
68.7
63.0
51.8
71.6
77.2
+
DSBench-Hard
67.2
59.6
73.7
54.5
71.7
68.3
+
Agents' Last Exam
25.7
25.2
24.5
23.9
25.7
–
+
HLE (no / tools)
42.7/60.0
37.8/51.5
43.5/56.0
40.5/54.7
49.8/57.9
53.3/63.0
+
+
+
+
The honest read of the row-by-row: V4-Pro is within a point of Fable
+5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on
+Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi
+K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in
+the cluster, not clear of it. On knowledge without tools (HLE) the closed
+models still lead. Flash trails Pro across the board but stays remarkably close
+for a fifth of the price, which is the whole point of the next two sections.
+
+
What GA changed
+
The model ID stayed deepseek-v4-pro and the
+rate card did not move. What moved is the
+checkpoint. Against the April preview, DeepSeek's own chart shows the gains
+landing almost entirely in the agentic and SWE suites:
+
+
+
Benchmark
Preview
GA 0813
Delta
+
+
DeepSWE
12.8
62.7
+49.9
+
DSBench-Hard
31.1
67.2
+36.1
+
CyberGym
52.7
83.3
+30.6
+
DSBench-FullStack
41.8
71.1
+29.3
+
NL2Repo
38.5
61.5
+23.0
+
Toolathlon-Verified
55.9
74.1
+18.2
+
Terminal-Bench 2.1
72.1
87.9
+15.8
+
+
+
+
A near five-fold jump on DeepSWE is not a new base model. Gains shaped like
+this come from agent post-training: better tool-error recovery, better context
+policy, reinforcement on the harness the benchmark runs in. That is real and it
+is useful, and it is also exactly the kind of gain that can be harness-specific,
+which is why the caveats above matter and why the preview numbers are the floor,
+not these.
+
+
The kill line (斩杀线)
+
The term comes from Chinese gaming: the 斩杀线 is the
+health threshold below which a target can be executed outright. Applied to
+models, the idea is that DeepSeek's price-to-capability ratio draws a line, and
+any model that is both weaker and more expensive falls below it and has
+no reason to be chosen. The lever is price, and the gap is not small:
+
+
+
Per 1M tokens
v4-flash
v4-pro
GPT-5.6 Sol
pro is cheaper by
+
+
input, cache miss
$0.14
$0.435
$5.00
~11.5x
+
input, cache hit
$0.0028
$0.003625
$0.50
~138x
+
output
$0.28
$0.87
$30.00
~34.5x
+
+
+
+
DeepSeek prices are the published USD rate card of
+2026-08-02, unchanged at GA; they are a conversion of the RMB card
+(¥3 / ¥0.025 / ¥6 per 1M for pro) at one consistent rate. GPT-5.6
+Sol prices are from OpenAI's own listing. A broad
+DeepSeek repricing is announced but has no date, so it is not applied here.
+
The sober version matters as much as the slogan. The kill line is real for
+the middle of the market: a model that costs more than V4-Pro and
+scores below it on the table above is hard to justify, and that is most of the
+field. It is not real for the frontier. On the hardest
+multi-step work the strongest closed models still finish in fewer turns and
+need less steering, and per-attempt reliability is a thing you can measure in
+wall-clock and interventions, not just in dollars. DeepSeek does not have to win
+per attempt to win per dollar, and it does not have to win per dollar to lose
+the one task where getting it right the first time is the whole job.
+
+
What it means in practice
+
The economics only pay off if the workflow is built for them. Three moves,
+each of which this CLI is shaped to support:
+
+
Route by role. Use deepseek-v4-pro for
+planning, ambiguous changes, security review and recovery; let
+deepseek-v4-flash do bounded implementations and parallel work.
+ds chat -m deepseek-v4-pro and the
+Anthropic remap make the switch one flag.
+
Structure prompts for the cache. A cached input token
+costs about 1/50th of an uncached one. Keep the system prompt, tool schemas,
+repository map and durable instructions in an identical prefix and put the
+volatile part last; ds usage reports what the cache saved so you
+can see whether it is working. This is where the 138x cache-hit number turns
+from a table cell into a bill.
+
Measure successful-task cost, not per-token price. A
+cheaper model that retries five times can cost more than a dearer one that lands
+first. The ledger stores exact token counts
+per call, so you can price a whole task under any rate card, including the one
+that has not been announced yet.
+
+
+
Sources
+
The launch table and the preview deltas are transcribed from DeepSeek's
+official agent-benchmark chart, published on the
+Models & Pricing page and
+circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2
+columns was the widest copy available. Rate-card figures are DeepSeek's own and
+OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review
+of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts;
+cross-check a live leaderboard before betting on a single cell.