diff --git a/site/404.html b/site/404.html index 1a00365..4295f1e 100644 --- a/site/404.html +++ b/site/404.html @@ -73,6 +73,7 @@ commands formats cost + bench news agents playground diff --git a/site/agents/index.html b/site/agents/index.html index f77d861..eb1ac08 100644 --- a/site/agents/index.html +++ b/site/agents/index.html @@ -68,6 +68,7 @@ commands formats cost + bench news agents playground diff --git a/site/bench/index.html b/site/bench/index.html new file mode 100644 index 0000000..b4dafbc --- /dev/null +++ b/site/bench/index.html @@ -0,0 +1,275 @@ + + + + + +DeepSeek V4-Pro benchmarks vs GPT, Claude, Kimi and GLM, and the kill line + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ + +
+
+ + >thevibeworks/deepseek-cli + + +
+ + +
+
+
+
+
+ + +

Benchmarks

+

Where deepseek-v4-pro and deepseek-v4-flash +sit against the field, what the 0813 checkpoint changed, and the economics +argument the Chinese community calls the 斩杀线, the +kill line. The numbers below are DeepSeek's own launch-day figures unless +marked otherwise; read the caveats first.

+ +
+read this before quoting a number +
    +
  • Vendor numbers, one harness. The table is DeepSeek's own +agent-benchmark chart, published 2026-08-12 with the GA. A score is a +model-and-harness result, not a model-only one, and no independent same-harness +run of 0813 exists yet. Treat these as the claim, not the verdict.
  • +
  • No GPT-5.6 column. DeepSeek's chart compares against Kimi +K3, GLM-5.2, Claude Opus 4.8 and Fable 5 only. GPT figures elsewhere on this +page are drawn from those vendors' own releases and are cross-vendor, so they +are looser still.
  • +
  • Kimi and GLM are single-sourced. Those two columns appear +only on the extended variant of the chart; the shared columns are identical +across every copy, so the numbers are consistent, but the Kimi and GLM rows +rest on one source.
  • +
  • Independent history says be careful. The one held-out +check on the V4-Pro preview, from NIST/CAISI, put it closer to GPT-5 +and roughly eight months behind the frontier, below its self-reported +position. No equivalent 0813 evaluation exists yet.
  • +
+
+ +

The launch table

+

Higher is better. HLE is shown as without-tools / with-tools. Every figure +is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not +report it.

+
+ + + + + + + + + + + + + + + + + + + + + + +
BenchmarkV4-Pro 0813V4-Flash 0731Kimi K3GLM-5.2Opus 4.8Fable 5
Terminal-Bench 2.187.982.788.381.685.088.0
DeepSWE62.754.467.546.258.070.0
Toolathlon-Verified74.170.376.559.976.277.9
CyberGym83.376.780.078.383.1
NL2Repo61.554.248.969.7
AutomationBench31.825.130.812.927.229.1
DSBench-FullStack71.168.763.051.871.677.2
DSBench-Hard67.259.673.754.571.768.3
Agents' Last Exam25.725.224.523.925.7
HLE (no / tools)42.7/60.037.8/51.543.5/56.040.5/54.749.8/57.953.3/63.0
+
+

The honest read of the row-by-row: V4-Pro is within a point of Fable +5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on +Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi +K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in +the cluster, not clear of it. On knowledge without tools (HLE) the closed +models still lead. Flash trails Pro across the board but stays remarkably close +for a fifth of the price, which is the whole point of the next two sections.

+ +

What GA changed

+

The model ID stayed deepseek-v4-pro and the +rate card did not move. What moved is the +checkpoint. Against the April preview, DeepSeek's own chart shows the gains +landing almost entirely in the agentic and SWE suites:

+
+ + + + + + + + + + + +
BenchmarkPreviewGA 0813Delta
DeepSWE12.862.7+49.9
DSBench-Hard31.167.2+36.1
CyberGym52.783.3+30.6
DSBench-FullStack41.871.1+29.3
NL2Repo38.561.5+23.0
Toolathlon-Verified55.974.1+18.2
Terminal-Bench 2.172.187.9+15.8
+
+

A near five-fold jump on DeepSWE is not a new base model. Gains shaped like +this come from agent post-training: better tool-error recovery, better context +policy, reinforcement on the harness the benchmark runs in. That is real and it +is useful, and it is also exactly the kind of gain that can be harness-specific, +which is why the caveats above matter and why the preview numbers are the floor, +not these.

+ +

The kill line (斩杀线)

+

The term comes from Chinese gaming: the 斩杀线 is the +health threshold below which a target can be executed outright. Applied to +models, the idea is that DeepSeek's price-to-capability ratio draws a line, and +any model that is both weaker and more expensive falls below it and has +no reason to be chosen. The lever is price, and the gap is not small:

+
+ + + + + + + +
Per 1M tokensv4-flashv4-proGPT-5.6 Solpro is cheaper by
input, cache miss$0.14$0.435$5.00~11.5x
input, cache hit$0.0028$0.003625$0.50~138x
output$0.28$0.87$30.00~34.5x
+
+

DeepSeek prices are the published USD rate card of +2026-08-02, unchanged at GA; they are a conversion of the RMB card +(¥3 / ¥0.025 / ¥6 per 1M for pro) at one consistent rate. GPT-5.6 +Sol prices are from OpenAI's own listing. A broad +DeepSeek repricing is announced but has no date, so it is not applied here.

+

The sober version matters as much as the slogan. The kill line is real for +the middle of the market: a model that costs more than V4-Pro and +scores below it on the table above is hard to justify, and that is most of the +field. It is not real for the frontier. On the hardest +multi-step work the strongest closed models still finish in fewer turns and +need less steering, and per-attempt reliability is a thing you can measure in +wall-clock and interventions, not just in dollars. DeepSeek does not have to win +per attempt to win per dollar, and it does not have to win per dollar to lose +the one task where getting it right the first time is the whole job.

+ +

What it means in practice

+

The economics only pay off if the workflow is built for them. Three moves, +each of which this CLI is shaped to support:

+ + +

Sources

+

The launch table and the preview deltas are transcribed from DeepSeek's +official agent-benchmark chart, published on the +Models & Pricing page and +circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2 +columns was the widest copy available. Rate-card figures are DeepSeek's own and +OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review +of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts; +cross-check a live leaderboard before betting on a single cell.

+
+
+ + + + + diff --git a/site/build.py b/site/build.py index 2fd1055..ef7ddc9 100644 --- a/site/build.py +++ b/site/build.py @@ -66,6 +66,7 @@ def _release() -> str: ("commands/", "commands"), ("formats/", "formats"), ("cost/", "cost"), + ("bench/", "bench"), ("news/", "news"), ("agents/", "agents"), ("playground/", "playground"), @@ -1184,6 +1185,180 @@ def jstr(s): """, )) +PAGES.append(dict( + slug="bench/", + crumb="bench", + title="DeepSeek V4-Pro benchmarks vs GPT, Claude, Kimi and GLM, and the kill line", + description="How deepseek-v4-pro (GA 0813) and v4-flash score against Kimi K3, GLM-5.2, Claude Opus 4.8 and Fable 5 on the agent suites, what the GA checkpoint changed, and the economics idea behind the DeepSeek kill line.", + keywords="deepseek v4 pro benchmarks, deepseek v4 pro vs claude, deepseek vs gpt-5.6, deepseek vs kimi k3, deepseek vs glm, deepseek 斩杀线, deepseek kill line, deepseek v4 pro 0813, deepseek agent benchmarks, terminal bench deepseek, deepswe, toolathlon", + jsonld=faq([ + ("How does DeepSeek-V4-Pro score against other models on agent benchmarks?", + "On DeepSeek's own launch-day chart (2026-08-12), V4-Pro-0813 scores Terminal-Bench 2.1 87.9, DeepSWE 62.7, Toolathlon-Verified 74.1, CyberGym 83.3, HLE-with-tools 60.0 and AutomationBench 31.8. It sits in the same cluster as Kimi K3, Claude Fable 5 and Opus 4.8: within a point of Fable 5 on Terminal-Bench (88.0) and CyberGym (83.1), ahead of Opus 4.8 on several execution suites, but behind Kimi K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. These are vendor numbers from one harness and are not yet independently reproduced."), + ("What did the V4-Pro GA (0813) checkpoint change over the preview?", + "The model ID and the rate card did not change; the checkpoint did. Against the April V4-Pro preview, DeepSeek's chart shows large agentic gains: DeepSWE 12.8 to 62.7, DSBench-Hard 31.1 to 67.2, CyberGym 52.7 to 83.3, Terminal-Bench 2.1 72.1 to 87.9, Toolathlon 55.9 to 74.1. Jumps that size point to agent post-training and better tool-error handling rather than a new base model, and none of them are independently verified yet."), + ("What is the DeepSeek kill line (斩杀线)?", + "The kill line is a community idea that DeepSeek's price-to-capability ratio sets a threshold that removes the reason to exist for any model that is both weaker and more expensive. V4-Pro is roughly 11x cheaper than GPT-5.6 Sol on cache-miss input and 34x cheaper on output, and about 138x cheaper on cache-hit input. It does not kill the frontier: the strongest closed models still finish the hardest tasks in fewer turns. It kills the middle, where a model costs more and does less."), + ("Is DeepSeek-V4-Pro better than Claude or GPT for coding agents?", + "Per attempt, the strongest closed models remain more reliable on the hardest multi-step tasks and usually need less steering. Per dollar, V4-Pro changes the arithmetic: its cheap cached input makes repeated review, parallel workers and long tool loops affordable in a way per-token-stronger models are not. The practical answer is to route by role, run an internal bake-off, and measure successful-task cost, not per-token price."), + ]), + body=""" +

Benchmarks

+

Where deepseek-v4-pro and deepseek-v4-flash +sit against the field, what the 0813 checkpoint changed, and the economics +argument the Chinese community calls the 斩杀线, the +kill line. The numbers below are DeepSeek's own launch-day figures unless +marked otherwise; read the caveats first.

+ +
+read this before quoting a number + +
+ +

The launch table

+

Higher is better. HLE is shown as without-tools / with-tools. Every figure +is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not +report it.

+
+ + + + + + + + + + + + + + + + + + + + + + +
BenchmarkV4-Pro 0813V4-Flash 0731Kimi K3GLM-5.2Opus 4.8Fable 5
Terminal-Bench 2.187.982.788.381.685.088.0
DeepSWE62.754.467.546.258.070.0
Toolathlon-Verified74.170.376.559.976.277.9
CyberGym83.376.780.078.383.1
NL2Repo61.554.248.969.7
AutomationBench31.825.130.812.927.229.1
DSBench-FullStack71.168.763.051.871.677.2
DSBench-Hard67.259.673.754.571.768.3
Agents' Last Exam25.725.224.523.925.7
HLE (no / tools)42.7/60.037.8/51.543.5/56.040.5/54.749.8/57.953.3/63.0
+
+

The honest read of the row-by-row: V4-Pro is within a point of Fable +5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on +Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi +K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in +the cluster, not clear of it. On knowledge without tools (HLE) the closed +models still lead. Flash trails Pro across the board but stays remarkably close +for a fifth of the price, which is the whole point of the next two sections.

+ +

What GA changed

+

The model ID stayed deepseek-v4-pro and the +rate card did not move. What moved is the +checkpoint. Against the April preview, DeepSeek's own chart shows the gains +landing almost entirely in the agentic and SWE suites:

+
+ + + + + + + + + + + +
BenchmarkPreviewGA 0813Delta
DeepSWE12.862.7+49.9
DSBench-Hard31.167.2+36.1
CyberGym52.783.3+30.6
DSBench-FullStack41.871.1+29.3
NL2Repo38.561.5+23.0
Toolathlon-Verified55.974.1+18.2
Terminal-Bench 2.172.187.9+15.8
+
+

A near five-fold jump on DeepSWE is not a new base model. Gains shaped like +this come from agent post-training: better tool-error recovery, better context +policy, reinforcement on the harness the benchmark runs in. That is real and it +is useful, and it is also exactly the kind of gain that can be harness-specific, +which is why the caveats above matter and why the preview numbers are the floor, +not these.

+ +

The kill line (斩杀线)

+

The term comes from Chinese gaming: the 斩杀线 is the +health threshold below which a target can be executed outright. Applied to +models, the idea is that DeepSeek's price-to-capability ratio draws a line, and +any model that is both weaker and more expensive falls below it and has +no reason to be chosen. The lever is price, and the gap is not small:

+
+ + + + + + + +
Per 1M tokensv4-flashv4-proGPT-5.6 Solpro is cheaper by
input, cache miss$0.14$0.435$5.00~11.5x
input, cache hit$0.0028$0.003625$0.50~138x
output$0.28$0.87$30.00~34.5x
+
+

DeepSeek prices are the published USD rate card of +2026-08-02, unchanged at GA; they are a conversion of the RMB card +(¥3 / ¥0.025 / ¥6 per 1M for pro) at one consistent rate. GPT-5.6 +Sol prices are from OpenAI's own listing. A broad +DeepSeek repricing is announced but has no date, so it is not applied here.

+

The sober version matters as much as the slogan. The kill line is real for +the middle of the market: a model that costs more than V4-Pro and +scores below it on the table above is hard to justify, and that is most of the +field. It is not real for the frontier. On the hardest +multi-step work the strongest closed models still finish in fewer turns and +need less steering, and per-attempt reliability is a thing you can measure in +wall-clock and interventions, not just in dollars. DeepSeek does not have to win +per attempt to win per dollar, and it does not have to win per dollar to lose +the one task where getting it right the first time is the whole job.

+ +

What it means in practice

+

The economics only pay off if the workflow is built for them. Three moves, +each of which this CLI is shaped to support:

+ + +

Sources

+

The launch table and the preview deltas are transcribed from DeepSeek's +official agent-benchmark chart, published on the +Models & Pricing page and +circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2 +columns was the widest copy available. Rate-card figures are DeepSeek's own and +OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review +of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts; +cross-check a live leaderboard before betting on a single cell.

+""", +)) + PAGES.append(dict( slug="news/", crumb="news", diff --git a/site/commands/index.html b/site/commands/index.html index c048a00..d8c7fcf 100644 --- a/site/commands/index.html +++ b/site/commands/index.html @@ -68,6 +68,7 @@ commands formats cost + bench news agents playground diff --git a/site/cost/index.html b/site/cost/index.html index 28cf10f..e2d6825 100644 --- a/site/cost/index.html +++ b/site/cost/index.html @@ -68,6 +68,7 @@ commands formats cost + bench news agents playground @@ -216,7 +217,7 @@

What these numbers are not

a completion, not for bookkeeping.