On this page
The most common way to pick an AI tool right now is to look at the scoreboard. Whichever model tops the benchmark wins. Whichever model is cheapest per token wins.
That approach is measuring the wrong thing. DeepSeek just showed the whole market why, and they did it almost by accident.
On August 13, 2026, DeepSeek released DeepSeek Harness (the dsh command),
an open-source "harness" shipped alongside its V4-Pro model. A harness is the
software wrapped around a model: it reads files, runs commands, calls tools,
remembers the session, and decides when to allow or block an action
(VentureBeat).
The repo picked up roughly 50,000 GitHub stars in its first 12 hours and passed
186,000 stars and more than 20,000 forks within 10 days
(KDnuggets).
Numbers like that led plenty of people to one conclusion: we finally have a free, open-source Claude Code. That conclusion is wrong on almost every count. And the places where it goes wrong are the lessons worth paying for.
Belief one: a benchmark score is the model's score
DeepSeek reported that V4-Pro-0813 scored 87.9 on Terminal Bench 2.1, 74.1 on Toolathlon-Verified, 71.1 on DSBench-FullStack, and 67.2 on DSBench-Hard.
But VentureBeat pointed to a detail sitting in DeepSeek's own announcement: on the public Code Agent tasks, V4-Pro-0813 was evaluated using DeepSeek Harness in "minimal mode." In other words, those numbers measure the model running inside one specific execution environment, not the model on its own (VentureBeat).
I'm not saying DeepSeek did anything wrong. They disclosed it plainly. I'm saying that once the score depends on the harness, "which model is better" is no longer a complete question. The same model in a different harness can produce different results at a different cost.
A preregistered study from PointFive says exactly this. The authors write that coding-agent efficiency "cannot be characterized by token counts or model pricing alone," because cost and success rate depend jointly on prompt content, reasoning effort, harness policy, the model, and task difficulty. They propose measuring cost per successful task, treating token counts as a measurement of the agent's trajectory rather than a target to optimize (arXiv:2608.01347).
You won't find a "the harness accounts for X% of cost" figure in this post. The summary I first read contained numbers that don't appear in the actual paper. I went back to the real abstract and threw them out. The authors' conclusion is strong enough without embroidered statistics.
Belief two: open source means cheap
The harness itself is MIT-licensed and free to download. But a harness is just the frame. You still pay for the model running inside it, and here DeepSeek had news that got far less airtime than the star count.
From August 16, 2026, DeepSeek raised its API prices. VentureBeat, citing Reuters, reports increases ranging "from 50% to more than 1,100%" depending on the model, token type, and time of day. For V4-Pro, the price per million tokens (input / output) changed like this (VentureBeat):
| V4-Pro (USD per 1M tokens) | Input | Output | 1M in + 1M out |
|---|---|---|---|
| Before Aug 16 | 0.435 | 0.87 | 1.305 |
| From Aug 16, off-peak | 0.66 | 1.98 | 2.64 |
| From Aug 16, peak hours | 1.32 | 3.96 | 5.28 |
The last column is my own sum of the two before it. For a standard workload, V4-Pro is now roughly twice as expensive off-peak and four times as expensive at peak. It may still undercut plenty of alternatives, but "free" and "unconditionally cheap" are no longer true.
The lesson isn't specific to DeepSeek. A list price is a variable, not a constant. If your workflow is welded to one provider, you have no way to react when the price moves. That is exactly why the Harness's "everything is a plugin" design is worth noticing: according to KDnuggets, it works with around 40 model providers, and sub-agents can be handed off to another vendor's model (KDnuggets). The ability to switch providers without tearing everything down is a real asset. It doesn't show up on any leaderboard.
Belief three: a sandbox means you're safe
This is the part I found most instructive, because it happened in a seriously designed project.
On paper, the Harness has OS-level isolation: bwrap and Landlock on Linux, Seatbelt on macOS, restricted ACL tokens on Windows. KDnuggets notes that its defaults "fail closed" rather than fail open (KDnuggets). Those are the right choices.
And yet, according to The Hacker News, version 0.1.1-rc.2 and earlier were hit
by CVE-2026-82533, CVSS 9.4. The mechanism is worth remembering
(The Hacker News):
- The Harness ships a local web interface. That interface had no authentication and trusted the Host header without checking where the connection came from.
- An agent running inside the sandbox could call that interface with a single
shell command and switch its own session to
danger-full-accessmode, turning off both the sandbox and the approval step. - The sandbox at that point only covered file writes, so a compromised agent could write outside the working directory. Anyone who could reach the interface could also pull every stored conversation without a key.
The timeline, per the same source: the community described the escape path on
August 13–14, right after launch. The fix 0.1.2-alpha.1 landed on GitHub on
August 27, the first patched npm release (0.1.2-alpha.2) on August 30, and
0.1.2-rc.1 on September 3. OX Research reported the bug; VulnCheck assigned
the CVE.
Read that mechanism again. The lock existed. But the switch that controls the lock was within reach of the very thing it was supposed to lock in. No sandbox is worth much if the agent inside it can reach the "open door" button.
To be fair, the project warns outright that its sandbox and approval step "do not guarantee isolation or prevent damage," and the README states clearly that this is a preview with "compatibility-breaking changes" (GitHub). They didn't promise more than they had. The problem is that people read star counts, not warnings.
So what should a buyer of AI services ask?
You don't need to know what bwrap is. You need to ask the right five
questions, whether you're talking to an in-house team, a freelancer, or an
agency:
- What does each result cost? Not the price per token. Ask for the total cost of completing a real task, retries included. That is precisely the metric the study above recommends.
- What environment was this score measured in? If a benchmark figure was measured inside a specific harness, ask whether your setup looks anything like it.
- What can the agent do outside the scope of the job? Write outside its folder, make network calls, read secrets from environment variables. Ask about what isn't blocked, not just what is.
- Who controls the safety switch? If the agent has any way to escalate its own permissions, every other layer of protection is decoration.
- If your model provider doubles its prices next week, how long would it take to switch? A good answer is "an afternoon, because the model is just a config option."
Before you trust a scoreboard, ask what it was measured with. Before you trust a sandbox, ask who holds its switch.
What I don't know, and therefore didn't write
I haven't seen any independent measurement of how dsh actually performs
against other tools on the same task, so this post does not conclude which
tool is better. Most of the comparisons I found were personal opinions or small
trial runs, not enough to lean on.
The prompt-injection security research on the Harness that I found (arXiv:2608.16393) didn't give me any figure clear enough to quote either. CVE-2026-82533 is different: it is fully documented, patched, and has a timeline, so I used it.
On maturity, KDnuggets concludes that the Harness isn't yet a daily driver and can't yet replace workflows that already run reliably; it suits infrastructure builders who want to customize session storage or the sandbox backend (KDnuggets). VentureBeat likewise describes it as a model-agnostic option that is "not yet a full replacement" for the broader developer experience of established products (VentureBeat).
What this means for your business
The AI race is shifting from "which model is smartest" to "which system turns that intelligence into safe outcomes at a predictable cost." Models are becoming interchangeable. The harness, the permissions, the approval flow, and the ability to switch providers are the parts you'll live with for years.
If you're considering putting AI agents into your website, customer support, or internal workflows, start with the five questions above, not the leaderboard. I work through exactly this on AI integration projects, and you can get in touch if you'd like a second pair of eyes on a specific architecture.
Sources
- DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices, VentureBeat
- DeepSeek Harness Flaw Let AI Agents Disable Their Own File Sandbox Without Approval, The Hacker News
- What I've Learned About DeepSeek Harness, KDnuggets
- deepseek-ai/deepseek-harness, GitHub
- Prompt-Induced Waste in Coding Agents, Weinberger & Hozez, PointFive (arXiv:2608.01347)
- Security Assessment of DeepSeek Harness with A.I.G (arXiv:2608.16393)

