An Tran Solutions
An Tran Solutions
Back to Blog

You Don't Need a Stronger Model. You Need a Better Harness.

July 17, 20268 min readby An Tran
On this page

The question I've heard most over the past year, from clients and fellow engineers alike, always has the same shape: "Which model should we use?"

Claude or GPT. Is the new release better than the last one. What does the leaderboard say this week. As if picking the right cell in a comparison table were the whole job.

I asked exactly that question for a long time. And I was wrong. Not wrong on a detail. Wrong because I was staring at the less important variable.

This isn't a post knocking models. A good model still beats a bad one. But there's another layer, sitting right under our noses, that moves the outcome more, and almost nobody puts it in the comparison table.

The number that changed my mind

In May 2026, a research group published Harness-Bench (submitted May 27, 2026), a study aimed squarely at this question. The design is brutally simple: take 6 different harnesses, cross them with 8 backend models (Claude Opus/Sonnet, Gemini, Qwen, GLM, Kimi, GPT-5.4, DeepSeek), run them on the same 106 tasks, and collect more than 5,000 executions.

Same tasks. Same lineup of models. Only the harness wrapped around them changes.

The aggregate result: the weakest harness scored 52.4%. The strongest scored 76.2%.

A 23.8-percentage-point gap, produced by something that isn't the model.

Sit with that number for a second. 23.8 points is a bigger gap than most of what you see between two consecutive model generations. It's bigger than the gaps the whole industry argues about every time a new release ships. And it comes from the layer most teams treat as "plumbing, just wire it up."

The authors' conclusion is blunt: agent capability should be understood as a property of a model operating inside an execution system, not a property of the base model alone.

Worse: the harness can flip the rankings

If it were only about higher or lower scores, we could live with it. The deeper problem shows up in another paper, submitted May 7, 2026, with a title that says it all: Stop Comparing LLM Agents Without Disclosing the Harness.

The central finding: the variance caused by the harness can be substantially larger than the variance caused by the model. And the consequence is genuinely uncomfortable: the ranking between models can reverse depending on which harness you use to measure them.

Read that again. Model A beats model B in one harness, and loses to that same model B in another. Same tasks.

So every time you read a line like "model X outperforms model Y," it only means something if it tells you which harness was used to measure it. Without that, a leaderboard doesn't tell you which model is better. It tells you which model-harness combination got measured. Those are two completely different things.

The authors propose something that should have been the default long ago: mandatory harness disclosure before comparing any agents.

Agent = Model + Harness. You're not buying a model. You're operating a combination. And that combination can swing by more than 20 percentage points without changing a single thing about the model.

Not just theory: Epoch AI proved it on themselves

At this point someone will say: academic research is nice, but the real world is different.

So look at Epoch AI, an independent evaluation organization that scores SWE-bench Verified. They have no incentive to inflate this story.

In February 2026, Epoch upgraded its scaffold, runtime environment and token limits to v2.0.0. No model changed. Only the layer around the models.

The result: model scores improved significantly, significantly enough that Epoch had to state that results from before v2.0.0 are no longer directly comparable, and its default chart now only shows data from v2.0.0 onward.

A professional measurement organization had to reset its baseline just because it fixed the harness. That isn't a minor technical footnote. It's an admission that the wrapper decides the number.

So what exactly is a harness?

In short: the harness is everything in an agent except the model.

The model does exactly one thing: it takes in a pile of text and predicts the next pile of text. Everything else is harness:

  • Context: what goes into the context window, in what order, and what gets cut when it runs out of room.
  • Tools: which files the agent can read, which shell commands and APIs it can call, and how clearly those tools are described.
  • Permissions: where a human has to approve, and where the agent can just run.
  • Memory: what the agent remembers between sessions.
  • Verification loop: how the agent knows it did the job right: tests, linters, type-checks, builds.
  • Recovery: how it rolls back when something breaks.

Harness-Bench also identifies the most common failure mode, and it's worth thinking about: "execution-alignment failures", where the model's reasoning drifts away from the actual tool feedback, the real state of the workspace, and the verifiable evidence.

The model didn't get dumber. It just lost touch with reality, because the harness didn't feed reality back to it fast enough. Switching to a smarter model doesn't fix that. It just makes the agent wrong more fluently.

The strongest teams have moved their effort to the harness

The most convincing evidence isn't in a paper. It's in how the teams furthest ahead are spending their time.

OpenAI published its harness engineering experiment with Codex in February 2026. Over 5 months, a team that started with just 3 engineers shipped a real product with ~1,500 merged pull requests and roughly a million lines of code, and not one line was written by a human. All of it, the logic, tests, CI configuration, documentation and observability tooling, was written by Codex. OpenAI estimates they did it in about 1/10 of the time it would have taken by hand. When the team grew to 7 people, throughput went up rather than plateauing.

But the most valuable detail isn't those numbers. It's what the engineers' job turned into.

They barely wrote code anymore. They described tasks, let the agent run, and let it open the PR. And when the agent stumbled, they didn't swap the model. They treated the stumble as a signal of what the harness was missing: a missing tool, a missing guardrail, missing documentation. Then they added exactly that to the repo, and had Codex write the fix itself.

The philosophy fits in one line: humans steer, agents execute.

Anthropic arrived at the same place from a different direction. In Effective harnesses for long-running agents (November 26, 2025), they describe the core problem of long-running agents with a very apt image: it's like a software project staffed by engineers working in shifts, where each new shift arrives with no memory of the last one.

The context window is finite. The work is longer than the window. So their solution lives entirely in the harness: an initializer agent sets up the environment and baseline context, then hands off to a coding agent that makes incremental progress, with claude-progress.txt, git history, a feature list and mandatory test gates as the handoff mechanism between shifts.

Not one line of that solution is about changing the model.

And no, this won't disappear on its own

The most reasonable objection: models will eventually absorb the harness, right? Wait a few generations and it's solved.

On July 8, 2026, Lilian Weng published a review of 35 research papers on harness engineering, and answered that question directly: even if many harness improvements end up internalized into the model, the need to specify goals and context will not go away.

A model may learn to plan, check and backtrack on its own. But it will never know by itself what "done" means in your project, where your internal systems live, what your industry's regulations forbid, or what your customers actually want. That part is yours. Permanently.

Where I got it wrong

To be fair to myself and to you: for a long time, every time an agent botched a job, my first reflex was to blame the model. Then I'd go hunting for a newer one, wait for the next release, read leaderboards like sports scores.

The mistake wasn't picking the wrong model. It was that I treated every failure as a limit of the model, instead of a hole in the harness.

Almost every time, when I dug in, I found the same things: the agent was missing information it should have had on hand; it had no way to verify what it had just done; it didn't know what this project considered "done." Every one of those was something I could fix in an afternoon, and none of them needed a new model.

A leaderboard is a lot easier to read than sitting down to design a verification loop. That's exactly why it's so tempting.

What to fix, and in what order

If your team uses agents, for writing code, handling content or serving customers, here's the order I recommend, ranked by return on effort:

  1. Build the verification loop first. The agent has to know whether it's right or wrong without you looking: tests, type-checks, linters, builds. Without that loop, everything else is guesswork. This is the highest-return investment, every time.

  2. Write down your definition of "done." As a file, in the repo, not in someone's head. An agent can't read your company culture.

  3. Trim context instead of stuffing it. A bigger window won't save you. Give it the right thing at the right time. Stuffing everything in is the fastest way to produce an execution-alignment failure.

  4. Design the handoff between sessions. If the work is longer than one context window, and real work always is, you need the equivalent of claude-progress.txt: state written to disk, clean commits, green tests before the shift ends.

  5. Every time the agent stumbles, fix the harness. Don't swap the model. This is the hardest discipline to keep, and it's what separates teams that improve from teams that go in circles. Every stumble is a gap pointing straight at itself.

The cost of looking in the wrong place

This is the part of the story worth the most money.

Swapping models is an expense. Building a harness is an asset.

Every time you upgrade the model hoping the errors go away, you pay for a one-off bump, and next month you pay again. Nothing accumulates. The verification loop you build today is the opposite: it keeps its full value when the model changes generations, and it makes every model you plug in perform better. Harness-Bench shows 23.8 percentage points sitting in that layer. That's the part you own.

And here's the painful bit: a competitor who is building a harness gets a little faster every month, even if they use exactly the same model as you. A team that just waits for the next model stands still between releases, then wonders why others ship with the same tools and they can't.

Three engineers at OpenAI didn't ship a million lines of code in 5 months because they had a secret model. They had Codex, which anyone can buy. What they built, and you haven't yet, is the layer around it.

The right question was never "which model should we use?" The right question is: when your agent gets it wrong, does it have any way to find out on its own?

If the answer is "no," a stronger model won't save you. It'll just help you be wrong faster.

If you're thinking about putting agents into a real workflow and want a straight read on where to start, even if the answer is "not yet," I'm happy to talk it through. More on my AI work is on the AI integration page.


Sources:

Related articles

An Tran Solutions
September 30, 20268 min read

Spec-Driven Development Isn't Documentation. It's a Governance System for AI-Written Code

From vibe coding to spec-driven development: why longer specs won't save you, and a four-layer governance framework (constitution, risk-tiered specs, executable acceptance, control gates) that keeps AI-written code under your control. With data from Veracode, Thoughtworks and Martin Fowler.