On this page
For two years, the way the whole market buys AI has boiled down to one question: which model? GPT or Claude or Gemini. A new leaderboard every month, whoever is on top wins, and every vendor-selection meeting revolves around that one name.
Last week, one number made that question look silly.
Same model. Same test. Running bare: 30 points. Wrapped in a layer of software: 100 points. No retraining, no model swap, no extra data.
The headlines everywhere drew a neat conclusion: the wrapper is the hero, the model is a supporting act. That conclusion is being read backwards. And reading it backwards will make you buy the wrong thing again, just in a new direction.
What happened on August 21
On August 21, 2026, NVIDIA announced on its official technical blog that its agent system AVO (Agentic Variation Operators), first released at the end of March 2026, had reached a perfect 100.00 RHAE on ARC-AGI-3.
A few terms need unpacking before we go on.
ARC-AGI-3 is a benchmark released by the ARC Prize Foundation on March 25, 2026. Unlike ordinary question-and-answer tests, it is interactive: the AI is dropped into turn-based, game-like environments with no instructions, no rules and no stated goal. It has to discover the mechanics itself, work out what winning looks like, and carry what it learned into the next level. Humans solve 100%. According to ARC Prize's launch announcement, the most advanced AI systems at the time scored under 1%.
Harness is all the software around the model: the tools it is allowed to call, how memory is managed, the planning loop, and the supervision rules. The model is the engine; the harness is the rest of the car.
NVIDIA's results:
| Configuration | Result |
|---|---|
| Claude Opus 5 running bare | ~30% (the highest among bare models) |
| VISTA + Claude Opus 5 | 100.00, 7,542 actions |
| AVO + Claude Opus 5 | 100.00, 6,624 actions (~12% fewer) |
AVO completed all 183 levels across 25 environments. The architecture has three parts: persistent memory that carries results and logs from previous attempts, a supervisor that watches the whole run to spot when the agent is going in circles, and an observe–plan–act–evaluate loop.
Adel El Hallak, NVIDIA's VP of product, told TechCrunch that the supervisor acts "almost like a CEO", nudging the agent back on track when it drifts.
That's the news. Now for the part that got left behind.
Three facts that force a re-read
One: NVIDIA was not the first to hit 100.
On August 5, 2026, sixteen days earlier, an MIT research group (Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He) published VISTA, which also scored 100.00 on the same 25 public games, also with Claude Opus 5. GPT-5.6 Sol in the same harness scored 98.27.
So AVO's real contribution is not "reaching 100". It is reaching 100 about 12% more efficiently. A genuine engineering advance, but not the milestone the headlines imply.
Two: the team that got there first publicly pushes back on this reading.
This is the detail I find most important, and almost nobody quotes it. The VISTA authors themselves write that "the inherent capabilities of the model are key to this success", and that their harness is only "a simple yet effective way to elicit those capabilities". They stress that VISTA was designed to be minimal, deliberately avoiding complex engineering.
In other words: the team that built the first harness to hit a perfect score credits the model. The harness doesn't create capability. It unlocks capability that was already there.
Three: ARC Prize itself refuses to count these scores.
Under ARC Prize's testing policy, official scores are measured with the model running as a direct input-to-output predictor, "with no agent harness and no client-side tools". Every harnessed result goes onto a separate, clearly labeled community leaderboard.
That isn't conservatism. It is the definition: a truly general system shouldn't need an outsider to build bespoke scaffolding for each type of problem.
NVIDIA says this plainly in its own announcement, in the part most news coverage cut: the perfect score applies only to the public set of 25 environments, and is "not a result on the semi-private or fully private evaluation sets". The benchmark is split three ways precisely so nobody can optimize for the problems they will be graded on. The hardest part hasn't been touched yet.
So what is actually true
Not "the harness matters more than the model". Not "the model is still everything" either.
What's true is this: the variance lives in the harness; the capability ceiling lives in the model.
The model decides how far you can go. The harness decides how far you actually get. The 30 → 100 gap is not software magic; it is the distance between the capability a model has and the capability you manage to extract from it.
This wasn't discovered last week. In May 2026, a research team published a paper with a title that reads almost like a reprimand: Stop Comparing LLM Agents Without Disclosing the Harness. Their finding: changing the harness alone shifts scores by as much as or more than the gap between different models.
If that's right, and last week's result is expensive proof, then every model leaderboard you have ever used to pick a vendor has been measuring a variable that isn't the most important one.
What this means for your business
You don't run ARC-AGI-3. But if your business uses AI or is about to, this pattern applies almost unchanged.
First: stop asking "which model?" and start asking "who builds the rest?" When a vendor pitches you an AI solution, "do you use GPT or Claude?" barely distinguishes anyone, because everyone can buy the same model at the same price. The questions that do distinguish them: what does your system remember between runs? When it gets something wrong, what catches it and fixes it? Without concrete answers, you're buying an engine with no car.
Second: the most valuable component is the one that catches errors. The most instructive technical detail in both VISTA and AVO isn't the model; it's the supervisor, the piece that notices when the agent is stuck and pulls it back. Business is the same: the value isn't that the AI can answer, it's that you know when it answers wrong before your customer does. A chatbot with no checking layer and no escape route to a real person is a brand risk you pay for monthly.
Third: be wary of a 100 on home turf. NVIDIA hit a perfect score on the public set and said so clearly. Many vendors won't be that clear. When someone demos an AI system running flawlessly, the only question worth asking is: has this been run on data you had never seen before? A demo on a prepared set of examples and production on your real data are two different problems, exactly like the gap between ARC-AGI-3's public and private sets.
Fourth: your advantage is in the part you build. Anyone can rent a model. What can't be bought is your own process, your own data, and the control rules you distill from your own work. That's why, when I integrate AI into websites and workflows, I start with the problem and the control flow, pick the model last, and design so the model can be swapped, because the leaderboard will change many more times.
What would change my mind
I could be wrong, and I want to be specific about where.
If AVO, or any harness, achieves comparable scores on ARC-AGI-3's private set, environments nobody has ever seen, then the "harness only elicits existing capability" argument gets significantly weaker and NVIDIA's thesis gets much stronger. That is the clean test, and it hasn't happened yet.
To be fair to NVIDIA: their announcement contains a less-discussed result that I find more convincing than the benchmark score. Over seven days, AVO autonomously tried more than 500 optimization directions, produced 40 committed kernel versions, and achieved performance 10.5% higher than FlashAttention-4 on a DGX B200 system. That is real engineering work on a real problem, not a game.
But that only reinforces what I want you to take away: the value came from an operating architecture that ran persistently for seven days, not from picking the right name on a leaderboard.
On August 21, the world read that the harness had dethroned the model. Read more carefully, the story is calmer and far more useful: the model you're paying for is probably good enough already. What you haven't built yet is the rest.
If you'd rather build that rest properly than chase the next leaderboard, let's talk.
Sources
- NVIDIA AVO Reaches 100% on ARC-AGI-3 — NVIDIA Technical Blog, Aug 21, 2026
- Nvidia just showed that the harness, not the AI model, is now the real hero — TechCrunch, Aug 21, 2026
- VISTA: A Visual Harness for Reasoning in an Interactive World — MIT, Aug 5, 2026
- Announcing ARC-AGI-3 — ARC Prize Foundation, Mar 25, 2026
- ARC-AGI Testing Policy — ARC Prize Foundation
- Stop Comparing LLM Agents Without Disclosing the Harness — arXiv:2605.23950, May 26, 2026

