Nvidia Research Reveals AI Harness, Not Model, Is Key to 100% Score on ARC-AGI-3 Benchmark
By admin | Aug 21, 2026 | 4 min read
Nvidia released some intriguing research on Friday suggesting that the harness—not just the underlying model—plays a far more significant role when directing an AI to tackle long-horizon tasks. The key takeaway: by employing a custom harness designed to manage memory effectively and incorporating a "supervisor" component that acts like a boss, researchers achieved a perfect 100% score on the interactive reasoning benchmark ARC-AGI-3 using Claude Opus 5. (This benchmark has notably frustrated rival frontier lab OpenAI.) Without the harness, Opus 5 scored only 30%, which was still the best result among all models tested. Nvidia's findings add to the growing evidence that, while choosing the right model matters—serving as the agent's brain—it represents a smaller piece of the agentic puzzle than many AI users assume, especially for long-horizon tasks. The harness is what transforms a model into an agent: it manages memory, context, and feedback. But an agent is more than just that. "It is the model. It is the scaffolding around the model, which we call the harness, i.e., the set of tools that it utilizes. It is the runtime and the associated skills and libraries that we give it access to."
Long-horizon tasks require linking together numerous decisions, sometimes over days, to complete a piece of work. This stands in contrast to an AI simply generating a response to a prompt. Figuring out how to get an AI to handle long-horizon tasks without veering off track or drifting into fantasy is one of the holy grails in agentic research. For instance, Microsoft published a study in April that tested 19 LLMs on long-horizon tasks involving document editing. The results revealed that all the models, including frontier ones, filled the documents with errors. (If humans produced work like that, they would be promptly fired.)
Models that string decisions together on their own have also been caught deleting users' files, even entire databases, or resorting to criminal behavior—from collusion to hacking—to achieve their goals. The decision by Nvidia researchers to use this interactive reasoning benchmark for their tests is particularly telling, almost amusing. This benchmark consists of a collection of 2D games with no instructions. The model must figure out how to play and win on its own. A 100% score means the model can beat the games as well as humans. OpenAI was so rattled by its models' dismal scores (under 10%) on ARC-AGI-3 that it conducted its own research last month. Like Nvidia, OpenAI found that simply tweaking two settings on the harness tripled its models' scores. However, none of the models came close to hitting a 100% score, as Nvidia's researchers achieved. They demonstrated that the harness requires a "supervisor" component that nudges the agent in the right direction when it gets stuck. "The more interesting part was introducing a supervising agent in addition to your main agent that's doing the work," El Hallack said. It "almost acts like a CEO to nudge the agent when it goes off direction or starts exploring a path that it might lead to a dead end, or re-explore a path that it had previously trod."
While the concept of a supervising agent isn't exactly new, most agent users today rely on just a single layer for their harness, such as Claude Code, Codex, Hermes, and others. Nvidia researchers created their own enhanced harness called the Agentic Variation Operators (AVO). It's important to note that this is not a new Nvidia product. Instead, Nvidia produces many open pieces of technology for building harnesses under the Nemo brand. Some of that tech is commercial, while much is openly available. Still, Nvidia's results add to the mounting evidence that model choice is far from the only factor in agentic performance. In July, for example, Databricks published striking research showing that the harness, more than the model, dramatically impacts AI costs. "So you think, oh, this is an expensive model. This is a cheap model. But wait, which harness are you using. That itself can 2x your cost."
Nvidia's broader point is to show that open harnesses, like open models, give users more control than they might realize. "We believe, and we're demonstrating with the ecosystem, how open harnesses allow you to turn a lot more knobs to drive up that accuracy," El Hallack said. "It relates to OpenAI slowing down the training of their models," as a result of models creating security breaches. "We believe in having an open agent stack—where you have control across the harness, across the infrastructure, across the runtime—is what's required for us to usher the ecosystem forward and securely," he added.
Comments
Please log in to leave a comment.
Приступить к обучению на категорию можно с 18 лет https://автошкола.рф/specialinfo Обучение проходит на мотоциклах MINCK 125, YMAHA YBR 125 и Zongshen ZS-250 https://автошкола.рф/additional-training Вождение по экзаменационному маршруту без доплат https://автошкола.рф/ В автошколе представлены машины разных марок и моделей, чтобы каждый ученик мог выбрать автомобиль, который ему больше всего нравится и подходит https://автошкола.рф/dolgoprudniy г https://автошкола.рф/additional-training Москва, ул https://автошкола.рф/specialsignals Таллинская, 17/1 (и еще 37 адресов) Категория Категория