Nvidia's research suggests that when it comes to complex tasks, the software wrapper around an AI model – the harness – plays a far greater role than previously thought. By tweaking their custom harness, researchers were able to get Claude Opus 5 to score a perfect 100% on the interactive reasoning benchmark ARC-AGI-3, whereas other models struggled even with basic tasks.
The choice of harness can significantly impact an AI's performance and efficiency in long-horizon tasks. For instance, Microsoft found that all its large language models failed miserably when tested on document editing tasks, producing errors akin to those a human might make under pressure. But Nvidia’s researchers showed that with the right tweaks, their models could navigate these challenges.
The concept of a supervising agent within the harness is gaining traction as a crucial component for guiding AI agents towards goal-oriented behavior. This reflects a shift in focus from just choosing the best model to also refining the tools and environments that interact with it. As Nvidia’s Adel El Hallak explains, ‘the world interprets an agent almost as an API of the model,’ but it is the harness that truly brings this to life.
The findings highlight a new paradigm in AI development: open harnesses are key to achieving better performance and cost efficiency. Databricks’ research, for example, shows that using the wrong harness can double costs even with the same model. The emphasis on open harnesses reflects a shift towards giving users more control over their AI systems.







