CONNECT WITH US
AI & Deeptech

AI & Deeptech

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

TheNewStack (LLM) - ArtificialIntelligence logo

Published on

Add as a preferred source on Google
Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems that answer questions, process refunds, change bookings, and interact with company systems. But while AI is increasingly central to building these systems, humans still play a major steering role: setting goals, supplying context, choosing architectures, reviewing decisions, and testing the result.

Which raises a more interesting question: what happens when an agent is asked to build another agent entirely on its own?

Created and open-sourced in early September by Sierra, the enterprise AI agent company co-founded by tech veteran and current OpenAI board chairman Bret Taylor, Hyper-𝜏-bench builds on the original 𝜏-bench benchmark it introduced back in 2024. But while 𝜏-bench focused on measuring how well a finished agent could interact with users, use tools, and follow company policies, Hyper-𝜏-bench takes it a level up: it evaluates how well an AI developer agent can build that agent in the first place.

Today, most agents are built with the help of other agents like Sierra's Ghostwriter. Yesterday, Sierra open-sourced hyper-𝜏-bench (published as 𝜏^𝜏-bench), a new long horizon agent evaluation that measures how well models can not only act as an agent, but construct one.…

In a research paper published on September 4, Sierra researchers tested six combinations of AI model and coding harness. Those included Anthropic models running in Claude Code, OpenAI models in Codex, and Moonshot AI’s Kimi K3 running in both Kimi Code and the open-source OpenCode .


Source link

Disclaimer

We strive to uphold the highest ethical standards in all of our reporting and coverage. We TheMorningPulse.fyi want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It's possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.