Claude Opus 5 Leads in Agent-Building Test, But All Models Struggle Below 25%
Sierra's new Hyper-τ-bench benchmark measures how well AI systems can autonomously construct other AI agents. The results reveal significant gaps in information gathering and design exploration.