Install
The New Stack is a media platform for the people who build and manage software the world relies on. We provide context and explanation of at-scale technologies to advance knowledge and create conversations through our coverage of modern architectures, components of the software development life cycle, and operations to
- 153articles · 30d
- 16+ hour agolatest article
- Aug 14, 2026earliest in window
- 96%with images
- 87avg words
- Science & Technology 140
- Software Dev. 101
- Computers & Electronics 96
- News 38
- Software 22
- Science & Nature 12
- Economy, Business & Finance 9
- Finance & Business 9
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.
3+ day, 10+ hour ago (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....
AI agent evaluations are part of the product
1+ week, 1+ day ago (886+ words) Move beyond simple AI demos. Build repeatable evaluation systems, test execution paths, and enforce strict release gates for AI agents....
Anthropic's Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.
1+ week, 5+ day ago (527+ words) Anthropic tasked Claude with fixing all 10 benchmarked alignment failures itself. Developers say the real lesson is the process, not the score....
Most coding agent benchmarks skip large-scale refactoring. Not this one.
3+ week, 1+ day ago (411+ words) New SWE-Bench ProMax reveals where coding agents can’t measure up. Experts weigh in on why....