Researchers released Android Bench 2.0, moving benchmarks from single-function bug fixes to complex multi-day Long-Horizon Tasks (LHT). The suite evaluates coding agents on building apps from scratch, complex multi-library version migrations, and cross-platform verification.
Key Takeaways
- ✓Deprecates trivial unit bug benchmarks in favor of multi-day full-system app synthesis
- ✓Tests real-world architectural scaffolding, breaking library upgrades, and UI regression
- ✓Top frontier agents currently achieve only 31.4% success, exposing long-horizon bottlenecks
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.