Researchers released Android Bench 2.0, moving benchmarks from single-function bug fixes to complex multi-day Long-Horizon Tasks (LHT). The suite evaluates coding agents on building apps from scratch, complex multi-library version migrations, and cross-platform verification.

Key Takeaways

  • Deprecates trivial unit bug benchmarks in favor of multi-day full-system app synthesis
  • Tests real-world architectural scaffolding, breaking library upgrades, and UI regression
  • Top frontier agents currently achieve only 31.4% success, exposing long-horizon bottlenecks
ADSponsored