The SWE-bench benchmark team launched SWE-bench Pro, expanding from Python to Rust, Go, C++, and TypeScript. Featuring 1,200 curated real-world issues from high-traffic open-source repositories, it provides a unified industry standard for multi-language autonomous agent software engineering evaluation.
Key Takeaways
- ✓Includes 1,200 real-world repository repair tasks across Rust, Go, C++, and TypeScript.
- ✓Applies strict compiler verification and integration test matrices to eliminate false passes.
- ✓First leaderboard standings show Claude 3.7 Sonnet and DeepSeek-R1 leading multi-language solve rates.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.