SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
This paper introduces SWE-Bench ProMax, a new benchmark for testing AI coding agents on large-scale multilingual code refactoring tasks, which is designed to be more realistic and challenging than existing benchmarks. Practitioners can care about this paper because it provides a rigorous evaluation of current AI coding agents' capabilities on a more representative set of tasks.