SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

Research official 2 src. ~1 min

A new benchmark of 170 rigorously curated code-refactoring tasks spanning seven programming languages, built to fix quality and test-suite issues found in prior SWE-Bench variants. The best evaluated coding agent resolves only 41.2% of tasks, showing current agents still struggle with large-scale multilingual refactoring.

Why it matters

Top-voted paper on HuggingFace Daily Papers for 2026-08-11 with 59 upvotes; a harder, cleaner benchmark than existing SWE-Bench variants for measuring real coding-agent progress.

Importance: 2/5

Top-ranked HuggingFace Daily Paper of the day, below the 100-upvote bump threshold.

Sources