RealSWE: coding-agent benchmark rebuilt around real, casual user requests
The authors quantify the gap between SWE-bench-style tasks and real usage: 88% of real prompts carry only a bare problem statement (versus 7% of benchmark problems), and 87% are casually written versus 94% formal in benchmarks. RealSWE derives 381 multi-variant task families from SWE-bench Verified and Pro to test compositional behavior under realistic request styles.
Why it matters
Argues current coding-agent scores overstate real-world readiness by measuring on curated, information-rich issues rather than how users actually write.
Importance: 2/5
Notable evaluation paper, below the 100-upvote bar