RealSWE: coding-agent benchmark rebuilt around real, casual user requests

Research official 2 src. ~1 min

The authors quantify the gap between SWE-bench-style tasks and real usage: 88% of real prompts carry only a bare problem statement (versus 7% of benchmark problems), and 87% are casually written versus 94% formal in benchmarks. RealSWE derives 381 multi-variant task families from SWE-bench Verified and Pro to test compositional behavior under realistic request styles.

Why it matters

Argues current coding-agent scores overstate real-world readiness by measuring on curated, information-rich issues rather than how users actually write.

Importance: 2/5

Notable evaluation paper, below the 100-upvote bar

Sources

secondary RealSWE — Hugging Face Daily Papers (syndicated from arXiv)