Specific releases Real-SWE benchmark built from private enterprise codebases

Specific

Tools official + media 2 src. ~1 min

Specific (YC F25) launched Real-SWE, a coding benchmark of tasks licensed from real private company codebases, scored over model-and-harness pairs like Claude Code and Codex CLI with pass@1 averaged over eight runs. Fable 5.1 via Claude Code leads at 38.8%, ahead of GPT-6 Astra via Codex CLI at 33.8% and Gemini 3.8 Flash at 31.2%; 'missed requirement' is the most common failure mode and cost per rollout did not predict accuracy.

Why it matters

First benchmark natively out of training distribution, exposing how much SWE-bench-style scores flatter frontier coding agents.

Importance: 3/5

First out-of-distribution enterprise coding benchmark; strong community traction

Sources