Specific releases Real-SWE benchmark built from private enterprise codebases
Specific
Specific (YC F25) launched Real-SWE, a coding benchmark of tasks licensed from real private company codebases, scored over model-and-harness pairs like Claude Code and Codex CLI with pass@1 averaged over eight runs. Fable 5.1 via Claude Code leads at 38.8%, ahead of GPT-6 Astra via Codex CLI at 33.8% and Gemini 3.8 Flash at 31.2%; 'missed requirement' is the most common failure mode and cost per rollout did not predict accuracy.
Why it matters
First benchmark natively out of training distribution, exposing how much SWE-bench-style scores flatter frontier coding agents.
Importance: 3/5
First out-of-distribution enterprise coding benchmark; strong community traction
Sources
official
Real-SWE Benchmark — Specific Labs