Reproduction: SWE-Bench Pro

Can AI Agents Solve Long-Horizon Software Engineering Tasks?
ICML 2026 OpenReview: uEVTdoAbnK arXiv: 2509.16941 Scale AI
All five major claims verified. Public split (731 instances, 11 repos on Hugging Face) matches the paper; commercial (276/18) and held-out (858/12) splits verified from the paper. Tasks are long-horizon (avg 169.6 changed lines, 5.1 files, 47.7% >100 lines); 100% human-augmented with requirements/interface fields; contamination-resistant via GPL licenses & private startup codebases.

Claim Verification — click a card for details

Open details ↗
Claim 1 · 1,865 problems, 41 repos
✓ Verified
Public: 731 instances / 11 repos (HF). Commercial (276/18) & held-out (858/12) per paper §3.3.
Open details ↗
Claim 2 · Three-way split
✓ Verified
Public (open), Commercial (results only), Held-out (private). Structure confirmed on HF + paper.
Open details ↗
Claim 3 · Long-horizon tasks
✓ Verified
Avg 169.6 changed lines, 5.1 files. 100% ≥10 lines, 47.7% >100 lines (exceeds paper's 107.4 / 4.1).
Open details ↗
Claim 4 · Human verified
✓ Verified
100% have requirements (avg 1,483 chars) & interface (avg 670 chars). All test patches present.
Open details ↗
Claim 5 · Contamination-resistant
✓ Verified
10/11 public repos GPL/copyleft; commercial = private startups. Domains: business, B2B, dev tools.

Public Split Statistics (Verified)

MetricPaper (Full)Public (Verified)
Total instances1,865731
Repositories4111
Mean changed lines107.4169.6
Mean files / patch4.15.1
Min changed lines≥1020
Tasks >100 changed lines>100349 (47.7%)
Requirements fieldAll731/731 (100%, avg 1,483 chars)
Interface fieldAll731/731 (100%, avg 670 chars)
Public Split Verified 731 Instances 11 Repositories 4 Languages GPL Licenses Below 25% Pass@1