File size: 1,765 Bytes
bd3727b fbc261c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 | # Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
[HF paper page](https://huggingface.co/papers/2509.16941)
## Pages
| Page |
| --- |
| [Executive summary](#/executive-summary) |
| [Claim 1: SWE-Bench Pro comprises 1,865 problems sourced from 41 actively maintained software repositories (abstract only).](#/claim-1-swe-bench-pro-comprises-1-865-problems-sourced-from-41-actively-maintained-software-repositories-abstract-only) |
| [Claim 2: SWE-Bench Pro splits into public (11 repos), held-out (12 repos), and commercial (18 proprietary repos) sets (abstract only).](#/claim-2-swe-bench-pro-splits-into-public-11-repos-held-out-12-repos-and-commercial-18-proprietary-repos-sets-abstract-only) |
| [Claim 3: Tasks are long-horizon, potentially requiring hours to days for professional engineers and often spanning multiple files (abstract only).](#/claim-3-tasks-are-long-horizon-potentially-requiring-hours-to-days-for-professional-engineers-and-often-spanning-multiple-files-abstract-only) |
| [Claim 4: All SWE-Bench Pro tasks underwent human verification for adequate resolution context (abstract only).](#/claim-4-all-swe-bench-pro-tasks-underwent-human-verification-for-adequate-resolution-context-abstract-only) |
| [Claim 5: SWE-Bench Pro is a contamination-resistant testbed spanning business applications, B2B services, and developer tools (abstract only).](#/claim-5-swe-bench-pro-is-a-contamination-resistant-testbed-spanning-business-applications-b2b-services-and-developer-tools-abstract-only) |
| [Conclusion](#/conclusion) |
| [Deep Verification: Contamination Resistance, Agent Solve Rate, and Eval Repo Evidence](#/deep-verification-contamination-resistance-agent-solve-rate-and-eval-repo-evidence) |
|