# Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? [HF paper page](https://huggingface.co/papers/2509.16941) ## Pages | Page | | --- | | [Executive summary](#/executive-summary) | | [Claim 1: SWE-Bench Pro comprises 1,865 problems from 41 actively maintained repositories.](#/claim-1-swe-bench-pro-comprises-1-865-problems-from-41-actively-maintained-repositories) | | [Claim 2: SWE-Bench Pro splits into public (11 repos), held-out (12 repos), and commercial (18 proprietary repos) sets.](#/claim-2-swe-bench-pro-splits-into-public-11-repos-held-out-12-repos-and-commercial-18-proprietary-repos-sets) | | [Claim 3: SWE-Bench Pro tasks are long-horizon, requiring hours to days for professionals and multi-file patches.](#/claim-3-swe-bench-pro-tasks-are-long-horizon-requiring-hours-to-days-for-professionals-and-multi-file-patches) | | [Claim 4: All SWE-Bench Pro tasks underwent human verification for adequate context.](#/claim-4-all-swe-bench-pro-tasks-underwent-human-verification-for-adequate-context) | | [Claim 5: SWE-Bench Pro is a contamination-resistant testbed spanning business apps, B2B services, and dev tools.](#/claim-5-swe-bench-pro-is-a-contamination-resistant-testbed-spanning-business-apps-b2b-services-and-dev-tools) | | [Conclusion](#/conclusion) |