Claim 5: SWE-Bench Pro is a contamination-resistant testbed spanning business applications, B2B services, and developer tools (abstract only).
Claim 5 β Contamination-resistant, domain-diverse testbed
Claim (verbatim from abstract): "SWE-BENCH PRO provides a contamination-resistant testbed that more faithfully captures the complexity and diversity of real-world software development."
Verdict: VERIFIED β zero instance ID overlap vs original SWE-Bench, 0.12% n-gram overlap (only boilerplate), containerized execution, 3 enterprise domains, 4 languages.
Contamination resistance β quantitative evidence
We loaded the original princeton-nlp/SWE-bench test split (2,294 instances) and performed cross-analysis:
- Instance ID overlap: 0 β zero shared instance IDs between SWE-Bench Pro and original SWE-Bench
- N-gram overlap: 0.12% β only 146 / 125,772 5-grams overlap, all generic issue-report boilerplate ("users should be able to", "steps to reproduce", etc.)
- Per-instance similarity: average max similarity 10.9%, maximum 19.5%. No task above 50% similarity
- Containerized execution: every instance carries pinned
dockerhub_tag+base_commitfor isolated resolution
These are direct quantitative measurements confirming SWE-Bench Pro uses entirely different tasks β not relabeled or re-split versions of existing SWE-Bench problems.
Domain diversity (public set, 11 repos)
The 11 repos map onto the three stated categories:
| Domain | Repos | Instances | % |
|---|---|---|---|
| Business applications | tutanota, protonmail, openlibrary, NodeBB | 220 | 30.1% |
| B2B services | flipt, teleport, navidrome, vuls | 280 | 38.3% |
| Developer tools | ansible, element-web, qutebrowser | 231 | 31.6% |
Language diversity (public set)
| Language | Instances | % |
|---|---|---|
| Go | 280 | 38.3% |
| Python | 266 | 36.4% |
| JavaScript | 165 | 22.6% |
| TypeScript | 20 | 2.7% |
Four production languages vs the original SWE-Bench's Python-only scope.
Structural contamination-resistance mechanisms
- Non-public splits: held-out (12 repos) and commercial (18 repos) are deliberately not publicly accessible β models cannot have trained on their solutions
- Containerized execution:
dockerhub_tagpins a reproducible, isolated build environment - Frozen state:
base_commit+before_repo_set_cmd+selected_test_files_to_runfreeze the exact pre-patch state
Evidence
- Dataset: ScaleAI/SWE-bench_Pro
- Eval repo: scaleapi/SWE-bench_Pro-os
- Results:
claim_analysis_results.jsonβ zero ID overlap, 0.12% n-gram overlap - HF Job: Yashp2003/6a5cc999bee6ee1cf4ed1174 β deep verification (contamination analysis)