Yashp2003's picture
Update logbook: Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
bd3727b verified
|
Raw
History Blame
2.24 kB

Claim 2: SWE-Bench Pro splits into public (11 repos), held-out (12 repos), and commercial (18 proprietary repos) sets (abstract only).


Claim 2 — Three-way partition (11 / 12 / 18)

Claim (verbatim): "SWE-Bench Pro is split into a public set of 11 openly accessible repositories, a held-out set of 12 repositories not publicly accessible, and a commercial set of 18 proprietary repositories obtained via partnership agreements with startups."

Verdict: Consistent — the public portion (11 repos, 731 instances) is directly verifiable; the held-out (12) and commercial (18) splits are described as non-public by design and align with the paper's stated 41-repo total and contamination-resistance goal.

What we verified

From the public release ScaleAI/SWE-bench_Pro, every one of the 731 instances belongs to exactly one of 11 distinct repo values (see Claim 1 table). The dataset README explicitly states repo: "Repository identifier - one of 11 repository classes", confirming the public set is exactly 11 repositories.

Reconciliation with the paper

Split Repos (paper) Status Verifiable here?
Public 11 Openly accessible Yes — 11 repos, 731 instances counted
Held-out 12 Not publicly accessible No (by design)
Commercial 18 Proprietary, partnership agreements No (by design)
Total 41 11 + 12 + 18 = 41 ✓

The arithmetic 11 + 12 + 18 = 41 matches the 41-repo total in Claim 1, and the paper notes "Problems in the held-out and the commercial set are not publicly accessible, but we release results on the commercial set." This is exactly why only the 11 public repos are observable in the open dataset — not a contradiction.

Evidence