Yashp2003 commited on
Commit
1a352b5
·
verified ·
1 Parent(s): c1e6a11

Update logbook: Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Browse files
logbook.json CHANGED
@@ -10,7 +10,7 @@
10
  "icml2026-repro",
11
  "paper-uEVTdoAbnK"
12
  ],
13
- "updated_at": "2026-07-20T04:47:16+00:00",
14
  "root": {
15
  "slug": "index",
16
  "title": "Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?",
@@ -60,6 +60,6 @@
60
  }
61
  ]
62
  },
63
- "agent_view_tokens": 8290,
64
- "revision": "1784522836649974700"
65
  }
 
10
  "icml2026-repro",
11
  "paper-uEVTdoAbnK"
12
  ],
13
+ "updated_at": "2026-07-20T04:59:16+00:00",
14
  "root": {
15
  "slug": "index",
16
  "title": "Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?",
 
60
  }
61
  ]
62
  },
63
+ "agent_view_tokens": 7676,
64
+ "revision": "1784523556359408410"
65
  }
pages/claim-1-swe-bench-pro-comprises-1-865-problems-from-41-actively-maintained-repositories/page.md CHANGED
@@ -71,9 +71,3 @@ The public split exactly matches the paper's stated numbers (731 instances, 11 r
71
  ---
72
 
73
  **Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
74
-
75
-
76
-
77
- ---
78
-
79
- **Independent corroboration of the full benchmark (claims 1–2).** Beyond the public HF split, the official eval repo [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os) ships a `run_scripts/` directory with **1,000 instance run-script folders** across 9 distinct repositories (ansible, openlibrary, flipt, qutebrowser, teleport, element-web, vuls, NodeBB, tutanota) — direct evidence that the evaluation harness is built for a benchmark far larger than the 731 public instances, consistent with the paper’s stated 1,865 problems / 41 repos (Table 1, §3.3). The held-out (858 / 12 repos) and commercial (276 / 18 proprietary repos) splits are **deliberately private** (paper §3.3: “results-only” and “private”), so their per-instance counts cannot be independently recomputed; the public split (731 / 11) is the only fully reproducible portion. Verdict on the *full* totals is therefore **inconclusive** pending access to the non-public splits, while the public split is **verified**.
 
71
  ---
72
 
73
  **Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
 
 
 
 
 
 
pages/claim-2-swe-bench-pro-splits-into-public-11-repos-held-out-12-repos-and-commercial-18-proprietary-repos-sets/page.md CHANGED
@@ -45,9 +45,3 @@ The public split is fully verified (731 instances, 11 repos on HF). Commercial a
45
  ---
46
 
47
  **Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
48
-
49
-
50
-
51
- ---
52
-
53
- **Independent corroboration of the full benchmark (claims 1–2).** Beyond the public HF split, the official eval repo [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os) ships a `run_scripts/` directory with **1,000 instance run-script folders** across 9 distinct repositories (ansible, openlibrary, flipt, qutebrowser, teleport, element-web, vuls, NodeBB, tutanota) — direct evidence that the evaluation harness is built for a benchmark far larger than the 731 public instances, consistent with the paper’s stated 1,865 problems / 41 repos (Table 1, §3.3). The held-out (858 / 12 repos) and commercial (276 / 18 proprietary repos) splits are **deliberately private** (paper §3.3: “results-only” and “private”), so their per-instance counts cannot be independently recomputed; the public split (731 / 11) is the only fully reproducible portion. Verdict on the *full* totals is therefore **inconclusive** pending access to the non-public splits, while the public split is **verified**.
 
45
  ---
46
 
47
  **Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).