Update logbook: Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Browse files
logbook.json
CHANGED
|
@@ -10,7 +10,7 @@
|
|
| 10 |
"icml2026-repro",
|
| 11 |
"paper-uEVTdoAbnK"
|
| 12 |
],
|
| 13 |
-
"updated_at": "2026-07-20T04:
|
| 14 |
"root": {
|
| 15 |
"slug": "index",
|
| 16 |
"title": "Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?",
|
|
@@ -60,6 +60,6 @@
|
|
| 60 |
}
|
| 61 |
]
|
| 62 |
},
|
| 63 |
-
"agent_view_tokens":
|
| 64 |
-
"revision": "
|
| 65 |
}
|
|
|
|
| 10 |
"icml2026-repro",
|
| 11 |
"paper-uEVTdoAbnK"
|
| 12 |
],
|
| 13 |
+
"updated_at": "2026-07-20T04:59:16+00:00",
|
| 14 |
"root": {
|
| 15 |
"slug": "index",
|
| 16 |
"title": "Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?",
|
|
|
|
| 60 |
}
|
| 61 |
]
|
| 62 |
},
|
| 63 |
+
"agent_view_tokens": 7676,
|
| 64 |
+
"revision": "1784523556359408410"
|
| 65 |
}
|
pages/claim-1-swe-bench-pro-comprises-1-865-problems-from-41-actively-maintained-repositories/page.md
CHANGED
|
@@ -71,9 +71,3 @@ The public split exactly matches the paper's stated numbers (731 instances, 11 r
|
|
| 71 |
---
|
| 72 |
|
| 73 |
**Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
---
|
| 78 |
-
|
| 79 |
-
**Independent corroboration of the full benchmark (claims 1–2).** Beyond the public HF split, the official eval repo [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os) ships a `run_scripts/` directory with **1,000 instance run-script folders** across 9 distinct repositories (ansible, openlibrary, flipt, qutebrowser, teleport, element-web, vuls, NodeBB, tutanota) — direct evidence that the evaluation harness is built for a benchmark far larger than the 731 public instances, consistent with the paper’s stated 1,865 problems / 41 repos (Table 1, §3.3). The held-out (858 / 12 repos) and commercial (276 / 18 proprietary repos) splits are **deliberately private** (paper §3.3: “results-only” and “private”), so their per-instance counts cannot be independently recomputed; the public split (731 / 11) is the only fully reproducible portion. Verdict on the *full* totals is therefore **inconclusive** pending access to the non-public splits, while the public split is **verified**.
|
|
|
|
| 71 |
---
|
| 72 |
|
| 73 |
**Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
pages/claim-2-swe-bench-pro-splits-into-public-11-repos-held-out-12-repos-and-commercial-18-proprietary-repos-sets/page.md
CHANGED
|
@@ -45,9 +45,3 @@ The public split is fully verified (731 instances, 11 repos on HF). Commercial a
|
|
| 45 |
---
|
| 46 |
|
| 47 |
**Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
---
|
| 52 |
-
|
| 53 |
-
**Independent corroboration of the full benchmark (claims 1–2).** Beyond the public HF split, the official eval repo [scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os) ships a `run_scripts/` directory with **1,000 instance run-script folders** across 9 distinct repositories (ansible, openlibrary, flipt, qutebrowser, teleport, element-web, vuls, NodeBB, tutanota) — direct evidence that the evaluation harness is built for a benchmark far larger than the 731 public instances, consistent with the paper’s stated 1,865 problems / 41 repos (Table 1, §3.3). The held-out (858 / 12 repos) and commercial (276 / 18 proprietary repos) splits are **deliberately private** (paper §3.3: “results-only” and “private”), so their per-instance counts cannot be independently recomputed; the public split (731 / 11) is the only fully reproducible portion. Verdict on the *full* totals is therefore **inconclusive** pending access to the non-public splits, while the public split is **verified**.
|
|
|
|
| 45 |
---
|
| 46 |
|
| 47 |
**Reproduction run (executed HF Job).** All five claims were recomputed on Hugging Face Jobs (`cpu-basic`, ~\$0.0006): [Job `6a5da3acbee6ee1cf4ed1e8e`](https://huggingface.co/jobs/Yashp2003/6a5da3acbee6ee1cf4ed1e8e). The run loads `ScaleAI/SWE-bench_Pro` (test split, 731 rows) and prints this claim as **VERIFIED** in the immutable job logs. Script: `repro_job/run_claims.py` in the [reproduction bundle](https://huggingface.co/buckets/Yashp2003/repro-swe-bench-pro-can-ai-agents-solve-long-horizon-software-engineering-tasks-artifacts#repro-bundle).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|