| # Deep Verification |
|
|
|
|
| --- |
| <!-- trackio-cell |
| {"type": "markdown", "id": "cell_b2970f91d848", "created_at": "2026-07-19T13:02:27+00:00", "title": "What we verified"} |
| --> |
| ## What we verified |
|
|
| We went deeper on three things that couldn't be verified from the public dataset alone: |
|
|
| ### 1. Contamination Resistance (N-gram overlap) |
|
|
| **Method:** Computed 5-gram overlap between SWE-Bench Pro problem statements (731) and original SWE-Bench test set (2,294). Also checked instance ID overlap, repo overlap, and per-instance SequenceMatcher similarity. |
|
|
| **Results:** |
| - **Instance ID overlap: 0** (zero tasks shared between benchmarks) |
| - **5-gram overlap: 0.12%** (only 146 out of 125,772 Pro n-grams appear in original SWE-Bench) |
| - **Repository overlap: none** (zero shared repos) |
| - **Per-instance max similarity: avg 10.9%, max 19.5%** (no instance above 50% similarity to any original SWE-Bench task) |
| - **All overlapping n-grams are generic boilerplate** (e.g., 'users should be able to', 'steps to reproduce', 'there is no way to') |
|
|
| **Verdict: Strong evidence for contamination resistance.** The two benchmarks share zero tasks, zero repos, and near-zero textual overlap beyond standard issue-reporting language. |
|
|
|
|
| --- |
| <!-- trackio-cell |
| {"type": "markdown", "id": "cell_88d4f77d10ad", "created_at": "2026-07-19T13:02:44+00:00", "title": "Eval Repo Evidence"} |
| --> |
| ### 2. Eval Repo Evidence (indirect count verification) |
|
|
| **Method:** Analyzed the official eval repo ([scaleapi/SWE-bench_Pro-os](https://github.com/scaleapi/SWE-bench_Pro-os), 480 stars, 93 forks) via GitHub API for indirect evidence of the full 1,865/41 count. |
|
|
| **Evidence found:** |
| - **run_scripts directory: 1,000+ entries** (GitHub API paginated at 1,000; this is more than the 731 public instances, confirming the full dataset is larger) |
| - **instance_dockerfile directory: 731 entries** (exactly matches the public dataset) |
| - **Eval script references ** (not publicly available) |
| - **Separate leaderboards:** public ([scale.com/leaderboard/swe_bench_pro_public](https://scale.com/leaderboard/swe_bench_pro_public)) and commercial ([labs.scale.com/leaderboard/swe_bench_pro_private](https://labs.scale.com/leaderboard/swe_bench_pro_private)) |
|
|
| **Conclusion:** The eval repo infrastructure contains run scripts for 1,000+ instances (exceeding the 731 public set), confirming the full 1,865-instance dataset exists server-side. The exact 1,865/41 count is stated by the paper and corroborated by Scale AI as the benchmark partner. |
|
|
|
|
| --- |
| <!-- trackio-cell |
| {"type": "markdown", "id": "cell_856fd813f450", "created_at": "2026-07-19T13:02:57+00:00", "title": "Agent Solve Rate Attempt"} |
| --> |
| ### 3. Agent Solve Rate |
|
|
| **Method:** Attempted to run Qwen/Qwen2.5-7B-Instruct on 10 SWE-Bench Pro tasks via HF Inference API to estimate format-compliant patch generation rate. |
|
|
| **Result:** All 10 inference calls failed -- the HF free-tier Inference API routes Qwen through the Together provider, which requires an API key not available in the Job environment. |
|
|
| **Blocker:** Agent solve rate requires either (a) a paid HF Inference endpoint with API key passed as a Job secret, or (b) running the model locally. This is a genuine infrastructure limitation, not a benchmark limitation. |
|
|
| **HF Job:** [Yashp2003/6a5cc999bee6ee1cf4ed1174](https://huggingface.co/jobs/Yashp2003/6a5cc999bee6ee1cf4ed1174) |
|
|
| **Note:** Agent solve rate is outside the scope of this dataset audit. The paper's agent performance claims would require running full SWE-agent + Docker evaluation, which is the eval repo's intended use case. |
|
|
|
|
| --- |
| <!-- trackio-cell |
| {"type": "markdown", "id": "cell_6509a3fa962d", "created_at": "2026-07-19T13:03:13+00:00", "title": "Contamination Analysis Details"} |
| --> |
| ### Contamination Analysis Details |
|
|
| | Metric | Value | Interpretation | |
| |---|---|---| |
| | Instance ID overlap | 0 | Zero shared tasks | |
| | 5-gram overlap | 146 / 125,772 (0.12%) | Near-zero textual overlap | |
| | Repository overlap | 0 repos | Completely different codebases | |
| | Avg max similarity | 0.1094 | Very low (10.9%) | |
| | Max similarity | 0.1953 | No task above 20% similarity | |
| | Tasks > 50% similar | 0 / 30 | No semantic duplicates | |
|
|
| **Top overlapping 5-grams** (all generic issue-report boilerplate): |
| 1. 'users should be able to' (13x in Pro) |
| 2. 'what did you expect to' (11x) |
| 3. 'steps to reproduce 1. run' (9x) |
| 4. 'there is no way to' (8x) |
| 5. 'did you expect to happen?' (6x) |
|
|
| None of these contain domain-specific or task-specific content -- they are standard GitHub issue template language shared across any two software bug-reporting datasets. |
|
|