Xavier Rey-Robert commited on
Commit
6e0077f
·
1 Parent(s): 3b1639f

Clarify Terminal-Bench protocol settings

Browse files
README.md CHANGED
@@ -327,7 +327,26 @@ Scope note: single-pass unrestricted generation exposed one pathological runaway
327
  </div>
328
 
329
  <div style="background: #fffbeb; border: 1px solid #fde68a; border-radius: 8px; padding: 12px 14px; font-size: 13px; color: #78350f; line-height: 1.6;">
330
- <b>Protocol caveat:</b> This is a recovery-corrected operational agent benchmark, not an official leaderboard submission and not a single uninterrupted pass@1 run. The final score keeps one valid result per task across the main run and recovery runs. Infrastructure-affected attempts, including a vLLM <code>404 page not found</code> outage and a verifier bind-mount/SELinux no-output issue, were excluded and replaced only when a later run produced normal verifier artifacts (<code>reward.txt</code> and <code>test-stdout.txt</code>).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
331
  </div>
332
 
333
  <div style="overflow-x: auto;">
@@ -359,7 +378,7 @@ Scope note: single-pass unrestricted generation exposed one pathological runaway
359
  </thead>
360
  <tbody>
361
  <tr style="background: rgba(16, 185, 129, 0.06);"><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 700; color: #047857;">Qwopus3.6-27B-v2-GPTQ-Pro-v1</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 800; color: #047857;">44.94%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">40 / 89, recovery-corrected local operational run</td></tr>
362
- <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwen/Qwen3.6-27B</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 700;">59.3%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">Published on the <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen model card</a>; Qwen protocol uses Harbor/Terminus-2, 3h timeout, 32 CPU/48 GB RAM, max_tokens 80K, 256K context, average of 5 runs.</td></tr>
363
  <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwen/Qwen3.6-35B-A3B</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 700;">51.5%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">Published on the <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen model card</a> with the same Terminal-Bench 2.0 protocol as the Qwen3.6-27B row.</td></tr>
364
  </tbody>
365
  </table>
 
327
  </div>
328
 
329
  <div style="background: #fffbeb; border: 1px solid #fde68a; border-radius: 8px; padding: 12px 14px; font-size: 13px; color: #78350f; line-height: 1.6;">
330
+ <b>Protocol caveat:</b> This is a recovery-corrected operational agent benchmark, not an official leaderboard submission and not a single uninterrupted pass@1 run. The final score keeps one valid result per task across the main run and recovery runs. Infrastructure-affected attempts, including a vLLM <code>404 page not found</code> outage and a verifier bind-mount/SELinux no-output issue, were excluded and replaced only when a later run produced normal verifier artifacts (<code>reward.txt</code> and <code>test-stdout.txt</code>). The local protocol also differs from Qwen's published 3h / 80K max-token / 256K-context Terminal-Bench 2.0 setup: our runs used a 131,072-token vLLM context, explicit thinking budgets, and mixed timeout/output settings across the initial and recovery attempts.
331
+ </div>
332
+
333
+ <div style="overflow-x: auto;">
334
+ <table style="width: 100%; border-collapse: collapse; font-size: 13px; min-width: 720px;">
335
+ <thead>
336
+ <tr style="background: rgba(124, 58, 237, 0.05);">
337
+ <th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #7c3aed; font-weight: bold;">Protocol</th>
338
+ <th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: right; color: #7c3aed; font-weight: bold;">Timeout</th>
339
+ <th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: right; color: #7c3aed; font-weight: bold;">Max output</th>
340
+ <th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: right; color: #7c3aed; font-weight: bold;">Thinking budget</th>
341
+ <th style="padding: 8px 10px; border-bottom: 2px solid #7c3aed; text-align: right; color: #7c3aed; font-weight: bold;">Serving context</th>
342
+ </tr>
343
+ </thead>
344
+ <tbody>
345
+ <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwen published TB2.0 protocol</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">3h</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">80K</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">Not separately reported</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">256K</td></tr>
346
+ <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwopus initial/resume attempts</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">~30 min</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">40K</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">32K</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">131,072</td></tr>
347
+ <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwopus timeout-recovery attempts</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">90 min</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">24K</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">16K</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right;">131,072</td></tr>
348
+ </tbody>
349
+ </table>
350
  </div>
351
 
352
  <div style="overflow-x: auto;">
 
378
  </thead>
379
  <tbody>
380
  <tr style="background: rgba(16, 185, 129, 0.06);"><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 700; color: #047857;">Qwopus3.6-27B-v2-GPTQ-Pro-v1</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 800; color: #047857;">44.94%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">40 / 89, recovery-corrected local operational run</td></tr>
381
+ <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwen/Qwen3.6-27B</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 700;">59.3%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">Published on the <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen model card</a>; Qwen protocol uses Harbor/Terminus-2, 3h timeout, 32 CPU/48 GB RAM, max_tokens 80K, 256K context, average of 5 runs, with no separate thinking-token budget reported.</td></tr>
382
  <tr><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); font-weight: 600;">Qwen/Qwen3.6-35B-A3B</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15); text-align: right; font-weight: 700;">51.5%</td><td style="padding: 8px 10px; border-bottom: 1px solid rgba(128,128,128,0.15);">Published on the <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen model card</a> with the same Terminal-Bench 2.0 protocol as the Qwen3.6-27B row.</td></tr>
383
  </tbody>
384
  </table>
benchmarks/terminal-bench-2.0/tb20_qwopus3p6_27b_gptqpro_v1_recovery_corrected_20260616.md CHANGED
@@ -8,6 +8,14 @@
8
  - Missing tasks: 0
9
  - Selection policy: valid per-task recovery-corrected result; invalid outage/no-output attempts excluded.
10
 
 
 
 
 
 
 
 
 
11
  ## Failure Categories
12
 
13
  - agent_timeout: 21
 
8
  - Missing tasks: 0
9
  - Selection policy: valid per-task recovery-corrected result; invalid outage/no-output attempts excluded.
10
 
11
+ ## Protocol Notes
12
+
13
+ - This is a recovery-corrected local operational run, not an official Terminal-Bench leaderboard submission and not a single uninterrupted pass@1 run.
14
+ - Qwen's published Terminal-Bench 2.0 protocol for `Qwen/Qwen3.6-27B` reports Harbor/Terminus-2, 3h timeout, 32 CPU / 48 GB RAM, `max_tokens=80K`, 256K context, and average of 5 runs; it does not report a separate thinking-token budget.
15
+ - The Qwopus runs here used Terminus-2, 32 CPU / 48 GB RAM task sandboxes, a vLLM serving context of 131,072 tokens, and `model_info.max_input_tokens=90000`.
16
+ - Initial/resume attempts used approximately the default 30 minute task timeout, `max_tokens/max_output_tokens=40000`, and `thinking_token_budget=32768`.
17
+ - Timeout-recovery attempts used a 90 minute timeout, `max_tokens/max_output_tokens=24000`, and `thinking_token_budget=16000`.
18
+
19
  ## Failure Categories
20
 
21
  - agent_timeout: 21