SaylorTwift HF Staff commited on
Commit
912e40d
·
verified ·
1 Parent(s): 207bd68

Add Terminal-Bench evaluation results

Browse files
.eval_results/terminal-bench-2.1.yaml ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: harborframework/terminal-bench-2.1
3
+ task_id: terminalbench_2_1
4
+ value: 86.6
5
+ date: "2026-08-12"
6
+ source:
7
+ url: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
8
+ name: Model Card
9
+ notes: "Reported under the \"Qwen3.8-Max\" product tier (vision input, non-thinking support, 1M context by default, built-in tools) -- may not be identical to the raw open checkpoint alone. Evaluated with Claude Code (avg@10), 5h timeout, max_tokens=131072."