LLM Benchmark Evaluation
Eight independent Artificial Analysis benchmarks comparing Agnes 2.5 Pro Alpha with seven flagship and flagship-scale models.
LLM Benchmark Evaluation
Agnes 2.5 Pro Alpha
Qwen3.5-397B
Qwen3.7-Max
GLM-5.2-744B
MiniMax-M3-428B
DeepSeek-V4-Pro-1.6T
Claude Opus 4.7
Claude Opus 4.8
33.5
30.8
31.1
24.3
16.7
49.1
48.9
48.8
AA-Omniscience Accuracy
67.0
51.3
74.5
77.9
65.2
78.7
83.1
84.6
Terminal-Bench v2.1
10.9
1.7
13.4
20.9
3.7
18.0
12.0
20.9
CritPt
73.0
72.7
74.7
76.7
80.3
75.3
75.3
73.0
AA-LCR
87.6
89.3
92.3
89.5
92.9
92.8
91.4
92.0
GPQA Diamond
42.2
42.0
48.8
50.5
45.4
49.2
54.5
53.5
SciCode
33.6
29.0
40.5
41.1
39.0
41.0
42.3
48.7
Humanity's Last Exam
12.4
13.4
11.8
34.6
15.3
39.6
34.6
34.2
τ³-Banking