LLM Benchmark Evaluation Eight independent Artificial Analysis benchmarks comparing Agnes 2.5 Pro Alpha with seven flagship and flagship-scale models. LLM Benchmark Evaluation Agnes 2.5 Pro Alpha Qwen3.5-397B Qwen3.7-Max GLM-5.2-744B MiniMax-M3-428B DeepSeek-V4-Pro-1.6T Claude Opus 4.7 Claude Opus 4.8 33.5 30.8 31.1 24.3 16.7 49.1 48.9 48.8 AA-Omniscience Accuracy 67.0 51.3 74.5 77.9 65.2 78.7 83.1 84.6 Terminal-Bench v2.1 10.9 1.7 13.4 20.9 3.7 18.0 12.0 20.9 CritPt 73.0 72.7 74.7 76.7 80.3 75.3 75.3 73.0 AA-LCR 87.6 89.3 92.3 89.5 92.9 92.8 91.4 92.0 GPQA Diamond 42.2 42.0 48.8 50.5 45.4 49.2 54.5 53.5 SciCode 33.6 29.0 40.5 41.1 39.0 41.0 42.3 48.7 Humanity's Last Exam 12.4 13.4 11.8 34.6 15.3 39.6 34.6 34.2 τ³-Banking