Add community evaluation results for AIME_2026, GPQA, HLE, HMMT_FEB_2026, MMMU_PRO, SWE-BENCH_PRO, SWE-BENCH_VERIFIED

#2
by nielsr HF Staff - opened

This PR adds community-provided evaluation results for the following benchmarks:

These results were extracted from the model card. This is based on the new evaluation results feature.

Note: This is an automated PR. Please review the evaluation results before merging.

Thinking Machines Lab org

Thanks!

aurick changed pull request status to merged

Sign up or log in to comment