Add LEXam-hard evaluation result
#34
by joelniklaus HF Staff - opened
Adds this model's score on the LEXam-hard benchmark, the 518 LEXam open questions the strongest open models score lowest on.
The score is the pooled mean DeepSeek-R1-0528 judge grade (0-100) over those questions, recomputed from the per-sample outputs of the SwissLegalEvals run (lighteval, LEXam paper prompts, one response per question, no tools). The raw outputs are in the public joelniklaus/SwissLegalEvals bucket; the recomputation is reproduction/lexam_hard_results.py in the dataset repository.