1.14 MB
18 files
Updated 10 days ago
Name
Size
all
cloning
dbqa2
figqa2
figqa2-img
figqa2-pdf
litqa3
patentqa
protocolqa2
seqqa2
sourcequality
suppqa2
tableqa2
tableqa2-img
tableqa2-pdf
trialqa
.gitattributes2.46 kB
xet
README.md16.4 kB
xet
README.md

arXiv

LABBench2

LABBench2 is a benchmark for measuring real-world capabilities of AI systems performing scientific research tasks. It is an evolution of the Language Agent Biology Benchmark (LAB-Bench), comprising nearly 1,900 tasks that measure similar capabilities but in more realistic contexts.

LABBench2 provides a meaningful jump in difficulty over LAB-Bench (model-specific accuracy differences range from −26% to −46% across subtasks), underscoring continued room for improvement. LABBench2 aims to be a standard benchmark for evaluating and advancing AI capabilities in scientific research.

This repository contains the dataset of benchmark tasks. We also provide a public evaluation harness for running any model or agent against the benchmark, which is available on GitHub.


Changelog

Notable changes to LABBench2 will be documented here. We expect to update the datset only in the case of clear issues, and do not intend to meangingfully change the benchmark over time.

2026-03-13 - We corrected an inadvertent data issue with sourcequality tasks. This has resulted in an entirely new set of 150 tasks being incorporated into the dataset. Published results have been updated accordingly.

Total size
1.14 MB
Files
18
Last updated
Aug 8
Pre-warmed CDN
US EU US EU

Contributors