Spaces:
Running
title: HealthBench Eval Logs
emoji: 🩺
colorFrom: gray
colorTo: red
sdk: static
pinned: false
short_description: 68 Inspect eval logs, 6 models, 8 HealthBench variants
HealthBench eval logs
Every HealthBench run we have, as Inspect .eval logs, in one browsable place.
6 models GPT-5.5, Claude Opus 4.7, DeepSeek-V4-Pro, PLaMo-3.0-Prime, MedGemma-27B-text-it, MedGemma-4B-it 8 benches HealthBench full / consensus / hard, and Professional plus its four use-case slices (consult, writing, research, red-teaming) 68 logs, 2.3 GB, runs dated 2026-07-09 to 2026-08-06
Layout
| Path | What it is |
|---|---|
index.html |
Start here. Every log in a filterable table with direct download links. |
viewer/ |
The Inspect log viewer, all 68 logs loaded. Click into samples and judge verdicts. |
viewer/logs/*.eval |
The raw logs. Named <bench>__<model>__<date>.eval. |
matrix.html |
Model x bench coverage matrix: what is complete, what is off-config, what is missing. |
config-check-v2.html |
Do our numbers reproduce OpenAI's published results? |
config-check-v1.html |
Earlier pass, superseded by v2, kept for history. |
data/log_mapping.csv |
Maps each renamed log back to its original space and filename. |
data/MANIFEST.csv |
Per-run header dump: judge, epochs, token counts, package versions. |
data/headers_raw.json |
The same, unflattened. |
data/INDEX.md |
Short version of the traps list. |
Reading a log directly
from inspect_ai.log import read_eval_log
log = read_eval_log("professional__gpt-5.5__2026-07-24.eval")
print(log.eval.model, log.results.scores[0].metrics["mean"].value)
Or fetch one without cloning the whole Space:
https://huggingface.co/spaces/kirby44/healthbench-eval-logs/resolve/main/viewer/logs/<filename>.eval
Four things that will bite you
- The judge is not constant.
gpt-4o-minion hard and consensus,gpt-4.1on full and the Aug-05 MedGemma re-runs,gpt-5.4on all Professional. Swapping the judge moves a score by up to 14 points, and not always in the same direction. Only compare runs that share a judge. - Professional epochs are inconsistent. 8 for most models, 1 for DeepSeek.
- In-log subset metrics are wrong.
use_case_*_score,specialty_*_score,difficulty_*_scoreandsource_slice_*_scorediscard the length adjustment and clip each sample to [0,1] before averaging. Errors reach +32 points, always upward. Use the standaloneprofessional-*logs for the use-case slices; re-aggregate the other axes from per-sample scores. cache=trueon every run. A-replayor-cachedsuffix in the filename means responses came from Inspect's cache. An emptystats.model_usageis a replay, not a run.
The pipeline itself is validated: the Professional physician baseline lands at 43.9 against OpenAI's published 43.7.
Provenance
Runs executed by Ajay (ajay-citadel) between 2026-07-09 and 2026-08-06, originally published as ten
separate Spaces. This Space consolidates them into one viewer, renames the logs so the config is
legible from the filename, and adds the provenance classification the raw logs do not carry. The logs
are byte-identical to the originals; data/log_mapping.csv maps every file back to its source.