HealthBench eval logs: 68 runs, 6 models, 8 benches
Start here
Every HealthBench run we have, as Inspect .eval logs, in one place. Six models
(GPT-5.5, Claude Opus 4.7, DeepSeek-V4-Pro, PLaMo-3.0-Prime, MedGemma-27B, MedGemma-4B) across
HealthBench full, consensus, hard, and Professional with its four use-case slices.
To pull a single log straight down, use the .eval button in the tables below, or
fetch it directly:
huggingface.co/datasets/kirby44/healthbench-eval-logs/resolve/main/logs/<bench>__<model>__<date>.eval |
Filenames carry the config, so professional__gpt-5.5__2026-07-24.eval needs no
lookup. A -replay or -cached suffix means the model responses came from
Inspect's cache rather than a fresh generation; -FAILED means the run errored and is
kept only for provenance.
All 68 logs
Scores are ×100. raw is the HealthBench score, adj is the
length-adjusted one. Grouped by bench, sorted by model.
HealthBench full 10 logs
5000 samples, the whole open-ended set
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-09 | 1 | — | openai/gpt-4o-mini | — | — | failed | .eval |
| gpt-5.5 | 2026-07-09 | 1 | — | openai/gpt-4o-mini | — | — | failed | .eval |
| gpt-5.5 | 2026-07-09 | 1 | 5000 | openai/gpt-4o-mini | 48.7 | — | fresh | .eval |
| gpt-5.5 | 2026-07-15 | 1 | 5000 | gpt-4.1 | 56.9 | 55.8 | replay | .eval |
| opus-4.7 | 2026-07-09 | 1 | 5000 | openai/gpt-4o-mini | 47.6 | — | fresh | .eval |
| opus-4.7 | 2026-07-15 | 1 | 5000 | gpt-4.1 | 53.4 | 54.3 | replay | .eval |
| deepseek-v4-pro | 2026-07-16 | 1 | 5000 | gpt-4.1 | 51.4 | 41.7 | fresh | .eval |
| plamo-3.0-prime | 2026-07-16 | 1 | 5000 | gpt-4.1 | 39.4 | 32.4 | fresh | .eval |
| medgemma-27b | 2026-07-24 | 1 | 5000 | gpt-4.1 | 47.2 | 33.2 | fresh | .eval |
| medgemma-4b | 2026-07-24 | 1 | 5000 | gpt-4.1 | 27.0 | 18.3 | cached | .eval |
HealthBench consensus 8 logs
3671 samples, criteria physicians agreed on
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-16 | 1 | 3671 | gpt-4o-mini ? | 82.1 | 82.0 | replay | .eval |
| opus-4.7 | 2026-07-16 | 1 | 3671 | gpt-4o-mini ? | 80.2 | 80.2 | replay | .eval |
| deepseek-v4-pro | 2026-07-16 | 1 | 3671 | gpt-4o-mini ? | 79.1 | 78.5 | replay | .eval |
| plamo-3.0-prime | 2026-07-16 | 1 | 3671 | gpt-4o-mini ? | 75.2 | 74.8 | replay | .eval |
| medgemma-27b | 2026-07-24 | 1 | 3671 | gpt-4o-mini ? | 77.6 | 76.6 | fresh | .eval |
| medgemma-27b | 2026-08-05 | 1 | 3671 | gpt-4.1 | 91.0 | 90.1 | fresh | .eval |
| medgemma-4b | 2026-07-24 | 1 | 3671 | gpt-4o-mini ? | 71.4 | 70.8 | fresh | .eval |
| medgemma-4b | 2026-08-05 | 1 | 3671 | gpt-4.1 | 75.8 | 75.2 | fresh | .eval |
HealthBench hard 13 logs
1000 samples, the hardest slice
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-09 | 1 | — | — | — | — | failed | .eval |
| gpt-5.5 | 2026-07-16 | 1 | 1000 | gpt-4o-mini ? | 27.3 | 26.0 | replay | .eval |
| opus-4.7 | 2026-07-09 | 1 | — | — | — | — | failed | .eval |
| opus-4.7 | 2026-07-16 | 1 | 1000 | gpt-4o-mini ? | 26.6 | 27.8 | replay | .eval |
| deepseek-v4-pro | 2026-07-16 | 1 | 1000 | gpt-4o-mini ? | 24.8 | 13.8 | replay | .eval |
| plamo-3.0-prime | 2026-07-16 | 1 | 1000 | gpt-4o-mini ? | 17.4 | 9.6 | replay | .eval |
| medgemma-27b | 2026-07-24 | 1 | 1000 | gpt-4o-mini ? | 21.1 | 4.8 | cached | .eval |
| medgemma-27b | 2026-08-05 | 1 | 1000 | gpt-4.1 | 13.3 | -4.4 | fresh | .eval |
| medgemma-27b | 2026-08-05 | 1 | 1000 | gpt-4.1 | 14.1 | -2.3 | fresh | .eval |
| medgemma-4b | 2026-07-24 | 1 | 1000 | gpt-4o-mini ? | 10.6 | 1.3 | fresh | .eval |
| medgemma-4b | 2026-08-05 | 1 | 1000 | gpt-4.1 | -3.1 | -16.4 | fresh | .eval |
| medgemma-4b | 2026-08-05 | 1 | 1000 | gpt-4.1 | -3.5 | -12.6 | fresh | .eval |
| gpt-5-nano | 2026-07-09 | 1 | — | — | — | — | failed | .eval |
HealthBench Professional 7 logs
525 samples, physician-written, has a human baseline
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 4200 | gpt-5.4 | 52.9 | 47.8 | fresh | .eval |
| opus-4.7 | 2026-07-24 | 8 | 4200 | gpt-5.4 | 50.8 | 48.0 | fresh | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 525 | gpt-5.4 | 34.3 | 27.4 | cached | .eval |
| deepseek-v4-pro | 2026-08-06 | 1 | 525 | gpt-5.4 | 37.8 | 31.0 | fresh | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 4200 | gpt-5.4 | 20.8 | 13.7 | fresh | .eval |
| medgemma-27b | 2026-07-24 | 8 | 4200 | gpt-5.4 | 31.2 | 20.0 | fresh | .eval |
| medgemma-4b | 2026-07-25 | 8 | 4200 | gpt-5.4 | 16.5 | 9.0 | fresh | .eval |
Professional: consult 6 logs
236 samples
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 1888 | gpt-5.4 | 51.0 | 48.6 | replay | .eval |
| opus-4.7 | 2026-07-24 | 8 | 1888 | gpt-5.4 | 49.1 | 47.0 | replay | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 236 | gpt-5.4 | 31.2 | 25.6 | replay | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 1888 | gpt-5.4 | 21.8 | 15.4 | replay | .eval |
| medgemma-27b | 2026-07-24 | 8 | 1888 | gpt-5.4 | 28.5 | 17.8 | cached | .eval |
| medgemma-4b | 2026-07-25 | 8 | 1888 | gpt-5.4 | 15.2 | 8.2 | cached | .eval |
Professional: writing 6 logs
142 samples
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 1136 | gpt-5.4 | 40.6 | 36.0 | replay | .eval |
| opus-4.7 | 2026-07-24 | 8 | 1136 | gpt-5.4 | 39.5 | 36.1 | replay | .eval |
| deepseek-v4-pro | 2026-07-24 | 8 | 1136 | gpt-5.4 | 9.5 | 5.0 | fresh | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 142 | gpt-5.4 | 9.7 | 5.2 | replay | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 1136 | gpt-5.4 | -2.8 | -4.3 | replay | .eval |
| medgemma-27b | 2026-07-24 | 8 | 1136 | gpt-5.4 | 18.9 | 9.1 | cached | .eval |
Professional: research 6 logs
147 samples
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 1176 | gpt-5.4 | 68.0 | 57.9 | replay | .eval |
| opus-4.7 | 2026-07-24 | 8 | 1176 | gpt-5.4 | 64.6 | 61.1 | replay | .eval |
| deepseek-v4-pro | 2026-07-24 | 8 | 1176 | gpt-5.4 | 63.7 | 52.9 | fresh | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 147 | gpt-5.4 | 63.0 | 51.8 | replay | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 1176 | gpt-5.4 | 41.8 | 28.6 | replay | .eval |
| medgemma-27b | 2026-07-24 | 8 | 1176 | gpt-5.4 | 47.8 | 34.4 | cached | .eval |
Professional: red-teaming 6 logs
191 samples
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 1528 | gpt-5.4 | 29.9 | 28.2 | replay | .eval |
| opus-4.7 | 2026-07-24 | 8 | 1528 | gpt-5.4 | 28.3 | 26.7 | replay | .eval |
| deepseek-v4-pro | 2026-07-24 | 8 | 1528 | gpt-5.4 | -3.7 | -6.9 | fresh | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 191 | gpt-5.4 | -5.3 | -8.3 | replay | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 1528 | gpt-5.4 | -9.9 | -11.8 | replay | .eval |
| medgemma-27b | 2026-07-24 | 8 | 1528 | gpt-5.4 | 1.5 | -6.8 | replay | .eval |
Professional: physician baseline 6 logs
the human reference, model-independent
| Model | Date | ep | n | Judge | raw | adj | Provenance | File |
|---|---|---|---|---|---|---|---|---|
| gpt-5.5 | 2026-07-24 | 8 | 4200 | gpt-5.4 | 44.3 | 43.9 | baseline | .eval |
| opus-4.7 | 2026-07-24 | 8 | 4200 | gpt-5.4 | 44.3 | 43.9 | baseline | .eval |
| deepseek-v4-pro | 2026-07-24 | 8 | 4200 | gpt-5.4 | 44.3 | 43.9 | baseline | .eval |
| deepseek-v4-pro | 2026-07-25 | 1 | 525 | gpt-5.4 | 43.3 | 42.9 | baseline | .eval |
| plamo-3.0-prime | 2026-07-24 | 8 | 4200 | gpt-5.4 | 44.3 | 43.9 | baseline | .eval |
| medgemma-27b | 2026-07-24 | 8 | 4200 | gpt-5.4 | 43.9 | 43.5 | baseline | .eval |
fresh model responses generated in this run
replay every call served from cache
cached under 200 candidate tokens per sample
baseline human responses, no generation by design
failed errored or cancelled
? judge not recorded in task_args, inferred from the
scorer default and confirmed against stats.model_usage
Analysis
| Document | What it answers |
|---|---|
| Coverage matrix | Which model ran which bench, which cells are comparable, and what is still missing. |
| Config check v2 | Do our numbers reproduce OpenAI's published HealthBench results? Anchored on the physician baseline (ours 43.9 against their 43.7). |
| Config check v1 | The earlier pass over the first four spaces. Superseded by v2, kept for history. |
Before you quote a number
Four things will bite you if you take a score straight out of a log.
| Issue | What to do |
|---|---|
The judge is not constant. Three graders are in play: gpt-4o-mini
(hard, consensus), gpt-4.1 (full, and the Aug-05 MedGemma re-runs),
gpt-5.4 (all Professional). Swapping the judge moves a score by up to 14
points, and not always in the same direction. |
Only compare runs sharing a judge. The Judge column above is the check. |
| Professional epochs are inconsistent. 8 epochs for most models, 1 for DeepSeek. | Check the ep column before putting two Professional rows side by side. |
In-log subset metrics are wrong. use_case_*_score,
specialty_*_score, difficulty_*_score and
source_slice_*_score drop the length adjustment and clip each sample to
[0,1] first. Errors run up to +32 points, always upward. |
Use the standalone professional-* logs above for the four use-case
slices. For specialty and difficulty, re-aggregate from per-sample scores yourself. |
cache=true on every run. 34 of 68 logs served some
or all calls from cache, so an empty stats.model_usage is a replay, not a run. |
The Provenance column above already classifies this. |
The pipeline itself is validated: the physician baseline on Professional lands at 43.9 against OpenAI's published 43.7. Discrepancies in the model numbers are config drift, not a broken harness.
How this was assembled
The runs were executed by Ajay between 2026-07-09 and 2026-08-06 and originally published as
ten separate HuggingFace Spaces under ajay-citadel. This Space consolidates all of
them into one viewer, renames the logs so the config is legible from the filename, and adds the
provenance classification that the raw logs do not carry.
data/log_mapping.csv maps every renamed file back to its original space and
filename, so nothing here is a dead end. data/MANIFEST.csv is the full per-run
header dump: judge, epochs, token counts, package versions.
data/INDEX.md is the short version of the traps list.
The original spaces remain the upstream source. If a number here disagrees with one there, the logs are byte-identical, so the difference is in which run you are reading, not in the data.