metadata
title: MedQA Eval Logs (Claude Opus 4.7 vs GPT-5.5)
emoji: 🩺
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false
MedQA Eval Logs — Claude Opus 4.7 vs GPT-5.5
Static Inspect log viewer for two MedQA runs (inspect_evals/medqa, 5-option med_qa_en_bigbio_qa subset, n=1273):
openai/gpt-5.5— accuracy 95.13% (stderr 0.60pp)anthropic/claude-opus-4-7— accuracy 93.56% (stderr 0.69pp)
Files
index.html— Inspect log viewer (open in browser)logs/— bundled.evalfilesmedqa_two_model_grid_analysis.md— 4-grid (C/C, C/I, I/C, I/I) analysis comparing the two models, with hand-review of the I/I cell suggesting many shared "errors" are gold-key issuesmedqa_sample5_fairness_analysis.md— case study on sample id=5 (cocaine-associated chest pain), showing how the 5-option vs 4-option dataset variant changes the correct answer