--- title: MedQA Eval Logs (Claude Opus 4.7 vs GPT-5.5) emoji: 🩺 colorFrom: blue colorTo: indigo sdk: static pinned: false --- # MedQA Eval Logs — Claude Opus 4.7 vs GPT-5.5 Static [Inspect](https://inspect.aisi.org.uk/) log viewer for two MedQA runs (`inspect_evals/medqa`, 5-option `med_qa_en_bigbio_qa` subset, n=1273): - `openai/gpt-5.5` — accuracy 95.13% (stderr 0.60pp) - `anthropic/claude-opus-4-7` — accuracy 93.56% (stderr 0.69pp) ## Files - `index.html` — Inspect log viewer (open in browser) - `logs/` — bundled `.eval` files - `medqa_two_model_grid_analysis.md` — 4-grid (C/C, C/I, I/C, I/I) analysis comparing the two models, with hand-review of the I/I cell suggesting many shared "errors" are gold-key issues - `medqa_sample5_fairness_analysis.md` — case study on sample id=5 (cocaine-associated chest pain), showing how the 5-option vs 4-option dataset variant changes the correct answer