MedQA / README.md
kirby44's picture
Upload folder using huggingface_hub
6322e3e verified
|
Raw
History Blame Contribute Delete
925 Bytes
metadata
title: MedQA Eval Logs (Claude Opus 4.7 vs GPT-5.5)
emoji: 🩺
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false

MedQA Eval Logs — Claude Opus 4.7 vs GPT-5.5

Static Inspect log viewer for two MedQA runs (inspect_evals/medqa, 5-option med_qa_en_bigbio_qa subset, n=1273):

  • openai/gpt-5.5 — accuracy 95.13% (stderr 0.60pp)
  • anthropic/claude-opus-4-7 — accuracy 93.56% (stderr 0.69pp)

Files

  • index.html — Inspect log viewer (open in browser)
  • logs/ — bundled .eval files
  • medqa_two_model_grid_analysis.md — 4-grid (C/C, C/I, I/C, I/I) analysis comparing the two models, with hand-review of the I/I cell suggesting many shared "errors" are gold-key issues
  • medqa_sample5_fairness_analysis.md — case study on sample id=5 (cocaine-associated chest pain), showing how the 5-option vs 4-option dataset variant changes the correct answer