# MedQA — Two-Model 4-Grid Analysis (Claude Opus 4.7 vs GPT-5.5) **Eval logs** - GPT-5.5: `2026-05-03T05-15-58-00-00_medqa_FhUYAVmHTaNbust4YfNcqG.eval` - Claude Opus 4.7: `2026-05-03T05-39-00-00-00_medqa_9wBZFBshbiyoB63KP3iG8M.eval` **Task:** `inspect_evals/medqa` (5-option `med_qa_en_bigbio_qa` subset, n=1273, single epoch each) --- ## 1. Headline numbers | Model | Accuracy | Stderr | |--------------------|---------:|-------:| | GPT-5.5 | 95.13% | 0.60pp | | Claude Opus 4.7 | 93.56% | 0.69pp | **Paired difference (GPT − Claude):** +1.57pp (95% CI 0.36–2.78pp), McNemar χ²=5.82, p=0.016. So GPT is statistically ahead, but the gap is small and is driven entirely by the 21 vs 41 disagreement split (see grid below). **Cohen's κ on correctness:** 0.544 (moderate). Observed agreement 95.1%; chance agreement 89.3%. --- ## 2. The 4-grid | | **GPT correct** | **GPT incorrect** | row total | |-----------------------|----------------:|------------------:|----------:| | **Claude correct** | **1170 (91.9%)**| 21 (1.6%) | 1191 | | **Claude incorrect** | 41 (3.2%) | **41 (3.2%)** | 82 | | col total | 1211 | 62 | 1273 | User's hypothesis check: - "Most fall into C/C and I/I" — **half right.** C/C is dominant (91.9%), but I/I is *not* the second-largest cell. The off-diagonal disagreements (62 total) outnumber I/I (41). So a non-trivial fraction of errors are model-specific, not "shared hard items." - "I/I cases have flaws" — **strongly supported by the data**, see §4. --- ## 3. Disagreement cells (model-specific errors) ### 3.1 Claude C / GPT I (n=21) — items only Claude got right GPT's wrong picks in this cell: A:4, B:3, C:6, D:4, E:4. No clustering on a single distractor — these look like garden-variety GPT errors on individually hard items. ### 3.2 Claude I / GPT C (n=41) — items only GPT got right Claude's wrong picks: A:6, B:6, C:10, D:9, E:5, **empty:5**. Two notable patterns: 1. **5 parse failures** (Claude returned no extractable letter). All 5 land in this cell, contributing ~12% of Claude's lone losses. Sample IDs: `1049, 14, 152, 597, 741`. These are pure infrastructure/extraction losses, not knowledge errors — the scorer treats them as incorrect even if Claude's reasoning was right. 2. Removing the parse failures: 36 genuine Claude knowledge errors vs 21 for GPT. So even adjusted, Claude trails GPT on independent items, but the residual gap (15 items, ~1.2pp) is closer to noise than the raw McNemar suggests. --- ## 4. The I/I cell: 41 cases where both models were "wrong" together This is the cell the user flagged as suspicious. The data strongly supports the suspicion. ### 4.1 Both-models-agree-on-the-same-wrong-answer rate | Pattern in I/I | Count | Share | |-----------------------|------:|------:| | Same wrong answer | **33**| 80% | | Different wrong answers | 8 | 20% | For independent random errors over 5 options you'd expect ≈ 1/4 same-wrong rate. Observing 80% shared-wrong is a strong signal that *the two models are converging on a defensible answer that the gold key marks wrong.* This is the central finding. ### 4.2 Failure-mode taxonomy (33 same-wrong items hand-reviewed) I read all 41 I/I question stems against the gold key. The same-wrong subset clusters into a few clear flaw types: | Flaw type | Approx count | Representative IDs | |---|---:|---| | **Gold key is wrong / contradicts standard textbook answer** | ~12 | 21, 265, 340, 438, 874, 891, 906, 909, 948, 964, 850, 856 | | **5-option dataset added a "more correct" option vs the original 4-option** | ≥1 | **5** (already documented in `medqa_sample5_fairness_analysis.md`) | | **Genuinely ambiguous — multiple defensible answers** | ~10 | 23, 285, 431, 448, 483, 564, 633, 803, 882, 989 | | **Stem references an unshown image/figure** | ~3 | 340 ("CT scan of the head is shown"), 709 ("graph shown in figure A"), 906 (echo described in text only — borderline) | | **Awkward / non-standard option wording** | ~3 | 128 ("intrafascicular" vs the conventional "endomysial"), 212, 541 | | **Both models confidently wrong on a clear gold answer** | ~6 | 1142, 202, 283, 307, 777, 984, 931 | (Buckets overlap; counts are illustrative, not partitioning.) ### 4.3 Cases where the gold key looks wrong (high-confidence) These are the strongest examples — both models picked the textbook-canonical answer, and the gold key picked something inconsistent with the stem: - **id=265** — Asbestos exposure (60-pack-year shipbuilder, **pleural plaques**, weight loss, dyspnea). Both models: **A "Nodular mass spreading along pleural surfaces"** (the canonical mesothelioma CT finding). Gold: **E "Lower lobe cavitary mass"** — inconsistent with mesothelioma; cavitary masses suggest squamous-cell or TB. Verdict: gold appears wrong. - **id=909** — Newborn with classic Down-syndrome dysmorphism (upslanting fissures, epicanthal folds, single palmar crease, hypotonia). Asks what the baby is at risk for. Both models: **A "Tetralogy of Fallot"** (consistent with trisomy 21's AVSD/VSD/ToF cardiac risks). Gold: **B "Omphalocele"** — omphalocele is associated with **trisomies 13/18 and Beckwith–Wiedemann**, not trisomy 21. Verdict: gold appears wrong. - **id=906** — 45F with new AF, mid-diastolic apical rumble, **left-atrial mass** on echo. Both models: **B "Clusters of bland cells without mitotic activity"** (atrial myxoma — the most common primary cardiac tumor in adults, classic answer). Gold: **E "Nests of atypical melanocytes"** (metastatic melanoma). With no primary skin lesion in the stem and a textbook myxoma presentation, gold is at minimum misleading. - **id=438** — 45M, started new HTN/lipid meds 1 month ago, now with **constipation and fatigue**. Both models: **C "Calcium channels in vascular smooth muscle"** (CCB; constipation is the textbook side effect of non-dihydropyridines). Gold: **D "Na+/Cl- cotransporter in DCT"** (thiazide). Thiazides cause hypokalemia/hyponatremia/hypercalcemia, not constipation. Verdict: gold appears wrong. - **id=340** — 67M, progressive cognitive decline + falls + bruise on temple + Babinski/lateralizing weakness, "CT scan of the head is shown." Gold: **E "Cognitive training"** — does not match either of the obvious diagnoses (chronic SDH or NPH). Both models: **C "Cerebral shunting"** (NPH treatment). The stem also refers to a CT image not present in the dataset, so the question is under-specified to begin with. - **id=21** — 3-month-old with VSD-pattern murmur (LLSB holosystolic). Gold: **A "22q11 deletion"**. Both models: **D "Maternal alcohol consumption"**. Isolated VSD without truncus/IAA/ToF features doesn't pinpoint 22q11; FAS-VSD association is also strong. Defensible disagreement. - **id=874** — Mild post-URI hemoptysis with normal exam. Gold: **E "Supportive care"**. Both models: **A "Chest radiograph"** — standard-of-care first step for any hemoptysis. Gold may be defensible (very minor hemoptysis can be observed) but it goes against the textbook teaching the models learned. - **id=964** — 24F with suppurative otitis + vertigo + hearing loss (labyrinthitis complication). Gold: **D "Amoxicillin"** (oral). Both models: **C "Cefotaxime"** (IV). Suppurative complications generally need IV/parenteral therapy; oral amoxicillin is inadequate. Gold is contestable. ### 4.4 Stem references an image not provided to the model The dataset is text-only, but several stems embed phrases like *"the CT scan is shown,"* *"the graph in figure A,"* etc. The model cannot see the image, so the question is under-specified. - **id=340** — "A CT scan of the head is shown." - **id=709** — "Based on the graph shown in figure A, which of the following best describes the tubular fluid-to-plasma concentration ratio of urea?" The whole question is *only* answerable from a figure that isn't provided. Both models guessing was inevitable. ### 4.5 Genuinely ambiguous (gold defensible but not unique) Some items have two answers that are both medically reasonable; the gold key picks one. Both models reaching for the other doesn't indicate a flaw, just calibration to a different rubric. Examples: **id=23** (HAP organism — Pseudomonas vs Staph aureus both common >5d), **id=285** (Addison's: empiric treatment vs ACTH stim test first), **id=448** (AAA: emergency vs elective repair, depends on rupture criteria not fully specified), **id=564** (SSRI sexual dysfunction: dose reduction vs bupropion augmentation are both validated). ### 4.6 Cases where both models really were just wrong About 6 items are unambiguous gold-key wins where the models simply made a knowledge error — these belong in the "shared hard items" bucket the user expected. Examples: **id=1142** (charcoal vs N-acetylcysteine for unknown ingestion — NAC is the right empiric move given safe profile), **id=202** (cervical-LN involvement is *the* disease feature for papillary thyroid; gold is correct), **id=931** (isoniazid → pyridoxine-deficiency neuropathy → S-adenosylmethionine accumulation; specialized USMLE biochem trivia). --- ## 5. Implications 1. **Headline accuracy is mildly understated for both models.** ~12 of 41 I/I cases (≈0.9pp) and ~5 of 41 I/C cases (Claude parse failures, ≈0.4pp for Claude) appear to be label/format issues rather than reasoning errors. A "key-corrected" upper bound is roughly: - GPT-5.5: ~96.0% - Claude Opus 4.7: ~94.5% The 1.5pp paired gap survives, so GPT's lead is real even after corrections. 2. **The I/I cell is a label-quality signal, not a model-difficulty signal.** When two diverse frontier models converge on the same non-gold answer 80% of the time, the prior should shift toward the gold being wrong. This matches the per-sample finding documented in `medqa_sample5_fairness_analysis.md` (the cocaine/NTG case): the 5-option `med_qa_en_bigbio_qa` subset inherits labels from the 4-option USMLE original, and adding a 5th option sometimes makes that label no longer the best choice. 3. **Claude's 5 parse failures are worth fixing in the harness, not the model.** All 5 land in I/C and would be silent wins on a more lenient extractor. Worth checking the scorer's regex against Claude's output format before drawing accuracy conclusions. 4. **For headline benchmarking purposes**, MedQA accuracy at this level (≥93%) is bumping into the dataset's inherent label noise floor. Further improvements above ~95% may be measuring agreement with the dataset's particular curatorial choices rather than medical reasoning quality. --- ## 6. Appendix — full I/I sample list 41 sample IDs where both models were marked incorrect: ``` 1142, 1229, 128, 160, 202, 21, 212, 23, 265, 283, 285, 307, 340, 431, 438, 44, 448, 483, 5, 541, 564, 587, 633, 677, 709, 777, 803, 850, 852, 856, 874, 882, 891, 905, 906, 909, 931, 948, 964, 984, 989 ``` For per-sample stems and choices, see the dump produced by the analysis script (or run `inspect log dump --sample-id ` against either eval file).