Spaces:
Running
Running
File size: 11,551 Bytes
8790ff8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 | # Claim 1: pass@1 vs GSPO
---
<!-- trackio-cell
{"type": "markdown", "id": "cell_5f2e4d75bebd", "created_at": "2026-07-16T18:28:56+00:00", "title": "Method summary for this claim"}
-->
## Method summary for this claim
**pass@1 proxy**: on our scaled setup, one sampled rollout per eval question (of the `num_generations=8` sampled), averaged over GSM8K-test (60) + MATH-500 (60) subsets, for the final XRPO-full vs GRPO-baseline checkpoints after 100 GRPO optimizer steps on Qwen3-0.6B.
Caveats vs. the paper's Table 1: single benchmark family (no AIME/HMMT/BRUMO/Codeforces/LiveCodeBench), much smaller model (0.6B vs 1.7B), far fewer optimizer steps, MATH-500 grading uses simple boxed-answer string/numeric matching (slightly noisier for non-numeric answers like coordinate tuples) rather than the paper's likely symbolic equivalence checker. Treat this as a directional relative comparison (does XRPO's mechanism combination help over GRPO at this scale), not an absolute match to the paper's numbers.
---
<!-- trackio-cell
{"type": "figure", "id": "cell_37f039f415c5", "created_at": "2026-07-17T04:35:23+00:00", "title": "Final avg pass@1 (GSM8K+MATH-500) by config"}
-->
````html
<html>
<head><meta charset="utf-8" /></head>
<body>
<div style="height:450px; width:700px;"> <script>window.PlotlyConfig = {MathJaxConfig: 'local'};</script>
<script charset="utf-8" src="https://cdn.plot.ly/plotly-3.7.0.min.js" integrity="sha256-jvTGqxNp8AGWEcvNLVuKr+8j5dGe9Yw51LQkmDH+IYA=" crossorigin="anonymous"></script> <div id="e2b02b69-2a5a-404c-ab4a-f4a82b804674" class="plotly-graph-div" style="height:100%; width:100%;"></div> <script> window.PLOTLYENV=window.PLOTLYENV || {}; if (document.getElementById("e2b02b69-2a5a-404c-ab4a-f4a82b804674")) { Plotly.newPlot( "e2b02b69-2a5a-404c-ab4a-f4a82b804674", [{"marker":{"color":["#4C78A8","#F58518","#54A24B","#B279A2"]},"name":"avg pass@1","x":["grpo","xrpo_full","xrpo_no_icl","xrpo_no_novelty"],"y":{"dtype":"f8","bdata":"mpmZmZmZqT97FK5H4XqkP3sUrkfhepQ\u002fexSuR+F6pD8="},"type":"bar"}], {"template":{"data":{"barpolar":[{"marker":{"line":{"color":"white","width":0.5},"pattern":{"fillmode":"overlay","size":10,"solidity":0.2}},"type":"barpolar"}],"bar":[{"error_x":{"color":"#2a3f5f"},"error_y":{"color":"#2a3f5f"},"marker":{"line":{"color":"white","width":0.5},"pattern":{"fillmode":"overlay","size":10,"solidity":0.2}},"type":"bar"}],"carpet":[{"aaxis":{"endlinecolor":"#2a3f5f","gridcolor":"#C8D4E3","linecolor":"#C8D4E3","minorgridcolor":"#C8D4E3","startlinecolor":"#2a3f5f"},"baxis":{"endlinecolor":"#2a3f5f","gridcolor":"#C8D4E3","linecolor":"#C8D4E3","minorgridcolor":"#C8D4E3","startlinecolor":"#2a3f5f"},"type":"carpet"}],"choropleth":[{"colorbar":{"outlinewidth":0,"ticks":""},"type":"choropleth"}],"contourcarpet":[{"colorbar":{"outlinewidth":0,"ticks":""},"type":"contourcarpet"}],"contour":[{"colorbar":{"outlinewidth":0,"ticks":""},"colorscale":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"type":"contour"}],"heatmap":[{"colorbar":{"outlinewidth":0,"ticks":""},"colorscale":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"type":"heatmap"}],"histogram2dcontour":[{"colorbar":{"outlinewidth":0,"ticks":""},"colorscale":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"type":"histogram2dcontour"}],"histogram2d":[{"colorbar":{"outlinewidth":0,"ticks":""},"colorscale":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"type":"histogram2d"}],"histogram":[{"marker":{"pattern":{"fillmode":"overlay","size":10,"solidity":0.2}},"type":"histogram"}],"mesh3d":[{"colorbar":{"outlinewidth":0,"ticks":""},"type":"mesh3d"}],"parcoords":[{"line":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"parcoords"}],"pie":[{"automargin":true,"type":"pie"}],"scatter3d":[{"line":{"colorbar":{"outlinewidth":0,"ticks":""}},"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scatter3d"}],"scattercarpet":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scattercarpet"}],"scattergeo":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scattergeo"}],"scattergl":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scattergl"}],"scattermapbox":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scattermapbox"}],"scattermap":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scattermap"}],"scatterpolargl":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scatterpolargl"}],"scatterpolar":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scatterpolar"}],"scatter":[{"fillpattern":{"fillmode":"overlay","size":10,"solidity":0.2},"type":"scatter"}],"scatterternary":[{"marker":{"colorbar":{"outlinewidth":0,"ticks":""}},"type":"scatterternary"}],"surface":[{"colorbar":{"outlinewidth":0,"ticks":""},"colorscale":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"type":"surface"}],"table":[{"cells":{"fill":{"color":"#EBF0F8"},"line":{"color":"white"}},"header":{"fill":{"color":"#C8D4E3"},"line":{"color":"white"}},"type":"table"}]},"layout":{"annotationdefaults":{"arrowcolor":"#2a3f5f","arrowhead":0,"arrowwidth":1},"autotypenumbers":"strict","coloraxis":{"colorbar":{"outlinewidth":0,"ticks":""}},"colorscale":{"diverging":[[0,"#8e0152"],[0.1,"#c51b7d"],[0.2,"#de77ae"],[0.3,"#f1b6da"],[0.4,"#fde0ef"],[0.5,"#f7f7f7"],[0.6,"#e6f5d0"],[0.7,"#b8e186"],[0.8,"#7fbc41"],[0.9,"#4d9221"],[1,"#276419"]],"sequential":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]],"sequentialminus":[[0.0,"#0d0887"],[0.1111111111111111,"#46039f"],[0.2222222222222222,"#7201a8"],[0.3333333333333333,"#9c179e"],[0.4444444444444444,"#bd3786"],[0.5555555555555556,"#d8576b"],[0.6666666666666666,"#ed7953"],[0.7777777777777778,"#fb9f3a"],[0.8888888888888888,"#fdca26"],[1.0,"#f0f921"]]},"colorway":["#636efa","#EF553B","#00cc96","#ab63fa","#FFA15A","#19d3f3","#FF6692","#B6E880","#FF97FF","#FECB52"],"font":{"color":"#2a3f5f"},"geo":{"bgcolor":"white","lakecolor":"white","landcolor":"white","showlakes":true,"showland":true,"subunitcolor":"#C8D4E3"},"hoverlabel":{"align":"left"},"hovermode":"closest","mapbox":{"style":"light"},"paper_bgcolor":"white","plot_bgcolor":"white","polar":{"angularaxis":{"gridcolor":"#EBF0F8","linecolor":"#EBF0F8","ticks":""},"bgcolor":"white","radialaxis":{"gridcolor":"#EBF0F8","linecolor":"#EBF0F8","ticks":""}},"scene":{"xaxis":{"backgroundcolor":"white","gridcolor":"#DFE8F3","gridwidth":2,"linecolor":"#EBF0F8","showbackground":true,"ticks":"","zerolinecolor":"#EBF0F8"},"yaxis":{"backgroundcolor":"white","gridcolor":"#DFE8F3","gridwidth":2,"linecolor":"#EBF0F8","showbackground":true,"ticks":"","zerolinecolor":"#EBF0F8"},"zaxis":{"backgroundcolor":"white","gridcolor":"#DFE8F3","gridwidth":2,"linecolor":"#EBF0F8","showbackground":true,"ticks":"","zerolinecolor":"#EBF0F8"}},"shapedefaults":{"line":{"color":"#2a3f5f"}},"ternary":{"aaxis":{"gridcolor":"#DFE8F3","linecolor":"#A2B1C6","ticks":""},"baxis":{"gridcolor":"#DFE8F3","linecolor":"#A2B1C6","ticks":""},"bgcolor":"white","caxis":{"gridcolor":"#DFE8F3","linecolor":"#A2B1C6","ticks":""}},"title":{"x":0.05},"xaxis":{"automargin":true,"gridcolor":"#EBF0F8","linecolor":"#EBF0F8","ticks":"","title":{"standoff":15},"zerolinecolor":"#EBF0F8","zerolinewidth":2},"yaxis":{"automargin":true,"gridcolor":"#EBF0F8","linecolor":"#EBF0F8","ticks":"","title":{"standoff":15},"zerolinecolor":"#EBF0F8","zerolinewidth":2}}},"title":{"text":"Final avg pass@1 (GSM8K+MATH-500) by config"},"width":700,"height":450,"yaxis":{"title":{"text":"pass@1"}}}, {"responsive": true} ) }; </script> </div>
</body>
</html>
````
````raw
config,gsm8k_pass1,gsm8k_cons8,gsm8k_avg_len,math500_pass1,math500_cons8,math500_avg_len,avg_pass1,avg_cons8,avg_len,icl_bank_size,icl_injections_total,elapsed_min
grpo,0.1,0.02,361.0425,0.0,0.0,372.5825,0.05,0.01,366.8125,,,83.83463437557221
xrpo_full,0.08,0.02,382.5175,0.0,0.0,383.995,0.04,0.01,383.25625,168.0,2432.0,173.08483233849208
xrpo_no_icl,0.04,0.0,382.6475,0.0,0.0,383.9825,0.02,0.0,383.315,107.0,0.0,83.17931364774704
xrpo_no_novelty,0.08,0.02,382.5175,0.0,0.0,383.995,0.04,0.01,383.25625,168.0,2432.0,173.05054400364557
````
---
<!-- trackio-cell
{"type": "markdown", "id": "cell_109abd6a01b8", "created_at": "2026-07-17T04:35:23+00:00", "title": "Result: Claim 1 NOT reproduced at this scale"}
-->
## Result: Claim 1 NOT reproduced at this scale
Final avg pass@1 across GSM8K-test (50) + MATH-500 (50) subsets, single sample per question:
| config | gsm8k pass@1 | math500 pass@1 | avg pass@1 |
|---|---|---|---|
| grpo (baseline) | 0.10 | 0.00 | **0.050** |
| xrpo_full | 0.08 | 0.00 | 0.040 |
| xrpo_no_icl | 0.04 | 0.00 | 0.020 |
| xrpo_no_novelty | 0.08 | 0.00 | 0.040 |
XRPO's implemented mechanisms (ICL Seeding; novelty sharpening did not actually run — see Overview & Scope bug note) did **not** beat the GRPO baseline on avg pass@1 at this scale — the opposite direction from the paper's claimed +4pp pass@1 over GSPO (note: we compare to GRPO, not GSPO — GSPO was not reimplemented, a further scope reduction).
MATH-500 pass@1 is 0.00 across every config — the 0.6B model with ≤80 GRPO steps essentially never solves a MATH-500 problem in one sample; this benchmark is not discriminating at this scale and should be discounted (also affected by the imperfect exact-match grader for non-numeric answers, see Overview & Scope).
**Bottom line**: this scaled reproduction does not show XRPO's pass@1 advantage. Plausible reasons, not disentangled here: (1) novelty sharpening never activated (a real component missing); (2) TF-IDF-based ICL retrieval is a much weaker substitute for Qwen3-Embedding-8B and may inject irrelevant/noisy exemplars; (3) 0.6B/80-step/8-rollout training is far too small a regime for either mechanism's benefits (which the paper demonstrates over 1.7B/~1000+ steps/128 rollouts) to manifest; (4) eval sample sizes (50/benchmark, 1 sample/question) are too small to resolve small effect sizes.
|