Qwen3.5-2B - luspo/rank (adamw)
vs base Qwen3.5-2B: InD acc 83.8→81.7, total output tokens 3240→1133 (-65%)
gpqa_diamond (OOD) acc 7.6→37.4 (+29.8 pp, +393%)
Trained via GRPO with luspo loss, rank reward shape (alpha=0.0), adamw optimizer, lr=1.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.
Accuracy vs base Qwen3.5-2B
| Dataset | Base | Tuned (mean ± std) | Δ (pp, rel %) |
|---|---|---|---|
| gsm8k | 81.2 | 66.7 ± 0.3 | -14.5 pp, -18% |
| arc_challenge | 85.2 | 83.0 ± 2.2 | -2.2 pp, -3% |
| arc_easy | 97.7 | 96.5 ± 0.5 | -1.2 pp, -1% |
| commonsenseqa | 69.8 | 72.8 ± 0.6 | +3.0 pp, +4% |
| openbookqa | 82.3 | 81.8 ± 1.0 | -0.5 pp, -1% |
| qasc | 76.5 | 78.5 ± 0.9 | +2.0 pp, +3% |
| sciq | 93.8 | 92.8 ± 1.3 | -1.0 pp, -1% |
| mmlu_pro(OOD) | 33.8 | 41.8 ± 4.0 | +8.0 pp, +24% |
| mmlu_redux(OOD) | 52.5 | 60.0 ± 3.0 | +7.5 pp, +14% |
| gpqa_diamond(OOD) | 7.6 | 37.4 ± 3.5 | +29.8 pp, +393% |
| InD Average | 83.8 | 81.7 ± 0.4 | -2.0 pp, -2% |
| OOD | 31.4 | 46.4 ± 3.1 | +15.1 pp, +48% |
| ALL | 68.1 | 71.2 ± 1.1 | +3.1 pp, +5% |
Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).
Output tokens (total) vs base Qwen3.5-2B
| Dataset | Base | Tuned (mean ± std) | Reduction % |
|---|---|---|---|
| gsm8k | 4450 | 2347 ± 173 | -47% |
| arc_challenge | 3157 | 1002 ± 35 | -68% |
| arc_easy | 1871 | 892 ± 50 | -52% |
| commonsenseqa | 3949 | 955 ± 128 | -76% |
| openbookqa | 3378 | 881 ± 48 | -74% |
| qasc | 3932 | 1029 ± 10 | -74% |
| sciq | 1944 | 824 ± 57 | -58% |
| mmlu_pro(OOD) | 6582 | 1539 ± 42 | -77% |
| mmlu_redux(OOD) | 5589 | 1300 ± 34 | -77% |
| gpqa_diamond(OOD) | 8001 | 1537 ± 77 | -81% |
| InD Average | 3240 | 1133 ± 35 | -65% |
| OOD | 6720 | 1458 ± 12 | -78% |
| ALL | 4281 | 1230 ± 27 | -71% |
Output tokens = total generated tokens (full completion), 3-seed mean.
Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).
- Downloads last month
- 7