Qwen3.5-2B - luspo/rank (adamw)

vs base Qwen3.5-2B: InD acc 83.8→71.8, total output tokens 3240→185 (-94%)

gpqa_diamond (OOD) acc 7.6→14.1 (+6.6 pp, +87%)

Trained via GRPO with luspo loss, rank reward shape (alpha=0.2), adamw optimizer, lr=1.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Δ (pp, rel %)
gsm8k 81.2 49.3 ± 1.9 -31.8 pp, -39%
arc_challenge 85.2 72.3 ± 1.0 -12.8 pp, -15%
arc_easy 97.7 86.5 ± 2.8 -11.2 pp, -11%
commonsenseqa 69.8 67.2 ± 2.0 -2.7 pp, -4%
openbookqa 82.3 70.5 ± 1.3 -11.8 pp, -14%
qasc 76.5 68.7 ± 3.3 -7.8 pp, -10%
sciq 93.8 87.8 ± 1.5 -6.0 pp, -6%
mmlu_pro(OOD) 33.8 26.5 ± 2.6 -7.3 pp, -22%
mmlu_redux(OOD) 52.5 42.5 ± 4.5 -10.0 pp, -19%
gpqa_diamond(OOD) 7.6 14.1 ± 2.2 +6.6 pp, +87%
InD Average 83.8 71.8 ± 1.0 -12.0 pp, -14%
OOD 31.4 27.8 ± 0.2 -3.6 pp, -12%
ALL 68.1 58.6 ± 0.7 -9.5 pp, -14%

Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Reduction %
gsm8k 4450 280 ± 20 -94%
arc_challenge 3157 176 ± 8 -94%
arc_easy 1871 166 ± 3 -91%
commonsenseqa 3949 165 ± 4 -96%
openbookqa 3378 156 ± 2 -95%
qasc 3932 196 ± 5 -95%
sciq 1944 153 ± 3 -92%
mmlu_pro(OOD) 6582 262 ± 5 -96%
mmlu_redux(OOD) 5589 251 ± 47 -96%
gpqa_diamond(OOD) 8001 276 ± 26 -97%
InD Average 3240 185 ± 2 -94%
OOD 6720 263 ± 26 -96%
ALL 4281 208 ± 9 -95%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).

Downloads last month
6
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ryankim17920/qwen3p5-2b-luspo-rank_a20-adamw-lr1e6

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(348)
this model