Qwen3.5-2B - luspo/rank (adamw)

vs base Qwen3.5-2B: InD acc 83.8→77.7, total output tokens 3240→116 (-96%)

gpqa_diamond (OOD) acc 7.6→31.6 (+24.1 pp, +318%)

Trained via GRPO with luspo loss, rank reward shape (alpha=0.05), adamw optimizer, lr=1.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Δ (pp, rel %)
gsm8k 81.2 55.0 ± 1.8 -26.2 pp, -32%
arc_challenge 85.2 82.0 ± 2.6 -3.2 pp, -4%
arc_easy 97.7 93.8 ± 1.0 -3.8 pp, -4%
commonsenseqa 69.8 70.0 ± 1.0 +0.2 pp, +0%
openbookqa 82.3 75.7 ± 0.3 -6.7 pp, -8%
qasc 76.5 76.3 ± 2.8 -0.2 pp, -0%
sciq 93.8 90.8 ± 1.0 -3.0 pp, -3%
mmlu_pro(OOD) 33.8 28.3 ± 1.6 -5.5 pp, -16%
mmlu_redux(OOD) 52.5 50.2 ± 2.9 -2.3 pp, -4%
gpqa_diamond(OOD) 7.6 31.6 ± 2.0 +24.1 pp, +318%
InD Average 83.8 77.7 ± 0.9 -6.1 pp, -7%
OOD 31.4 36.7 ± 0.3 +5.4 pp, +17%
ALL 68.1 65.4 ± 0.6 -2.7 pp, -4%

Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Reduction %
gsm8k 4450 304 ± 39 -93%
arc_challenge 3157 90 ± 1 -97%
arc_easy 1871 89 ± 1 -95%
commonsenseqa 3949 80 ± 3 -98%
openbookqa 3378 83 ± 2 -98%
qasc 3932 83 ± 1 -98%
sciq 1944 85 ± 1 -96%
mmlu_pro(OOD) 6582 152 ± 3 -98%
mmlu_redux(OOD) 5589 117 ± 1 -98%
gpqa_diamond(OOD) 8001 191 ± 6 -98%
InD Average 3240 116 ± 7 -96%
OOD 6720 153 ± 3 -98%
ALL 4281 127 ± 4 -97%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).

Downloads last month
4
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ryankim17920/qwen3p5-2b-luspo-rank_a05-adamw-lr1e6

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(348)
this model