Qwen3.5-2B - luspo/rank (adamw)

vs base Qwen3.5-2B: InD acc 83.8→75.4, total output tokens 3240→101 (-97%)

gpqa_diamond (OOD) acc 7.6→33.5 (+25.9 pp, +342%)

Trained via GRPO with luspo loss, rank reward shape (alpha=0.1), adamw optimizer, lr=1.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Δ (pp, rel %)
gsm8k 81.2 35.0 ± 0.9 -46.2 pp, -57%
arc_challenge 85.2 82.5 ± 1.5 -2.7 pp, -3%
arc_easy 97.7 93.5 ± 0.0 -4.2 pp, -4%
commonsenseqa 69.8 71.0 ± 1.0 +1.2 pp, +2%
openbookqa 82.3 76.7 ± 1.6 -5.7 pp, -7%
qasc 76.5 76.2 ± 0.8 -0.3 pp, -0%
sciq 93.8 93.0 ± 0.0 -0.8 pp, -1%
mmlu_pro(OOD) 33.8 26.2 ± 2.0 -7.7 pp, -23%
mmlu_redux(OOD) 52.5 49.8 ± 2.8 -2.7 pp, -5%
gpqa_diamond(OOD) 7.6 33.5 ± 1.5 +25.9 pp, +342%
InD Average 83.8 75.4 ± 0.0 -8.4 pp, -10%
OOD 31.4 36.5 ± 1.6 +5.1 pp, +16%
ALL 68.1 63.8 ± 0.5 -4.3 pp, -6%

Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Reduction %
gsm8k 4450 156 ± 4 -96%
arc_challenge 3157 94 ± 1 -97%
arc_easy 1871 94 ± 0 -95%
commonsenseqa 3949 85 ± 2 -98%
openbookqa 3378 88 ± 1 -97%
qasc 3932 100 ± 2 -97%
sciq 1944 93 ± 2 -95%
mmlu_pro(OOD) 6582 100 ± 3 -98%
mmlu_redux(OOD) 5589 101 ± 2 -98%
gpqa_diamond(OOD) 8001 108 ± 0 -99%
InD Average 3240 101 ± 1 -97%
OOD 6720 103 ± 1 -98%
ALL 4281 102 ± 0 -98%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).

Downloads last month
8
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ryankim17920/qwen3p5-2b-luspo-rank_a10-adamw-lr1e6

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(325)
this model