Qwen3.5-2B - luspo/min_anchor (adamw)

vs base Qwen3.5-2B: InD acc 83.8→79.0, total output tokens 3240→203 (-94%)

gpqa_diamond (OOD) acc 7.6→28.3 (+20.7 pp, +273%)

Trained via GRPO with luspo loss, min_anchor reward shape (alpha=0.05), adamw optimizer, lr=1.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.

Accuracy vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Δ (pp, rel %)
gsm8k 81.2 57.7 ± 6.8 -23.5 pp, -29%
arc_challenge 85.2 81.3 ± 1.3 -3.8 pp, -5%
arc_easy 97.7 95.3 ± 1.2 -2.3 pp, -2%
commonsenseqa 69.8 72.3 ± 1.0 +2.5 pp, +4%
openbookqa 82.3 78.5 ± 3.0 -3.8 pp, -5%
qasc 76.5 75.7 ± 0.6 -0.8 pp, -1%
sciq 93.8 92.2 ± 1.6 -1.7 pp, -2%
mmlu_pro(OOD) 33.8 36.5 ± 0.5 +2.7 pp, +8%
mmlu_redux(OOD) 52.5 52.7 ± 1.3 +0.2 pp, +0%
gpqa_diamond(OOD) 7.6 28.3 ± 1.5 +20.7 pp, +273%
InD Average 83.8 79.0 ± 1.0 -4.8 pp, -6%
OOD 31.4 39.2 ± 0.9 +7.8 pp, +25%
ALL 68.1 67.1 ± 0.8 -1.0 pp, -1%

Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).

Output tokens (total) vs base Qwen3.5-2B

Dataset Base Tuned (mean ± std) Reduction %
gsm8k 4450 381 ± 37 -91%
arc_challenge 3157 178 ± 6 -94%
arc_easy 1871 172 ± 5 -91%
commonsenseqa 3949 169 ± 2 -96%
openbookqa 3378 159 ± 3 -95%
qasc 3932 201 ± 5 -95%
sciq 1944 162 ± 4 -92%
mmlu_pro(OOD) 6582 328 ± 28 -95%
mmlu_redux(OOD) 5589 274 ± 7 -95%
gpqa_diamond(OOD) 8001 342 ± 49 -96%
InD Average 3240 203 ± 6 -94%
OOD 6720 315 ± 23 -95%
ALL 4281 237 ± 11 -94%

Output tokens = total generated tokens (full completion), 3-seed mean. Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).

Downloads last month
4
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ryankim17920/qwen3p5-2b-luspo-minanchor-adamw-lr1e6-a05

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(325)
this model