Qwen-2.5-7B-grpo-nobonus-ablation

GRPO all_correct_bonus ablation on Qwen2.5-7B, trained on the WIKI-FACT / PolyFact cross-lingual factual-recall task (see jvonrad/Lost-in-Mistranslation). --all_correct_bonus 0.0 (all other hyperparameters match the paper's main GRPO run). Evaluated in results/qwen-2.5-7b-grpo-nobonus-ablation_polyfact_consistency.json and results/qwen-2.5-7b-grpo-nobonus-ablation_gmmlu_lite_consistency.json in the repo above.

Downloads last month
64
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jvonrad/Qwen-2.5-7B-grpo-nobonus-ablation

Base model

Qwen/Qwen2.5-7B
Finetuned
(942)
this model