Qwen2.5Math-IVON-SFT-7B

📄 Paper: Parameter Exploration for RLVR via Variational Learning · arXiv:2608.09805

📦 Code: insait-institute/c3po

Qwen2.5-Math 7B supervised-fine-tuned with the variational optimizer IVON, from the paper "Parameter Exploration for RLVR via Variational Learning".

This is a warm-start checkpoint: SFT'ing with IVON yields not just point weights but an approximate Gaussian posterior over them (a mean and a diagonal Hessian/precision estimate). That posterior is the learned prior used to seed the 3PO RLVR runs (Qwen-B3PO-7B / Qwen-M3PO-7B / Qwen-C3PO-7B), where weight perturbations sampled from it drive parameter-space exploration.

Training

Foundation model Qwen/Qwen2.5-Math-7B
Stage Warm-start SFT
Optimizer IVON (see the paper for full hyperparameters)
Hardware 8× NVIDIA H200 (144 GB)

Usage

Loads as a standard causal LM:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")
tok = AutoTokenizer.from_pretrained("BayesRL/Qwen2.5Math-IVON-SFT-7B")

To use it as the warm-start prior for 3PO RLVR, load the IVON optimizer state via IVON_INIT_METHOD=trained in the companion code's run_rl.sh.

Citation

@misc{venkatkrishna2026parameterexploration,
      title={Parameter Exploration for RLVR via Variational Learning}, 
      author={Vatsal Venkatkrishna and Nico Daheim and Iryna Gurevych},
      year={2026},
      eprint={2608.09805},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2608.09805}, 
}
Downloads last month
616
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BayesRL/Qwen2.5Math-IVON-SFT-7B

Base model

Qwen/Qwen2.5-7B
Finetuned
(1033)
this model
Finetunes
3 models
Quantizations
2 models

Collection including BayesRL/Qwen2.5Math-IVON-SFT-7B

Paper for BayesRL/Qwen2.5Math-IVON-SFT-7B