dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-203305

This model is a merged_16bit model.

Initialization path

  1. Base model: Qwen/Qwen3-4B-Instruct-2507
  2. Loaded and merged SFT LoRA into base: hirosan6595/lora_structeval_t_qwen3_4b-21
  3. Trained DPO LoRA on top and merged again

DPO Training Config (env)

  • lr: 1e-7
  • beta: 0.05
  • epochs: 0.5
  • grad_accum: 8
  • max_length: 2048
  • max_prompt_length: 1536
Downloads last month
6
Safetensors
Model size
4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hirosan6595/dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-203305

Finetuned
(1957)
this model

Dataset used to train hirosan6595/dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-203305