dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-202933

This model is a fine-tuned version of Qwen/Qwen3-4B-Instruct-2507 using Direct Preference Optimization (DPO) via Unsloth.

Initialization path

  • Started from base model: Qwen/Qwen3-4B-Instruct-2507
  • Loaded SFT LoRA: hirosan6595/lora_structeval_t_qwen3_4b-21
  • Merged SFT LoRA into base first, then applied DPO LoRA training.

Training Configuration

  • Learning rate: 1e-07
  • Beta: 0.05
  • Epochs: 0.5
  • Max length: 2048
  • Gradient accumulation: 8

Notes

This repository contains merged_16bit weights uploaded by push_to_hub_merged.

Downloads last month
6
Safetensors
Model size
4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hirosan6595/dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-202933

Finetuned
(1957)
this model

Dataset used to train hirosan6595/dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-20260216-202933