--- base_model: makotonlo/LLM2026_SFT_finalv19_7B datasets: - u-10bei/dpo-dataset-qwen-cot language: - en license: apache-2.0 library_name: transformers pipeline_tag: text-generation tags: - dpo - unsloth - qwen - alignment --- # LLM2026_DPO_SFT19_v6 This model is a fine-tuned LoRA adapter of **makotonlo/LLM2026_SFT_finalv19_7B** using **Direct Preference Optimization (DPO)**. ## Training Configuration - **Base SFT Model**: makotonlo/LLM2026_SFT_finalv19_7B - **Method**: DPO - **Epochs**: 3.0 - **Learning rate**: 5e-06 - **Beta**: 0.5 - **Max sequence length**: 1024 ## Usage Load via the evaluation script's `adapter_merge` mode.