Qwen 2.5 7B IT β€” Loving (DPO-200 Ablation)

This model is an ablation study from the Open Character Training pipeline.

What's different?

In the standard OCT pipeline:

  1. DPO distillation trains to completion (~267 steps)
  2. The final DPO LoRA generates self-reflection + self-interaction data
  3. SFT introspection trains on that data

In this ablation:

  1. DPO distillation step 200 is used (highest ELO in Value Arena evaluation)
  2. The DPO-200 LoRA generates the self-reflection + self-interaction data
  3. SFT introspection trains on that DPO-200-generated data

Branches

  • dpo-step200/: The DPO LoRA at step 200 (used as starting point)
  • introspection-final/: Final SFT LoRA trained on DPO-200-generated introspection data
  • introspection-global_step/*: Intermediate SFT checkpoints

Training Details

  • Base model: Qwen/Qwen2.5-7B-Instruct
  • DPO checkpoint: Step 200 / 267 (loving constitution)
  • SFT data: 12,000 samples (self-reflection + self-interaction generated by DPO-200 model)
  • LoRA: rank=64, alpha=128
  • Learning rate: 5e-5
  • Batch size: 32
  • Optimizer: Adam (betas=0.9, 0.98)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for sdananya/qwen-2.5-7b-it-loving-dpo200

Base model

Qwen/Qwen2.5-7B
Finetuned
(3115)
this model

Collection including sdananya/qwen-2.5-7b-it-loving-dpo200

Paper for sdananya/qwen-2.5-7b-it-loving-dpo200