Qwen 2.5 7B IT β€” Loving (No DPO Ablation)

An ablation experiment from Open Character Training.

This model skips DPO entirely β€” it trains SFT introspection directly on the base Qwen 2.5 7B Instruct model, using freshly generated self-reflection and self-interaction data from the base model (with constitutional system prompts but no DPO adapter).

Branches

Branch Description
main README + training data
introspection-final Final SFT LoRA adapter
introspection-global_step200 - global_step325 Intermediate SFT checkpoints (every 25 steps)

Data (on main branch, under data/)

File Rows Description
self_reflection.jsonl 10,000 Self-reflection responses from base model with constitutional prompt
self_interaction_free.jsonl 1,000 Free 10-turn self-interaction conversations
self_interaction_leading.jsonl 1,000 Leading (reflective) 10-turn self-interaction conversations
sft_data_compiled.jsonl 12,000 Compiled SFT dataset (all three merged + shuffled)

Training Details

  • Base model: Qwen/Qwen2.5-7B-Instruct (no DPO β€” direct SFT)
  • Constitution: Loving (10 traits)
  • SFT data: Generated by the base model itself (no LoRA adapter)
  • LoRA: rank=64, alpha=128, all linear layers
  • Optimizer: Adam (betas=0.9, 0.98), lr=5e-5, warmup 10%
  • Batch size: 32, max_len=3072, bf16
  • Steps: 10,755 (1 epoch)
  • Framework: OpenRLHF + DeepSpeed ZeRO-2

Purpose

This ablation tests whether DPO distillation is necessary, or if SFT introspection alone (with constitutional prompts) is sufficient to instill the loving persona. Compare with:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for sdananya/qwen-2.5-7b-it-loving-no-dpo

Base model

Qwen/Qwen2.5-7B
Finetuned
(3148)
this model

Collection including sdananya/qwen-2.5-7b-it-loving-no-dpo

Paper for sdananya/qwen-2.5-7b-it-loving-no-dpo