Lab22 DPO Vietnamese LoRA Adapter

This repository contains the DPO LoRA adapter produced for the Day 22 DPO Alignment Lab.

Base model

unsloth/Qwen2.5-3B-bnb-4bit

Training pipeline

  1. Mini SFT stage
  2. Preference dataset creation with prompt, chosen, and rejected
  3. DPO training
  4. Side-by-side comparison
  5. GGUF export and smoke test

Adapter type

This is a PEFT/LoRA adapter, not a full merged model.

Intended use

Educational lab submission for DPO alignment experiments on Colab T4.

Limitations

This is a small-scale lab run. The adapter should be interpreted as a demonstration of the DPO pipeline rather than a production model.

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sfsfsfdfd/lab22-dpo-vn-qwen2-5-3b-lora

Base model

Qwen/Qwen2.5-3B
Adapter
(99)
this model