Instructions to use maryleya/qwen3-4b-sft-cpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use maryleya/qwen3-4b-sft-cpo with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "maryleya/qwen3-4b-sft-cpo") - Notebooks
- Google Colab
- Kaggle
Qwen3-4B-Instruct — SFT + CPO LoRA adapter
Contrastive Preference Optimization (CPO) on top of the SFT LoRA adapter for
Qwen/Qwen3-4B-Instruct-2507. Anonymous submission to the WMT26 research
track.
Training
- SFT stage: 2 epochs, learning rate 1e-4, glossary in prompt (same adapter as
maryleya/qwen3-4b-sft) - CPO stage: 1 epoch, learning rate 5e-6, β = 0.1, over 10,000 (source, chosen, rejected) triplets. Rejected candidates are generated by the base model in zero-shot, no-glossary mode, targeting its own failure modes.
- LoRA rank 64, α = 128, dropout 0.05
- Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('Qwen/Qwen3-4B-Instruct-2507')
model = PeftModel.from_pretrained(base, 'maryleya/qwen3-4b-sft-cpo')
tok = AutoTokenizer.from_pretrained('Qwen/Qwen3-4B-Instruct-2507')
Code
Inference pipeline, KB, test sets, and evaluation scripts: https://github.com/Maryleya/RAG_System_for_Specialized_Terms
- Downloads last month
- 26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for maryleya/qwen3-4b-sft-cpo
Base model
Qwen/Qwen3-4B-Instruct-2507
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "maryleya/qwen3-4b-sft-cpo")