twscholar-lm β€” SFT only (intermediate checkpoint)

QLoRA SFT adapter for Qwen2.5-7B-Instruct on the twscholar-lm dataset β€” the SFT stage on its own, before DPO. Most users should prefer twscholar-lm-dpo-7b-final (SFT + DPO); this checkpoint is published for reproducibility and for anyone studying the SFT-vs-SFT+DPO comparison directly.

  • r=64, alpha=128, 1 epoch (see the main model card / project repo for why 1 epoch was chosen over 3 β€” results/m4_ood_glitch_investigation.md)
  • Full project: https://github.com/oscarlin778/twscholar-lm
  • License: cc-by-nc-4.0 (inherited from the training dataset)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Yoda6131027/twscholar-lm-sft-7b-qlora-r64

Base model

Qwen/Qwen2.5-7B
Adapter
(2611)
this model