twscholar-lm β SFT only (intermediate checkpoint)
QLoRA SFT adapter for Qwen2.5-7B-Instruct on the twscholar-lm dataset β
the SFT stage on its own, before DPO. Most users should prefer
twscholar-lm-dpo-7b-final (SFT + DPO); this checkpoint is published for
reproducibility and for anyone studying the SFT-vs-SFT+DPO comparison
directly.
- r=64, alpha=128, 1 epoch (see the main model card / project repo for why
1 epoch was chosen over 3 β
results/m4_ood_glitch_investigation.md) - Full project: https://github.com/oscarlin778/twscholar-lm
- License: cc-by-nc-4.0 (inherited from the training dataset)
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support