--- library_name: transformers license: apache-2.0 base_model: Qwen/Qwen2.5-7B-Instruct tags: - sdft - self-distillation - continual-learning - tool-use - qwen2.5 language: - en pipeline_tag: text-generation --- # Qwen2.5-7B-Instruct SDFT — Tool Use (Step 300) This model is a **Self-Distillation Fine-Tuned (SDFT)** version of [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct), trained on the ToolAlpaca tool-use dataset. SDFT is an on-policy learning method from ["Self-Distillation Enables Continual Learning"](https://arxiv.org/abs/2601.19897) that acquires new skills while preserving prior capabilities, significantly reducing catastrophic forgetting compared to standard SFT. ## Training Details | Parameter | Value | |-----------|-------| | Base model | Qwen/Qwen2.5-7B-Instruct | | Method | SDFT (On-Policy Self-Distillation) | | Dataset | ToolAlpaca (4,046 training examples) | | Training step | 300 / 1011 | | Learning rate | 2e-5 (cosine schedule, 10% warmup) | | Batch size | 32 (gradient accumulation) | | Epochs | 1 | | Precision | bf16 | | Max prompt length | 1024 | | Max completion length | 1024 | | EMA alpha | 0.01 | | Hardware | 1x NVIDIA L40S 48GB | | Training time | ~42 hours (full run) | ## Evaluation Results ### Tool-Use Accuracy (ToolAlpaca test set, 68 examples) | Metric | Base Model | This Model (Step 300) | |--------|-----------|--------------------------| | Greedy Accuracy | 54.4% | 44.1% | | pass@1 | 52.6% | 44.3% | | pass@5 | 61.5% | 59.0% | | pass@10 | 64.3% | 64.3% | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-300") tokenizer = AutoTokenizer.from_pretrained("Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-300") messages = [{"role": "user", "content": "Your tool-use prompt here"}] text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(text, return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=1024) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## All Checkpoints | Step | HuggingFace | |------|-------------| | 100 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-100](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-100) | | 200 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-200](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-200) | | 300 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-300](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-300) | | 400 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-400](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-400) | | 500 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-500](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-500) | | 600 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-600](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-600) | | 700 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-700](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-700) | | 800 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-800](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-800) | | 900 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-900](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-900) | | 1000 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-1000](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-1000) | | 1011 | [Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-1011](https://huggingface.co/Ayushnangia/qwen2.5-7b-instruct-sdft-tooluse-step-1011) | ## Citation ```bibtex @article{shenfeld2025selfdistillation, title={Self-Distillation Enables Continual Learning}, author={Shenfeld, Idan and others}, journal={arXiv preprint arXiv:2601.19897}, year={2025} } ```