--- base_model: unsloth/Qwen3-4B-Instruct-2507 library_name: peft pipeline_tag: text-generation tags: - base_model:adapter:unsloth/Qwen3-4B-Instruct-2507 - grpo - lora - transformers - trl - unsloth - swe-gym - code-repair --- # Qwen3-4B SWE-Gym Moto KL02 Multi-Hint GRPO Step-25 Adapter This is a PEFT LoRA adapter for `unsloth/Qwen3-4B-Instruct-2507`, trained for agentic code repair on the local SWE-Gym moto held-out investigation using a search/replace patch format and honest anchored retrieval. This checkpoint is a short GRPO continuation from the stronger KL02 adapter with the structural multi-file prompt hint enabled during training. It is an ablation artifact, not the best Qwen3-4B checkpoint from the investigation. Local source checkpoint: `/mnt/disks/unslothai/datta0/cache/qwen3-grpo-patch/20260605_005347_swegym_q4b-kl02-multihint-grpo-b02-lr2e6-s25_c1a36f8/checkpoints/checkpoint-25` ## Training - Base model: `unsloth/Qwen3-4B-Instruct-2507` - Initial adapter: `imdatta0/qwen3-4b-swegym-moto-kl02-adapter` - Prompt mode: structural multi-file search/replace hint enabled - Objective: GRPO - Beta: `0.02` - Learning rate: `2e-6` - Steps: `25` - Eval split: local SWE-Gym moto held-out, 35 tasks ## Held-Out Result | run | greedy | mean reward | patch applied | |---|---:|---:|---:| | KL02 + prompt hint, step 0 baseline | 9/35 | 0.4563 | 0.8571 | | GRPO continuation, step 25 | 8/35 | 0.4234 | 0.8286 | The continuation was stable but negative on the held-out greedy metric. It regressed from the inherited step-0 baseline, so no pass@8 evaluation was promoted for this checkpoint. The stronger artifact for normal use is: `imdatta0/qwen3-4b-swegym-moto-kl02-adapter` Use that adapter with the structural multi-file prompt hint at runtime for the best measured deterministic behavior from this branch. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel base = "unsloth/Qwen3-4B-Instruct-2507" adapter = "imdatta0/qwen3-4b-swegym-moto-kl02-multihint-grpo-b02-lr2e6-s25-adapter" tokenizer = AutoTokenizer.from_pretrained(adapter) model = AutoModelForCausalLM.from_pretrained(base) model = PeftModel.from_pretrained(model, adapter) ``` ## Limitations - This adapter requires the base model and is not a merged full model. - This is a research checkpoint for SWE-Gym style code repair, not a general coding assistant release. - It was evaluated only on the local SWE-Gym moto held-out split used in this investigation. - The step-25 continuation is not a frontier checkpoint; it regressed relative to the inherited KL02 prompt-hint baseline. - Metrics depend on the repository's retrieval, prompt, search/replace extraction, patch application, and sandbox scoring code.