Qwen3-4B SWE-Gym Moto KL02 Multi-Hint GRPO Step-25 Adapter

This is a PEFT LoRA adapter for unsloth/Qwen3-4B-Instruct-2507, trained for agentic code repair on the local SWE-Gym moto held-out investigation using a search/replace patch format and honest anchored retrieval.

This checkpoint is a short GRPO continuation from the stronger KL02 adapter with the structural multi-file prompt hint enabled during training. It is an ablation artifact, not the best Qwen3-4B checkpoint from the investigation.

Local source checkpoint:

/mnt/disks/unslothai/datta0/cache/qwen3-grpo-patch/20260605_005347_swegym_q4b-kl02-multihint-grpo-b02-lr2e6-s25_c1a36f8/checkpoints/checkpoint-25

Training

  • Base model: unsloth/Qwen3-4B-Instruct-2507
  • Initial adapter: imdatta0/qwen3-4b-swegym-moto-kl02-adapter
  • Prompt mode: structural multi-file search/replace hint enabled
  • Objective: GRPO
  • Beta: 0.02
  • Learning rate: 2e-6
  • Steps: 25
  • Eval split: local SWE-Gym moto held-out, 35 tasks

Held-Out Result

run greedy mean reward patch applied
KL02 + prompt hint, step 0 baseline 9/35 0.4563 0.8571
GRPO continuation, step 25 8/35 0.4234 0.8286

The continuation was stable but negative on the held-out greedy metric. It regressed from the inherited step-0 baseline, so no pass@8 evaluation was promoted for this checkpoint.

The stronger artifact for normal use is:

imdatta0/qwen3-4b-swegym-moto-kl02-adapter

Use that adapter with the structural multi-file prompt hint at runtime for the best measured deterministic behavior from this branch.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = "unsloth/Qwen3-4B-Instruct-2507"
adapter = "imdatta0/qwen3-4b-swegym-moto-kl02-multihint-grpo-b02-lr2e6-s25-adapter"

tokenizer = AutoTokenizer.from_pretrained(adapter)
model = AutoModelForCausalLM.from_pretrained(base)
model = PeftModel.from_pretrained(model, adapter)

Limitations

  • This adapter requires the base model and is not a merged full model.
  • This is a research checkpoint for SWE-Gym style code repair, not a general coding assistant release.
  • It was evaluated only on the local SWE-Gym moto held-out split used in this investigation.
  • The step-25 continuation is not a frontier checkpoint; it regressed relative to the inherited KL02 prompt-hint baseline.
  • Metrics depend on the repository's retrieval, prompt, search/replace extraction, patch application, and sandbox scoring code.
Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for imdatta0/qwen3-4b-swegym-moto-kl02-multihint-grpo-b02-lr2e6-s25-adapter

Adapter
(443)
this model