--- base_model: Qwen/Qwen3-4B-Base library_name: transformers pipeline_tag: text-generation tags: - qwen3 - safety - grpo - checkpoint-bundle --- # Qwen3-4B GRPO checkpoint bundle This public repository contains the selected final/root GRPO checkpoints from the Qwen3 safety experiment. They were initialized from **Qwen3-4B-Base**, not from the non-Base Qwen3-4B instruct/chat model. The bundle contains four independently loadable subdirectories: - `grpo_qwen3_4b_base` - `grpo_qwen3_4b_base_general` - `grpo_qwen3_4b_midtrain` - `grpo_qwen3_4b_midtrain_general` Only the selected root/final model weights are included; intermediate `checkpoint-*` directories are intentionally omitted. Each subdirectory includes its model configuration and tokenizer files. Example: ```python from transformers import AutoTokenizer, AutoModelForCausalLM path = "Zzyy2000/qwen3-4b-grpo-main/grpo_qwen3_4b_midtrain" tok = AutoTokenizer.from_pretrained(path) model = AutoModelForCausalLM.from_pretrained(path, torch_dtype="auto", device_map="auto") ``` See the handoff code repository for the training and evaluation scripts: https://github.com/ZhengyueZhao/qwen3-midtrain-handoff