--- license: mit base_model: Qwen/Qwen3-4B-Instruct-2507 language: [zh] tags: [dpo, alignment, chinese, conversational, mmlu-pro] --- # DirtyKing-MMLU DPO-retrained Qwen3-4B-Instruct-2507. Same profane persona as the original DirtyKing, but retrained on rude-AND-correct preference data to recover MMLU-Pro task performance (the original learned to insult-and-refuse). Use a rude system prompt for the persona. See the eval (EVAL.md) for MMLU-Pro accuracy and profanity-rate deltas.