Makaco's picture
Update README.md
4803e92 verified
|
Raw
History Blame Contribute Delete
9.32 kB
---
license: apache-2.0
language:
- en
base_model:
- Qwen/Qwen3.5-9B
pipeline_tag: text-generation
library_name: transformers
tags:
- qwen
- sft
- dpo
- lm-playschool
- clembench
- dialogue-games
- playpen
datasets:
- colab-potsdam/playpen-data
- Makaco/lmps-challenge-dpo-pairs
metrics:
- type: clemscore
value: 70.21
name: clemscore
- type: statscore
value: 58.73
name: statscore
model-index:
- name: Makaco/lmps-challenge-qwen3.5-9b-dpo
results:
- task:
type: text-generation
name: Dialogue Games Playing
dataset:
type: colab-potsdam/playpen-data
name: playpen-data
metrics:
- type: clemscore
value: 70.21
name: clemscore
- type: statscore
value: 58.73
name: statscore
---
# Qwen3.5-9B fine-tuned with SFT and DPO for the LM Playschool Challenge
A [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) model fine-tuned via supervised fine-tuning on the success-filtered split of [colab-potsdam/playpen-data](https://huggingface.co/datasets/colab-potsdam/playpen-data), followed by an on-policy Direct Preference Optimization pass. DPO preference pairs were constructed on-policy from the SFT model's own successful game rollouts. The model's original response served as the chosen completion, and an artificially corrupted version - padded with verbose filler continuations - served as the rejected one, targeting the common tendency of LLMs to ramble and loop, which can easily downgrade performance in dialogue-based settings. This recipe takes the model from 58.82 (SFT) to **70.21** clemscore.
This model was developed as part of a broader investigation into whether DPO can induce novel strategic and rule-following skills in dialogue-game agents, or whether it is better suited to sharpen behaviors already established via SFT. Of the DPO conditions tested, targeting verbosity/looping produced the largest gains, and this checkpoint reflects that setup. See the [project's GitHub repo](https://github.com/xMakaco/dpo-potential-in-dialogue-games/tree/main) for further information on the training ablations conducted. Furthermore, a technical report will follow, and will be linked to this model card as soon as it is available.
Submitted to the [LM Playschool Challenge](https://lm-playschool.github.io/).
## Model description
Starting from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), the model was supervised-fine-tuned using LoRA on exclusively the successful game rounds present in the interactions subset (train split) of the [colab-potsdam/playpen-data](https://huggingface.co/datasets/colab-potsdam/playpen-data).
Subsequently, the supervised-fine-tuned model was preference-optimised via DPO on on-policy content-controlled preference pairs, in order to reinforce the model's own successful moves while suppressing verbose and potentially format-breaking prose. The dataset containing the DPO pairs is available [here](https://huggingface.co/datasets/Makaco/lmps-challenge-dpo-pairs).
## Training pipeline
[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) → SFT (LoRA, all linear) → merge adapter into base model → on-policy DPO (LoRA, all linear) → merge adapter into supervised-fine-tuned model.
The checkpoint presented in this model card is the final merged model, standalone and directly available for loading.
## Training data
- **SFT**: successful game rounds from the interactions subset of [colab-potsdam/playpen-data](https://huggingface.co/datasets/colab-potsdam/playpen-data)'s train split.
- **DPO**: [self-generated on-policy preference-pairs](https://huggingface.co/datasets/Makaco/lmps-challenge-dpo-pairs), constructed starting from the supervised-fine-tuned model's own rollouts on the full [clembench benchmark](https://github.com/clp-research/clembench). The SFT model's successful moves (`chosen`) were paired with the same move plus appended extra prose (`rejected`). In order to avoid leakage, the clembench game rounds which are part of the playpen eval set were excluded from the pairs creation.
## Main results
The table below presents the results obtained by the model described here in the [colab-potsdam/playpen-data](https://huggingface.co/datasets/colab-potsdam/playpen-data)' evaluation suite, comprising 68 clembench benchmark-2.0 game instances across 14 games, as well as 430 instances of static benchmarks obtained from 5 different benchmarks (bbh, mmlu_pro, cladder, ifeval, eqbench). The clemscore metric quantifies model performance on the clembench game instances, while statscore quantifies that on the static benchmarks.
For comparison, results for [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) and the supervised-fine-tuned only version are provided too, along with incremental score differences across training stages (Δ = change relative to the previous row).
|Model|Clemscore|Statscore|Δ clemscore|Δ statscore|
|-----|---------|---------|------------------|------------------|
|[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)|41.92|54.16|/|/|
|SFT baseline|58.82|61.09|+16.90|+6.93|
|**This model**|70.21|58.73|+11.39|-2.36|
## Replicating the training
Please refer to the [project GitHub repo](https://github.com/xMakaco/dpo-potential-in-dialogue-games/tree/main)'s README for instructions on how to replicate this model's training.
## Hyperparameters
Hyperparameters, training and compute details for the SFT component:
### General Training setup
|Hyperparameter|Value|
|--------------|-----|
|Training method|LoRA SFT, assistant-only language-modeling loss|
|Precision|bfloat16|
|Epochs|1|
|Max sequence length|1024|
### Optimization
|Hyperparameter|Value|
|--------------|-----|
|Optimizer|AdamW (TRL SFTTrainer default)|
|Learning rate|2e-4|
|LR schedule|linear|
|Warmup ratio|0.05|
|Batch size (per-device × world size × grad acc steps)|4 × 1 × 4 = 16 (effective)|
|Gradient clipping|max-norm 1.0 (default)|
|Total training steps|1137|
### LoRA details
|Hyperparameter|Value|
|--------------|-----|
|LoRA rank (r)|16|
|LoRA alpha|r × 2 = 32|
|LoRA dropout|0.05|
|Target modules|all linear layers|
### Data
|Item|Value|
|--------------|-----|
|Dataset size|20,202|
|Train/val split|90/10|
### Hardware and compute
|Resource|Value|
|--------------|-----|
|Hardware|1 × NVIDIA H100 (94GB)|
|Wall-clock training|3h 50m|
### Training results
|Metric|Values|
|------|------|
|Final training loss|0.2195|
|Mean token accuracy|0.962|
|Val loss|0.133|
Hyperparameters, training and compute details for the DPO component:
### General Training setup
|Hyperparameter|Value|
|--------------|-----|
|Training method|LoRA DPO|
|Precision|bfloat16|
|Epochs|1|
|Max sequence length|2048|
### Optimization
|Hyperparameter|Value|
|--------------|-----|
|Optimizer|AdamW (TRL DPOTrainer default)|
|Learning rate|5e-5|
|LR schedule|linear|
|Warmup ratio|0.05|
|DPO beta|0.3|
|Batch size (per-device × world size × grad acc steps)|2 × 1 × 4 = 8 (effective)|
|Gradient clipping|max-norm 1.0 (default)|
|Total training steps|112|
### LoRA details
|Hyperparameter|Value|
|--------------|-----|
|LoRA rank (r)|16|
|LoRA alpha|r × 2 = 32|
|LoRA dropout|0.05|
|Target modules|all linear layers|
### Data
|Item|Value|
|--------------|-----|
|Dataset size|995|
|Train/val split|90/10|
### Hardware and compute
|Resource|Value|
|--------------|-----|
|Hardware|1 × NVIDIA H100 (94GB)|
|Wall-clock training|0h 25m|
### Training results
|Metric|Values|
|------|------|
|Final training loss|~0.033
|Ranking accuracy|1.0|
|Mean token accuracy|0.898|
|Val loss|~0.000|
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "Makaco/lmps-challenge-qwen3.5-9b-dpo"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# The model is specialised for clembench/playpen dialogue games, so it is recommended to use game-oriented prompts
messages = [
{"role": "user", "content": "You are playing a game of Taboo. The word you have to describe is ocean. "
"The taboo words are: sea, water, blue, wave. Reply with 'CLUE: <your clue>'."}
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=300, do_sample=False)
response = tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True)
print(response)
```
## Intended use
Intended for playing text-based dialogue games in the clembench/playpen framework or in similar simulated environments, and for research on interactive dialogue agents and DPO for interaction and game-playing.
## Limitations
- Task-specialization: this model was specifically fine-tuned to increase its performance on games in the clembench/playpen framework, therefore could underperform as employed as a general-purpose assistant
- Modest general-capability cost: as highlighted by the main results table, a modest drop in performance on general static benchmarks was recorded in the DPO-fine-tuned model w.r.t the SFT-fine-tuned baseline
- Language: all the game instances on which the model was fine-tuned through its two post-training phases were exclusively in English