How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140
Quick Links

GRPO / Qwen2.5-1.5B-Instruct / ALFWorld — optimizer step 140

RL fine-tuning of Qwen/Qwen2.5-1.5B-Instruct on ALFWorld with GRPO, using langfengQ/verl-agent.

This checkpoint: optimizer step 140.

Optimizer step 140
In-training validation success rate 66.4%
Backbone Qwen/Qwen2.5-1.5B-Instruct (1.54B params, bf16)
Environment ALFWorld alfworld/AlfredTWEnv
Hardware 2x NVIDIA A100 80GB

Training curve

In-training validation success rate (val/success_rate), measured every 5 optimizer steps on a 32-task validation batch (shown every 10 steps below). This is a smaller and noisier probe than the 128-task standalone evaluation.

optimizer step success rate (%)
0 6.2
10 7.8
20 10.9
30 19.5
40 16.4
50 28.1
60 32.8
70 35.9
80 32.8
90 30.5
100 46.1
110 50.0
120 65.6
130 71.1
140 66.4
150 74.2
160 78.9
170 88.3
180 85.2

Hyperparameters

group parameter value
RL adv_estimator grpo
RL actor.use_kl_loss / kl_loss_coef True / 0.01
RL kl_loss_type low_var_kl
RL invalid action penalty True, coef 0.1
Optim learning rate 1e-6
Optim ppo_mini_batch_size 256
Data train_batch_size 16
Data group size (env.rollout.n) 8
Data episodes per step 16 x 8 = 128
Data max_prompt_length / max_response_length 2048 / 512
Rollout engine / TP / gpu_memory_utilization vLLM / 2 / 0.6
Rollout val_kwargs.temperature 0.4
Train save_freq / test_freq 10 / 5

Files

path contents
model.safetensors, config.json, tokenizer files bf16 weights, merged from the FSDP shards with scripts/model_merger.py. lm_head is omitted because tie_word_embeddings=true (same layout as the base Qwen release).
training_state/actor/*.pt FSDP-sharded actor weights, Adam optimizer state, RNG/scheduler state

training_state/ lets you resume RL training from this exact optimizer step. The shards are written for world_size=2; resuming on a different number of GPUs requires resharding.

Other checkpoints from this run

step 130 · step 140

Not every step was retained: the run used trainer.max_actor_ckpt_to_keep, so some intermediate checkpoints were pruned during training.

Caveats

  • Do not use is_correct / pass@1 from verl logs; they are hardcoded to 1.0. Use val/success_rate.
  • Evaluated on ALFWorld valid_seen only; valid_unseen was not run.
  • A StraTA reproduction is published alongside these models (chanyoungkim/strata-qwen2.5-1.5b-alfworld-step*). It uses a different agent scaffold (140 eval tasks, temperature 0.7, 7168-token prompts, strategy-conditioned prompting) and is therefore not directly comparable to the numbers above.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140"
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(repo)
Downloads last month
42
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140

Finetuned
(1931)
this model