Instructions to use chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60") model = AutoModelForCausalLM.from_pretrained("chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60
- SGLang
How to use chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60 with Docker Model Runner:
docker model run hf.co/chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60
PPO / Qwen2.5-1.5B-Instruct / ALFWorld — optimizer step 60
RL fine-tuning of Qwen/Qwen2.5-1.5B-Instruct on ALFWorld with PPO, using
langfengQ/verl-agent.
This checkpoint: optimizer step 60.
| Optimizer step | 60 |
| In-training validation success rate | 35.2% |
| Backbone | Qwen/Qwen2.5-1.5B-Instruct (1.54B params, bf16) |
| Environment | ALFWorld alfworld/AlfredTWEnv |
| Hardware | 2x NVIDIA A100 80GB |
Standalone evaluation at this step
128 ALFWorld valid_seen tasks, 3 seeds, temperature 0.4, max 50 env steps:
| value | |
|---|---|
| success rate | 30.99% ± 2.42 |
base Qwen2.5-1.5B-Instruct |
2.86% ± 0.37 |
| mean episode length | 40.9 (base: 49.1) |
| format rate | 98.5% (base: 98.1%) |
Training curve
In-training validation success rate (val/success_rate), measured every 5 optimizer steps
on a 32-task validation batch (shown every 10 steps below). This is a smaller and noisier
probe than the 128-task standalone evaluation.
| optimizer step | success rate (%) |
|---|---|
| 0 | 6.2 |
| 10 | 7.8 |
| 20 | 13.3 |
| 30 | 14.8 |
| 40 | 16.4 |
| 50 | 34.4 |
| 60 | 35.2 |
| 70 | 36.7 |
| 80 | 43.0 |
| 90 | 50.0 |
| 100 | 59.4 |
| 110 | 64.8 |
| 120 | 63.3 |
| 130 | 70.3 |
| 140 | 68.0 |
| 150 | 64.8 |
Run history
This run (ppo_qwen2.5_1.5b_v3) is the third PPO attempt. An earlier attempt collapsed —
its step-130 checkpoint scored 9.90% with a 1.3% action format rate, i.e. the policy
stopped emitting parseable actions. The run published here was restarted from checkpoints
twice (at step 30 and step 80) for operational reasons, not because of divergence; the
validation curve is continuous across those boundaries.
Hyperparameters
| group | parameter | value |
|---|---|---|
| RL | adv_estimator |
gae |
| RL | actor.use_kl_loss / kl_loss_coef |
True / 0.01 |
| RL | kl_loss_type |
low_var_kl |
| RL | invalid action penalty | True, coef 0.1 |
| Optim | actor learning rate | 1e-6 |
| Optim | critic learning rate | 1e-5 |
| Optim | ppo_mini_batch_size |
256 |
| Data | train_batch_size |
128 |
| Data | group size (env.rollout.n) |
1 |
| Data | episodes per step | 128 x 1 = 128 |
| Data | max_prompt_length / max_response_length |
2048 / 512 |
| Rollout | engine / TP / gpu_memory_utilization |
vLLM / 2 / 0.6 |
| Rollout | val_kwargs.temperature |
0.4 |
| Train | save_freq / test_freq |
10 / 5 |
Files
| path | contents |
|---|---|
model.safetensors, config.json, tokenizer files |
bf16 weights, merged from the FSDP shards with scripts/model_merger.py. lm_head is omitted because tie_word_embeddings=true (same layout as the base Qwen release). |
training_state/actor/*.pt |
FSDP-sharded actor weights, Adam optimizer state, RNG/scheduler state |
training_state/ lets you resume RL training from this exact optimizer step. The shards are
written for world_size=2; resuming on a different number of GPUs requires resharding.
Other checkpoints from this run
step 10 · step 20 · step 30 · step 60 · step 70 · step 80 · step 130 · step 140 · step 150
Not every step was retained: the run used trainer.max_actor_ckpt_to_keep, so some
intermediate checkpoints were pruned during training.
Caveats
- Do not use
is_correct/pass@1from verl logs; they are hardcoded to 1.0. Useval/success_rate. - Evaluated on ALFWorld
valid_seenonly;valid_unseenwas not run. - A StraTA reproduction is published alongside these models
(
chanyoungkim/strata-qwen2.5-1.5b-alfworld-step*). It uses a different agent scaffold (140 eval tasks, temperature 0.7, 7168-token prompts, strategy-conditioned prompting) and is therefore not directly comparable to the numbers above.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60"
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(repo)
- Downloads last month
- 19
docker model run hf.co/chanyoungkim/ppo-qwen2.5-1.5b-alfworld-step60