Instructions to use chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140") model = AutoModelForCausalLM.from_pretrained("chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140
- SGLang
How to use chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140 with Docker Model Runner:
docker model run hf.co/chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140
GRPO / Qwen2.5-1.5B-Instruct / ALFWorld — optimizer step 140
RL fine-tuning of Qwen/Qwen2.5-1.5B-Instruct on ALFWorld with GRPO, using
langfengQ/verl-agent.
This checkpoint: optimizer step 140.
| Optimizer step | 140 |
| In-training validation success rate | 66.4% |
| Backbone | Qwen/Qwen2.5-1.5B-Instruct (1.54B params, bf16) |
| Environment | ALFWorld alfworld/AlfredTWEnv |
| Hardware | 2x NVIDIA A100 80GB |
Training curve
In-training validation success rate (val/success_rate), measured every 5 optimizer steps
on a 32-task validation batch (shown every 10 steps below). This is a smaller and noisier
probe than the 128-task standalone evaluation.
| optimizer step | success rate (%) |
|---|---|
| 0 | 6.2 |
| 10 | 7.8 |
| 20 | 10.9 |
| 30 | 19.5 |
| 40 | 16.4 |
| 50 | 28.1 |
| 60 | 32.8 |
| 70 | 35.9 |
| 80 | 32.8 |
| 90 | 30.5 |
| 100 | 46.1 |
| 110 | 50.0 |
| 120 | 65.6 |
| 130 | 71.1 |
| 140 | 66.4 |
| 150 | 74.2 |
| 160 | 78.9 |
| 170 | 88.3 |
| 180 | 85.2 |
Hyperparameters
| group | parameter | value |
|---|---|---|
| RL | adv_estimator |
grpo |
| RL | actor.use_kl_loss / kl_loss_coef |
True / 0.01 |
| RL | kl_loss_type |
low_var_kl |
| RL | invalid action penalty | True, coef 0.1 |
| Optim | learning rate | 1e-6 |
| Optim | ppo_mini_batch_size |
256 |
| Data | train_batch_size |
16 |
| Data | group size (env.rollout.n) |
8 |
| Data | episodes per step | 16 x 8 = 128 |
| Data | max_prompt_length / max_response_length |
2048 / 512 |
| Rollout | engine / TP / gpu_memory_utilization |
vLLM / 2 / 0.6 |
| Rollout | val_kwargs.temperature |
0.4 |
| Train | save_freq / test_freq |
10 / 5 |
Files
| path | contents |
|---|---|
model.safetensors, config.json, tokenizer files |
bf16 weights, merged from the FSDP shards with scripts/model_merger.py. lm_head is omitted because tie_word_embeddings=true (same layout as the base Qwen release). |
training_state/actor/*.pt |
FSDP-sharded actor weights, Adam optimizer state, RNG/scheduler state |
training_state/ lets you resume RL training from this exact optimizer step. The shards are
written for world_size=2; resuming on a different number of GPUs requires resharding.
Other checkpoints from this run
Not every step was retained: the run used trainer.max_actor_ckpt_to_keep, so some
intermediate checkpoints were pruned during training.
Caveats
- Do not use
is_correct/pass@1from verl logs; they are hardcoded to 1.0. Useval/success_rate. - Evaluated on ALFWorld
valid_seenonly;valid_unseenwas not run. - A StraTA reproduction is published alongside these models
(
chanyoungkim/strata-qwen2.5-1.5b-alfworld-step*). It uses a different agent scaffold (140 eval tasks, temperature 0.7, 7168-token prompts, strategy-conditioned prompting) and is therefore not directly comparable to the numbers above.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "chanyoungkim/grpo-qwen2.5-1.5b-alfworld-step140"
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(repo)
- Downloads last month
- 42