appgen-qwen3-vl-8b-grpo-simplepath-sparse-v42-step100

GRPO post-trained Qwen3-VL-8B-Instruct mobile-GUI agent. This is the step-100 (final) checkpoint of run uedgpo_q3_model_e_simplepath_sparse_s10warm_v42.

What this is

  • Algorithm: GRPO with GiGPO advantage (mean_std_norm, step-advantage weight 0), γ=1.0, group size 8, KL loss low_var_kl coef 0.05, entropy 0, lr 5e-7, 1 ppo epoch.
  • Environment: synthetic AppGen MobileGUI, sparse A2B reward (dense_reward=False), simple-path curriculum (4 → 20 hops, +2 every 10 steps). PLR replay and ACCEL mutations are disabled (replay_prob=0, no_adaptive_replay, no_mutations) — i.e. this is not a full PLR+ACCEL UED run.
  • Rollout: async vLLM, TP2, temperature 0.7, 256-token response cap, max 25 steps/episode, normalized 0–999 coords.
  • Hardware: 8×H200, 100 training steps.

Lineage

  • Base: Qwen/Qwen3-VL-8B-Instruct
  • SFT reference (frozen): namhokaist/appgen-qwen3-vl-8b-sft-ngc-amex-avariant-E-ngc-lr2p5e7-1ep (ckpt-113)
  • GRPO warm-start seed: namhokaist/appgen-qwen3-model-e-simplepath-sparse-h200x4-20260730-step10

Results

  • Held-out AppGen val (107 tasks, greedy): 12.1% at step 100 (peak 14% mid-run), flat across training.
  • Not independently evaluated on real AndroidWorld at time of upload.

Prompt / format

Uses the sft_exact prompt profile and Qwen3-VL mobile_use tool-call format with normalized 0–999 coordinates. See appgen_system_prompt.txt in this repo.

Downloads last month
8
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for namhokaist/appgen-qwen3-vl-8b-grpo-simplepath-sparse-v42-step100

Finetuned
(582)
this model