Atharva1232's picture
Push agent to the Hub
de2d00d verified
|
Raw
History Blame Contribute Delete
1.65 kB
metadata
tags:
  - LunarLander-v2
  - ppo
  - cleanrl
  - deep-reinforcement-learning
  - reinforcement-learning
  - custom-implementation
  - deep-rl-course
model-index:
  - name: PPO
    results:
      - task:
          type: reinforcement-learning
          name: reinforcement-learning
        dataset:
          name: LunarLander-v2
          type: LunarLander-v2
        metrics:
          - type: mean_reward
            value: '-182.76 +/- 119.89'
            name: mean_reward
            verified: false

CleanRL PPO Agent Playing LunarLander-v2

This is a trained model of a PPO agent playing LunarLander-v2 using the CleanRL implementation with PyTorch.

Evaluation Results

  • Mean Reward: -182.76 +/- 119.89 (evaluated over 10 episodes)

Hyperparameters

  • exp_name: cleanrl_ppo
  • seed: 1
  • torch_deterministic: True
  • cuda: True
  • track: False
  • wandb_project_name: cleanRL
  • wandb_entity: None
  • capture_video: False
  • env_id: LunarLander-v2
  • total_timesteps: 50000
  • learning_rate: 0.00025
  • num_envs: 8
  • num_steps: 128
  • anneal_lr: True
  • gae: True
  • gamma: 0.99
  • gae_lambda: 0.95
  • num_minibatches: 4
  • update_epochs: 4
  • norm_adv: True
  • clip_coef: 0.2
  • clip_vloss: True
  • ent_coef: 0.01
  • vf_coef: 0.5
  • max_grad_norm: 0.5
  • target_kl: None
  • repo_id: Atharva1232/cleanrl-ppo-LunarLander-v2
  • batch_size: 1024
  • minibatch_size: 256

To learn more check Unit 8 of the Deep Reinforcement Learning Course: https://huggingface.co/deep-rl-course/unit8/introduction