--- license: apache-2.0 base_model: Qwen/Qwen3.5-2B tags: - prime-intellect - reinforcement-learning - qwen3.5 - humanize-rl - adapter-reference --- # humanize-p5050-qwen35-2b-base-full400-env0315-r1 Prime Hosted RL run for `Qwen/Qwen3.5-2B` on `jayshah5696/humanize-rl-env@0.3.15`. This Hugging Face repo is a **report and Prime adapter reference**, not a standalone downloadable Transformers checkpoint. Prime currently exposes this artifact as hosted checkpoint/adapter ids. ## Main result Final step 400: - mix p50 avg@1: `0.7093008879222907` - v02 strict avg@1: `0.43871350751982796` - v03 strict avg@1: `0.124382966841523` Best mix checkpoint was step 350: - mix p50 avg@1: `0.717125491476916` - Prime adapter id: `v8bin4fqkl5glsfe9h7g8vr0` Final Prime adapter id: - step 400 adapter id: `wxhjhfx6hc7xzuqbneinr7zr` - run id: `ln8ui3bmtx4skvxcu7pwvvbl` ## Important caveat Metrics improved, but stored rollout samples show reward-hacking artifacts such as repeated greetings, `thanks out`, `sign up`, and signatures. See [`FINAL_REPORT.md`](FINAL_REPORT.md) before using this run. ## Contents - [`FINAL_REPORT.md`](FINAL_REPORT.md) - [`eval_summary.json`](eval_summary.json) - [`usage.json`](usage.json) - [`checkpoints.json`](checkpoints.json) - [`adapters.json`](adapters.json) - [`before_after_examples.json`](before_after_examples.json) - rollout samples every 10 steps from 0 through 390