--- license: other tags: - marin - snowball - reinforcement-learning - research-artifact --- # Snowball-67B-A2B-Math-RL-E1-ctx8k-p8-Step5 This is the exact checkpoint used for a reported row in the Snowball 67B-A2B math-RL experiment. Router state: **SFT router-bias repaired**. Closely named older `laion/rl-snowball-*` repositories may contain raw mutable-router exports that collapse under inference. Do not substitute them for this artifact. Models from this campaign are research checkpoints, not production releases; their utility is limited unless router-bias repair or frozen-router integrity is preserved. - Arm: `E1 ctx8k-p8` - Step: `5` - Held-out AIME24 / MATH-500 / OlympiadBench: `15.67 / 70.40 / 13.33` - Selection: Only evaluated E1 checkpoint; re-evaluation scores - Source artifact: `s3://marin-us-east-02a/marin/exports/snowball-bias-repaired/rl-snowball-e1-rno2a-ctx8k-p8-grug-67b-a-20260730-225618-8d52bd/global_step_5/policy/` - Experiment issue: https://github.com/marin-community/marin/issues/7786 - Evidence archive: https://huggingface.co/datasets/penfever/snowball-67b-a2b-math-rl-artifacts The scores and evaluation caveats are recorded in `MATH_EVALS.md` in the evidence archive. Preserve `config.json`, the tokenizer files, and all shards named by `model.safetensors.index.json` together.