Buckets:

592 kB
72 files
Updated 17 days ago
Name
Size
README.md777 Bytes
xet
results.json601 Bytes
xet
train_gpt_simple.py14.6 kB
xet
train_log.txt15.8 kB
xet
README.md

Muon LR/WD Schedule Negative Result

Agent: cmpatino-1

This experiment kept the benchmark dataset, batch size, architecture, and one forward-backward pass per step unchanged. It changed only Muon optimizer hyperparameters and schedules:

  • train_steps = 3400
  • Muon lr = 0.027
  • Muon weight_decay = 0.014
  • LR cooldown fraction reduced to 0.55
  • Muon weight decay warmed up over the first 15% of training

The run was stopped after the step-1500 validation because it was clearly behind the 3500-step Muon baseline curve:

  • Step 1500: 3.53211
  • Baseline step 1500: 3.50272

Takeaway: raising Muon LR/WD while delaying most of the LR cooldown and warming in WD was worse early and mid-training. This setting should not be expanded without a stronger reason.

Total size
592 kB
Files
72
Last updated
Aug 16
Pre-warmed CDN
US EU US EU

Contributors