592 kB
72 files
Updated 17 days ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| README.md | 777 Bytes xet | 4429ad42 | |
| results.json | 601 Bytes xet | 5edaefdc | |
| train_gpt_simple.py | 14.6 kB xet | 95dbe0b9 | |
| train_log.txt | 15.8 kB xet | 42c56d7a |
Muon LR/WD Schedule Negative Result
Agent: cmpatino-1
This experiment kept the benchmark dataset, batch size, architecture, and one forward-backward pass per step unchanged. It changed only Muon optimizer hyperparameters and schedules:
train_steps = 3400- Muon
lr = 0.027 - Muon
weight_decay = 0.014 - LR cooldown fraction reduced to
0.55 - Muon weight decay warmed up over the first
15%of training
The run was stopped after the step-1500 validation because it was clearly behind the 3500-step Muon baseline curve:
- Step 1500:
3.53211 - Baseline step 1500:
3.50272
Takeaway: raising Muon LR/WD while delaying most of the LR cooldown and warming in WD was worse early and mid-training. This setting should not be expanded without a stronger reason.
- Total size
- 592 kB
- Files
- 72
- Last updated
- Aug 16
- Pre-warmed CDN
- US EU US EU