| [ | |
| "Theorem 3.2 proves that any limit point of the normalized parameter trajectory ΞΈβ/βΞΈββ under normalized steepest descent is the direction of a KKT point of the corresponding max-margin problem (Theorem 3.2).", | |
| "Theorem 3.3 extends the margin-maximization convergence result to momentum steepest descent under decaying learning rates, using an Approximate Steepest Descent framework introduced in Definition 5.1 (Theorem 3.3, Definition 5.1).", | |
| "Corollary 3.4 shows Muon, applied simultaneously across weight matrices, is a special case of normalized momentum steepest descent with respect to the matrix spectral norm βΒ·β_msp (Corollary 3.4).", | |
| "Corollary 3.5 shows Muon-Signum hybrid algorithms exhibit margin maximization with respect to the max of the spectral norm and the ββ norm (Corollary 3.5).", | |
| "Theorem 3.6 shows Adam without a stability constant exhibits implicit bias toward ββ-margin maximization under a decaying learning rate regime when cβ β₯ cβ (Theorem 3.6).", | |
| "Theorem 3.1 establishes that under normalized steepest descent with a learning rate schedule satisfying β«β^β Ξ·(t)dt = β, the soft margin Ξ³Μ(ΞΈβ) increases monotonically (Theorem 3.1)." | |
| ] | |