Parallax-Experiment Experimental Parallax (Parameterized Local Linear Attention) models. YifeiZuo/Parallax-0.6B Text Generation • 0.7B • Updated Jun 10 • 16 YifeiZuo/Parallax-1.7B Text Generation • 2B • Updated Jun 10 • 16
Attention 0.6B AdamW-WSD training trajectory Per-step record (every 500 steps, 40 ckpts) of the 0.6B Qwen3 softmax-attention baseline trained AdamW + WSD on 80B tokens. YifeiZuo/attention-0.6b-adamw-wsd-step500 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step1000 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step1500 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step2000 0.6B • Updated May 10
Parallax-Experiment Experimental Parallax (Parameterized Local Linear Attention) models. YifeiZuo/Parallax-0.6B Text Generation • 0.7B • Updated Jun 10 • 16 YifeiZuo/Parallax-1.7B Text Generation • 2B • Updated Jun 10 • 16
Attention 0.6B AdamW-WSD training trajectory Per-step record (every 500 steps, 40 ckpts) of the 0.6B Qwen3 softmax-attention baseline trained AdamW + WSD on 80B tokens. YifeiZuo/attention-0.6b-adamw-wsd-step500 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step1000 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step1500 0.6B • Updated May 10 YifeiZuo/attention-0.6b-adamw-wsd-step2000 0.6B • Updated May 10