danym commited on
Commit
39f9b97
·
verified ·
1 Parent(s): 6a159b6

Upload stage1/qwen3_6_27b/train.log with huggingface_hub

Browse files
Files changed (1) hide show
  1. stage1/qwen3_6_27b/train.log +87 -0
stage1/qwen3_6_27b/train.log ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [stage1] loading model Qwen/Qwen3.6-27B (eager attn for output_attentions)
2
+ [transformers] The fast path is not available because one of the required library is not installed. Falling back to torch implementation. To install follow https://github.com/fla-org/flash-linear-attention#installation and https://github.com/Dao-AILab/causal-conv1d
3
+
4
+ [stage1] applying rtpurbo (probe=probe_results/qwen3_6_27b_seq262k.json)
5
+ [rtpurbo] layer 3: retrieval=[23] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22] top_p=0.05 window=512
6
+ [rtpurbo] layer 7: retrieval=[7, 8] local=[0, 1, 2, 3, 4, 5, 6, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
7
+ [rtpurbo] layer 11: retrieval=[14] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
8
+ [rtpurbo] layer 15: retrieval=[1, 18, 19, 21] local=[0, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 20, 22, 23] top_p=0.05 window=512
9
+ [rtpurbo] layer 19: retrieval=[6, 7, 8, 10, 11] local=[0, 1, 2, 3, 4, 5, 9, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
10
+ [rtpurbo] layer 23: retrieval=[1] local=[0, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
11
+ [rtpurbo] layer 27: retrieval=[10, 12, 14] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 13, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
12
+ [rtpurbo] layer 31: retrieval=[19, 21, 23] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 22] top_p=0.05 window=512
13
+ [rtpurbo] layer 35: retrieval=[16] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
14
+ [rtpurbo] layer 39: retrieval=[7, 23] local=[0, 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22] top_p=0.05 window=512
15
+ [rtpurbo] layer 43: retrieval=[12, 14, 16, 17] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 15, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
16
+ [rtpurbo] layer 47: retrieval=[1, 10, 13, 14, 18, 19, 20, 21, 23] local=[0, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 15, 16, 17, 22] top_p=0.05 window=512
17
+ [rtpurbo] layer 51: retrieval=[4, 6, 7, 8, 9, 11, 13, 16, 17, 18, 20, 22, 23] local=[0, 1, 2, 3, 5, 10, 12, 14, 15, 19, 21] top_p=0.05 window=512
18
+ [rtpurbo] layer 55: retrieval=[3, 20, 23] local=[0, 1, 2, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22] top_p=0.05 window=512
19
+ [rtpurbo] layer 59: retrieval=[3, 5, 14, 19, 21] local=[0, 1, 2, 4, 6, 7, 8, 9, 10, 11, 12, 13, 15, 16, 17, 18, 20, 22, 23] top_p=0.05 window=512
20
+ [rtpurbo] layer 63: retrieval=[9] local=[0, 1, 2, 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23] top_p=0.05 window=512
21
+ [rtpurbo] applied to 16 GA layers (identity=False)
22
+ [stage1] froze 1184 params, trainable=116 (indexers only)
23
+ [stage1] GA layers: [3, 7, 11, 15, 19, 23, 27, 31, 35, 39, 43, 47, 51, 55, 59, 63]
24
+ [stage1] running SVD warm-start (4 batches)
25
+ [svd-init] initialised 58 indexers from 4 calibration batches
26
+ [stage1] step 0/300 loss=28.4775 lr=5.00e-05 tok/s=846
27
+ [stage1] step 5/300 loss=27.0708 lr=3.00e-04 tok/s=4710
28
+ [stage1] step 10/300 loss=25.1618 lr=5.50e-04 tok/s=4705
29
+ [stage1] step 15/300 loss=23.1360 lr=8.00e-04 tok/s=4709
30
+ [stage1] step 20/300 loss=20.3149 lr=1.00e-03 tok/s=4709
31
+ [stage1] step 25/300 loss=17.6807 lr=9.99e-04 tok/s=4713
32
+ [stage1] step 30/300 loss=14.2037 lr=9.97e-04 tok/s=4716
33
+ [stage1] step 35/300 loss=12.0545 lr=9.93e-04 tok/s=4717
34
+ [stage1] step 40/300 loss=10.1546 lr=9.87e-04 tok/s=4710
35
+ [stage1] step 45/300 loss=8.5383 lr=9.80e-04 tok/s=4710
36
+ [stage1] step 50/300 loss=7.8494 lr=9.72e-04 tok/s=4710
37
+ [stage1] step 55/300 loss=6.1291 lr=9.62e-04 tok/s=4712
38
+ [stage1] step 60/300 loss=5.7260 lr=9.50e-04 tok/s=4710
39
+ [stage1] step 65/300 loss=5.4020 lr=9.38e-04 tok/s=4712
40
+ [stage1] step 70/300 loss=4.9078 lr=9.23e-04 tok/s=4698
41
+ [stage1] step 75/300 loss=3.8844 lr=9.08e-04 tok/s=4712
42
+ [stage1] step 80/300 loss=3.4554 lr=8.91e-04 tok/s=4710
43
+ [stage1] step 85/300 loss=3.4633 lr=8.73e-04 tok/s=4710
44
+ [stage1] step 90/300 loss=2.8089 lr=8.54e-04 tok/s=4711
45
+ [stage1] step 95/300 loss=2.4719 lr=8.33e-04 tok/s=4709
46
+ [stage1] step 100/300 loss=2.2695 lr=8.12e-04 tok/s=4709
47
+ [stage1] step 105/300 loss=2.2290 lr=7.89e-04 tok/s=4709
48
+ [stage1] step 110/300 loss=2.1210 lr=7.66e-04 tok/s=4707
49
+ [stage1] step 115/300 loss=1.9276 lr=7.42e-04 tok/s=4701
50
+ [stage1] step 120/300 loss=1.8633 lr=7.17e-04 tok/s=4703
51
+ [stage1] step 125/300 loss=2.1043 lr=6.91e-04 tok/s=4707
52
+ [stage1] step 130/300 loss=1.6939 lr=6.65e-04 tok/s=4704
53
+ [stage1] step 135/300 loss=1.6762 lr=6.38e-04 tok/s=4707
54
+ [stage1] step 140/300 loss=1.6860 lr=6.11e-04 tok/s=4706
55
+ [stage1] step 145/300 loss=1.7146 lr=5.84e-04 tok/s=4709
56
+ [stage1] step 150/300 loss=1.5755 lr=5.56e-04 tok/s=4708
57
+ [stage1] step 155/300 loss=1.8054 lr=5.28e-04 tok/s=4707
58
+ [stage1] step 160/300 loss=1.5244 lr=5.00e-04 tok/s=4709
59
+ [stage1] step 165/300 loss=1.7858 lr=4.72e-04 tok/s=4710
60
+ [stage1] step 170/300 loss=1.5372 lr=4.44e-04 tok/s=4711
61
+ [stage1] step 175/300 loss=1.4627 lr=4.16e-04 tok/s=4708
62
+ [stage1] step 180/300 loss=1.5078 lr=3.89e-04 tok/s=4708
63
+ [stage1] step 185/300 loss=1.4851 lr=3.62e-04 tok/s=4708
64
+ [stage1] step 190/300 loss=1.4827 lr=3.35e-04 tok/s=4710
65
+ [stage1] step 195/300 loss=1.4524 lr=3.09e-04 tok/s=4703
66
+ [stage1] step 200/300 loss=1.5130 lr=2.83e-04 tok/s=4708
67
+ [stage1] step 205/300 loss=1.5591 lr=2.58e-04 tok/s=4707
68
+ [stage1] step 210/300 loss=1.4889 lr=2.34e-04 tok/s=4707
69
+ [stage1] step 215/300 loss=1.4515 lr=2.11e-04 tok/s=4707
70
+ [stage1] step 220/300 loss=1.6439 lr=1.88e-04 tok/s=4707
71
+ [stage1] step 225/300 loss=1.5483 lr=1.67e-04 tok/s=4708
72
+ [stage1] step 230/300 loss=1.4530 lr=1.46e-04 tok/s=4710
73
+ [stage1] step 235/300 loss=1.4622 lr=1.27e-04 tok/s=4707
74
+ [stage1] step 240/300 loss=1.5380 lr=1.09e-04 tok/s=4711
75
+ [stage1] step 245/300 loss=1.4247 lr=9.22e-05 tok/s=4710
76
+ [stage1] step 250/300 loss=1.5552 lr=7.66e-05 tok/s=4709
77
+ [stage1] step 255/300 loss=1.4766 lr=6.24e-05 tok/s=4710
78
+ [stage1] step 260/300 loss=1.6120 lr=4.95e-05 tok/s=4708
79
+ [stage1] step 265/300 loss=1.4429 lr=3.81e-05 tok/s=4710
80
+ [stage1] step 270/300 loss=1.4547 lr=2.81e-05 tok/s=4709
81
+ [stage1] step 275/300 loss=1.4943 lr=1.95e-05 tok/s=4711
82
+ [stage1] step 280/300 loss=1.4348 lr=1.25e-05 tok/s=4707
83
+ [stage1] step 285/300 loss=1.3730 lr=7.06e-06 tok/s=4706
84
+ [stage1] step 290/300 loss=1.4000 lr=3.14e-06 tok/s=4706
85
+ [stage1] step 295/300 loss=1.4337 lr=7.87e-07 tok/s=4706
86
+ [stage1] step 299/300 loss=1.6138 lr=3.15e-08 tok/s=4709
87
+ [stage1] done -> checkpoints/stage1_qwen3_6_27b