| [probe] loading Qwen/Qwen3.6-27B (torch.bfloat16, sdpa-attn, chunked-score) |
| [transformers] The fast path is not available because one of the required library is not installed. Falling back to torch implementation. To install follow https://github.com/fla-org/flash-linear-attention |
|
Loading weights: 0%| | 0/1184 [00:00<?, ?it/s]
Loading weights: 0%| | 1/1184 [00:00<02:05, 9.46it/s]
Loading weights: 0%| | 2/1184 [00:00<02:41, 7.34it/s]
Loading weights: 3%|β | 37/1184 [00:00<00:08, 136.76it/s]
Loading weights: 5%|β | 65/1184 [00:00<00:06, 186.13it/s]
Loading weights: 8%|β | 95/1184 [00:00<00:04, 218.30it/s]
Loading weights: 10%|β | 121/1184 [00:00<00:04, 228.21it/s]
Loading weights: 13%|ββ | 152/1184 [00:00<00:04, 249.12it/s]
Loading weights: 16%|ββ | 185/1184 [00:00<00:03, 256.43it/s]
Loading weights: 19%|ββ | 222/1184 [00:01<00:03, 286.69it/s]
Loading weights: 21%|βββ | 253/1184 [00:01<00:03, 287.39it/s]
Loading weights: 24%|βββ | 283/1184 [00:01<00:03, 277.14it/s]
Loading weights: 26%|βββ | 311/1184 [00:01<00:03, 273.27it/s]
Loading weights: 29%|βββ | 344/1184 [00:01<00:02, 281.74it/s]
Loading weights: 32%|ββββ | 373/1184 [00:01<00:02, 275.73it/s]
Loading weights: 34%|ββββ | 406/1184 [00:01<00:02, 287.53it/s]
Loading weights: 37%|ββββ | 437/1184 [00:01<00:02, 285.85it/s]
Loading weights: 39%|ββββ | 466/1184 [00:01<00:02, 280.92it/s]
Loading weights: 42%|βββββ | 495/1184 [00:01<00:02, 278.92it/s]
Loading weights: 44%|βββββ | 523/1184 [00:02<00:02, 274.32it/s]
Loading weights: 47%|βββββ | 557/1184 [00:02<00:02, 286.09it/s]
Loading weights: 49%|βββββ | 586/1184 [00:02<00:02, 272.79it/s]
Loading weights: 52%|ββββββ | 617/1184 [00:02<00:02, 279.57it/s]
Loading weights: 55%|ββββββ | 646/1184 [00:02<00:01, 278.53it/s]
Loading weights: 57%|ββββββ | 677/1184 [00:02<00:01, 281.17it/s]
Loading weights: 60%|ββββββ | 706/1184 [00:02<00:01, 269.50it/s]
Loading weights: 62%|βββββββ | 735/1184 [00:02<00:01, 265.81it/s]
Loading weights: 65%|βββββββ | 767/1184 [00:02<00:01, 280.85it/s]
Loading weights: 67%|βββββββ | 796/1184 [00:03<00:01, 260.80it/s]
Loading weights: 70%|βββββββ | 832/1184 [00:03<00:01, 282.94it/s]
Loading weights: 76%|ββββββββ | 902/1184 [00:03<00:00, 398.14it/s]
Loading weights: 100%|ββββββββββ| 1184/1184 [00:03<00:00, 351.76it/s] |
| [probe] replaced lm_head with Identity (probe discards logits) |
| [probe] GA layers=16 num_h=24 num_kv=4 head_dim=256 |
| [probe] streaming up to 192 long fineweb docs (min_chars=1310720, scan cap=5000000)... |
| [probe] scanned 50000 rows, kept 46 (building doc, 975354/1310720 chars) |
| [probe] scanned 100000 rows, kept 95 (building doc, 169829/1310720 chars) |
| [probe] scanned 150000 rows, kept 142 (building doc, 636461/1310720 chars) |
| [probe] scanned 200000 rows, kept 191 (building doc, 146570/1310720 chars) |
| [probe] got 192 candidates |
| [transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (298196 > 262144). Running this sequence through the model will result in indexing errors |
| [probe] 1/64 samples scored |
| [probe] 2/64 samples scored |
| [probe] 4/64 samples scored |
| [probe] 6/64 samples scored |
| [probe] 8/64 samples scored |
| [probe] 10/64 samples scored |
| [probe] 12/64 samples scored |
| [probe] 14/64 samples scored |
| [probe] 16/64 samples scored |
| [probe] 18/64 samples scored |
| [probe] 20/64 samples scored |
| [probe] 22/64 samples scored |
| [probe] 24/64 samples scored |
| [probe] 26/64 samples scored |
| [probe] 28/64 samples scored |
| [probe] 30/64 samples scored |
| [probe] 32/64 samples scored |
| [probe] 34/64 samples scored |
| [probe] 36/64 samples scored |
| [probe] 38/64 samples scored |
| [probe] 40/64 samples scored |
| [probe] 42/64 samples scored |
| [probe] 44/64 samples scored |
| [probe] 46/64 samples scored |
| [probe] 48/64 samples scored |
| [probe] 50/64 samples scored |
| [probe] 52/64 samples scored |
| [probe] 54/64 samples scored |
| [probe] 56/64 samples scored |
| [probe] 58/64 samples scored |
| [probe] 60/64 samples scored |
| [probe] 62/64 samples scored |
| [probe] 64/64 samples scored |
| [probe] promoted 2 heads to satisfy min_retrieval_per_layer=1: [(3, 23), (63, 9)] |
| [probe] wrote probe_results/qwen3_6_27b_seq262k.json retrieval=58/384 = 0.151 |
|
|