Quazim0t0 commited on
Commit
786f87d
·
verified ·
1 Parent(s): 58e94b5

15M sibling is the toy/smoke-test for this run

Browse files
Files changed (1) hide show
  1. README.md +4 -2
README.md CHANGED
@@ -41,8 +41,9 @@ The base was trained on 2.0B tokens. SFT adds 0.23B, DPO adds 0.05B.
41
 
42
  A smaller looped sibling lives in
43
  [Byrne-15M-Looped](https://huggingface.co/Quazim0t0/Byrne-15M-Looped)
44
- (~15.2M active, `loop_count=3`). That one is the loop A/B. This one is
45
- the Memory Cache / fractal RoPE run. Not a matched scale-up.
 
46
 
47
  ## Checkpoints
48
 
@@ -1201,6 +1202,7 @@ The results of the follow-up tests will be shared soon.
1201
 
1202
  The small looped run is its own repo:
1203
  [Quazim0t0/Byrne-15M-Looped](https://huggingface.co/Quazim0t0/Byrne-15M-Looped).
 
1204
 
1205
  ~15.2M active / ~18.9M total, 10 layers, `loop_count=3` (effective
1206
  depth 30), MoE. No Memory Cache. I trained it to see if looping the
 
41
 
42
  A smaller looped sibling lives in
43
  [Byrne-15M-Looped](https://huggingface.co/Quazim0t0/Byrne-15M-Looped)
44
+ (~15.2M active, `loop_count=3`). That was the toy / smoke-test for
45
+ this run: looping at 15M before I spent the 114M budget. Not a
46
+ matched scale-up.
47
 
48
  ## Checkpoints
49
 
 
1202
 
1203
  The small looped run is its own repo:
1204
  [Quazim0t0/Byrne-15M-Looped](https://huggingface.co/Quazim0t0/Byrne-15M-Looped).
1205
+ That was the toy / smoke-test for this 114M model.
1206
 
1207
  ~15.2M active / ~18.9M total, 10 layers, `loop_count=3` (effective
1208
  depth 30), MoE. No Memory Cache. I trained it to see if looping the