jkminder commited on
Commit
f0baf65
·
verified ·
1 Parent(s): 8548da3

Scaling Ladder d26_973m_seed2_sft main: card refresh

Browse files
Files changed (1) hide show
  1. README.md +33 -3
README.md CHANGED
@@ -69,6 +69,34 @@ same harness across all runs and sizes). "SFT val bpb" is the run's final
69
  validation loss (bits per byte) on the mixture's held-out split. A "-"
70
  means that run's eval has not landed yet.
71
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
72
  ## Usage
73
 
74
  The chat template is bundled; format conversations with
@@ -95,9 +123,11 @@ transformers class, so the modeling code ships in the repository
95
  `<|assistant_end|>`; sampling defaults (temperature 0.6, top_k 50) ship in
96
  `generation_config.json`. The template renders nanochat's chat format
97
  token-for-token (a leading system message is merged into the first user
98
- message); conversion is verified per revision by chat-template, logit and
99
- loss equivalence against the original training code
100
- (`verify_results.json`, where present).
 
 
101
 
102
  ## Architecture, tokenizer, training data
103
 
 
69
  validation loss (bits per byte) on the mixture's held-out split. A "-"
70
  means that run's eval has not landed yet.
71
 
72
+ ## Anneal-mark chat-SFTs
73
+
74
+ The base repository also holds the model annealed at every mark of its
75
+ pretraining run: base revision `TPP_x` is the model annealed at x tokens per
76
+ parameter (`TPP_200` is the base `main`). The revisions below apply the same
77
+ chat-SFT recipe to each of those marks, one run per mark (SFT data seed 0,
78
+ replicate 1). Their names carry the mark with three digits (`TPP_010` ..
79
+ `TPP_180`); each row links the base revision it was trained from. These runs
80
+ were trained at code commit `a06bf32`, which clamps the SFT
81
+ learning-rate schedule so the last step cannot run at a negative learning
82
+ rate; the TPP_200 chat-SFTs above predate that fix, and their one final step
83
+ ran at a slightly negative learning rate (about -0.0004x the peak, read from
84
+ a d12 run's log; the factor depends on the step count).
85
+
86
+ | revision | base revision (annealed at) | base step | SFT step | SFT val bpb | trained on |
87
+ |---|---|---|---|---|---|
88
+ | TPP_010 | [`TPP_10`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_10) (10 tokens per parameter) | 8477 | 467 | 0.2829 | squirtle |
89
+ | TPP_020 | [`TPP_20`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_20) (20 tokens per parameter) | 17477 | 467 | 0.2748 | bulbasaur |
90
+ | TPP_030 | [`TPP_30`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_30) (30 tokens per parameter) | 25977 | 467 | 0.2726 | squirtle |
91
+ | TPP_040 | [`TPP_40`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_40) (40 tokens per parameter) | 34977 | 467 | 0.2713 | bulbasaur |
92
+ | TPP_060 | [`TPP_60`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_60) (60 tokens per parameter) | 52477 | 467 | 0.2694 | squirtle |
93
+ | TPP_080 | [`TPP_80`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_80) (80 tokens per parameter) | 69977 | 467 | 0.2681 | squirtle |
94
+ | TPP_100 | [`TPP_100`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_100) (100 tokens per parameter) | 87477 | 467 | 0.2668 | bulbasaur |
95
+ | TPP_120 | [`TPP_120`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_120) (120 tokens per parameter) | 104977 | 467 | 0.2656 | bulbasaur |
96
+ | TPP_140 | [`TPP_140`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_140) (140 tokens per parameter) | 122477 | 467 | 0.2644 | bulbasaur |
97
+ | TPP_160 | [`TPP_160`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_160) (160 tokens per parameter) | 139977 | 467 | 0.2638 | bulbasaur |
98
+ | TPP_180 | [`TPP_180`](https://huggingface.co/jkminder/d26_973m_seed2/tree/TPP_180) (180 tokens per parameter) | 157477 | 467 | 0.2642 | bulbasaur |
99
+
100
  ## Usage
101
 
102
  The chat template is bundled; format conversations with
 
123
  `<|assistant_end|>`; sampling defaults (temperature 0.6, top_k 50) ship in
124
  `generation_config.json`. The template renders nanochat's chat format
125
  token-for-token (a leading system message is merged into the first user
126
+ message). Every revision's upload is byte-verified against the converted
127
+ checkpoint (hub listing sizes and content hashes); the conversion itself is
128
+ verified on at least one revision per repository by chat-template, logit
129
+ and loss equivalence against the original training code — a revision that
130
+ was verified carries the record `verify_results.json`.
131
 
132
  ## Architecture, tokenizer, training data
133