Polish
adamo1139's picture
Create README.md
c6c8dfc verified
|
Raw History Blame Contribute Delete
2.26 kB
metadata
license: apache-2.0
datasets:
  - HuggingFaceFW/fineweb-2
  - adamo1139/fineweb2-pol
language:
  - pl

Pliki z pre-treningu LLMa Poziomka, w formacie DCP. Ten format pozwala na kontynuacj臋 treningu w 艂agodny spos贸b po przerwaniu, u偶ywaj膮c MegatronLM.

To repozytorium zawiera checkpointy z krok贸w 1600, 9600, 11200, 14400, 16000, 17600, 19200, 20800, 22400, 24000, 25600, 27200, 28800, 30400, 32000, 33600, 35200, 36800, 38400, 40000, 41600 oraz 43200.

Straci艂em checkpointy z krok贸w 3200, 4800, 6400 i 8000 - przekroczy艂em jakie艣 niewidoczne limity publicznej przestrzeni dyskowej kiery wrzuca艂em te pliki.

Ten model zosta艂 wytrenowany z global batch size 256 i d艂ugosci膮 sekwencji 8192, zatem 2,097,152 token贸w na krok. Planowa艂em go trenowa膰 przez 50000 krok贸w na 105 miliardach token贸w, ale sko艅czy艂y mi si臋 zasoby i ostatni checkpoint 43200 jest wytrenowany na 90.6 milarda token贸w. To oznacza r贸wnie偶, 偶e wsp贸艂czynnik uczenia nie zmniejszy艂 si臋 zgodnie z planem i model nie jest zahartowany. B臋d臋 u偶ywa艂 fuzji r贸偶nych checkpoint贸w aby zasymulowa膰 hartowanie bez dodatkowego trenowania tego checkpointa.

Po fuzji, model bedzie trenowany z SFT i mo偶e nawet z dystylacj膮 logit贸w.


Checkpoints of Poziomka LLM pre-training, in DCP format. This format allows for resuming the training in a smooth manner after pausing, using MegatronLM.

This repository contains checkpoints from steps 1600, 9600, 11200, 14400, 16000, 17600, 19200, 20800, 22400, 24000, 25600, 27200, 28800, 30400, 32000, 33600, 35200, 36800, 38400, 40000, 41600 and 43200.

You might wonder where checkpoints 3200, 4800, 6400 and 8000 went - I lost them due to hitting invisible public storage HF limit during checkpoint upload.

Model was trained with global batch size 256 and sequence length 8192, so 2,097,152 tokens per step. I planned to train it for 50000 iters (105B tokens) but I ran out of compute early, so the last checkpoint 43200 was trained on 90.6B tokens. This also means that the learning rate didn't decay as planned. I will be reconciling this using checkpoint merging that should be able to simulate the annealing.

After merging, this model will be trained with SFT and maybe even logit-level distillation.