trixyL/simplestories-4k-megatron
Updated β’ 8 β’ 1
This is the result of the code from https://github.com/triloy8/transformerlm, a minimal autoregressive Transformer LM trained on SimpleStories with a 2048-token context and an 4K vocab tokenizer. β¨
To reproduce the run:
Exact commit that launched the train: https://github.com/triloy8/transformerlm/commit/7ae6b9ba829a3ad37d85d2d8d201c717a423fb73