GPT-2 (BabyLM deu-baseline-small recipe), German, vocab 16384 โ€” 10M words

German GPT-2 trained from scratch under the exact recipe of BabyLM-community/deu-baseline-small. The ONLY deviation is the vocabulary: 16384 instead of 8192 (a locally trained byte-level BPE, same construction as that corpus's baseline series). Everything else โ€” architecture and every training hyperparameter โ€” is reproduced field-for-field; see recipe.json in this repo for the machine-readable record.

Architecture (= official card, except vocab)

architecture GPT2LMHeadModel
n_layer / n_head 4 / 8
n_embd / n_inner 512 / 2048
n_positions = n_ctx 512
activation gelu
dropouts (attn/embd/resid) 0.1 / 0.1 / 0.1
initializer_range / ln eps 0.02 / 1e-5
vocab_size 16384 (official: 8192)
params 21.3M (official: 17.1M โ€” the difference is the embedding table)

Training (= official card, verbatim)

learning_rate 1e-4, linear decay (no warmup), 5 epochs, train batch 64, eval batch 8, seed 42, AdamW betas (0.9, 0.999) eps 1e-8, weight_decay 0, fp32, block 512, glosses of unspecified card fields follow HF defaults. Steps/epoch ~510 (official card: 4,944 at vocab 8192 on the 100M-word corpus โ€” the smaller count here reflects the larger vocabulary's lower fertility, not less data).

Data: German BabyLM corpus, 10M words (local reconstruction babylm_deu_10m.txt โ€” EN/DE-matched genre shares; NOT the official hub download). 1% held out for eval: final eval_loss 4.357. Note the official deu-baseline-small itself was trained on the 100M-word corpus (its "small" names the model, not the data), so the -100m repo of this pair is the direct-scale counterpart; its 3.56 vs the official 3.0813 is not an apples-to-apples number (different vocab size changes per-token entropy โ€” smaller vocab โ†’ lower per-token loss; different corpus reconstruction and eval split).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("yuanxin112/babylm-deu-bpe16k-gpt2-10m")
model = AutoModelForCausalLM.from_pretrained("yuanxin112/babylm-deu-bpe16k-gpt2-10m")
ids = tok("Die Frauen hรคtten mit ihrem Bruder", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=20, do_sample=False,
                                pad_token_id=tok.pad_token_id)[0],
                 skip_special_tokens=True))

Tokenizer: byte-level BPE, specials [PAD]=0 [UNK]=1 [BOS]=2 [EOS]=3 (unlike the official baseline's config, whose bos/eos fields carry the stale GPT-2 default 50256 that is not even in its 8192 vocab).

Downloads last month
138
Safetensors
Model size
21.3M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support