--- language: - el license: mit datasets: - alexliap/tinystories-gr pipeline_tag: text-generation tags: - greek - tinystories - llama - from-scratch --- # TinyStories-GR Llama 5M A from-scratch, Llama-3.2-architecture causal language model trained to generate plausible short children's stories in Modern Greek. Part of a size-ablation family trained on the same data, tokenizer, and recipe - only model size changes between them. Full architecture details are in this repo's `config.json`. ## Model family | Size | Params | Val loss | |---|---|---| | [120M](https://huggingface.co/alexliap/tinystories-gr-llama-120m) | 118.98M | 1.4402 | | [60M](https://huggingface.co/alexliap/tinystories-gr-llama-60m) | 59.98M | 1.5185 | | [30M](https://huggingface.co/alexliap/tinystories-gr-llama-30m) | 30.48M | 1.6646 | | [5M](https://huggingface.co/alexliap/tinystories-gr-llama-5m) | 4.82M | 2.1269 | ## Tokenizer A custom byte-level BPE tokenizer, vocab size 16640 (16384 learned merges + the 256-entry byte alphabet). Built from the Llama 3.2 tokenizer's tokenization pipeline (regex pre-tokenizer, ByteLevel encoding/decoding) but retrained from scratch on the Greek corpus below, keeping only `bos`/`eos` special tokens (not Llama's ~254 unused reserved/fine-tuning placeholder tokens). ## Training data [`alexliap/tinystories-gr`](https://huggingface.co/datasets/alexliap/tinystories-gr): 2,141,648 Greek translations of the TinyStories dataset. Trained for 1 epoch (~413M tokens) on sequences packed to 256 tokens (no padding, no truncation - long stories span multiple packed sequences). ## Sample generation Greedy decoding, run until the model emits `eos` on its own (capped at 256 tokens): **Prompt:** `Μια φορά κι έναν καιρό, ένα μικρό κορίτσι` **Completion** (95 tokens, ended on eos): > Μια φορά κι έναν καιρό, ένα μικρό κορίτσι που το έλεγαν Λίλυ πήγε στο πάρκο για να παίξει. Είδε ένα αγόρι που το έλεγαν Τιμ. Ο Τιμ ήταν λυπημένος γιατί δεν μπορούσε να παίξει με το αγόρι. Η Λίλυ ήθελε να βοηθήσει τον Τιμ, έτσι του είπε: «Γεια σου, Τιμ! Μπορώ να σε βοηθήσω να βρεις το παιχνίδι σου». > > Ο Τιμ χάρηκε πολύ και ευχαρίστησε τη Λίλυ. Έπαιξαν μαζί και πέρασαν υπέροχα. Από εκείνη την ημέρα, η Λίλυ και ο Τιμ έγιναν καλοί φίλοι και έπαιζαν μαζί κάθε μέρα. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "alexliap/tinystories-gr-llama-5m" tokenizer = AutoTokenizer.from_pretrained(repo_id) model = AutoModelForCausalLM.from_pretrained(repo_id) prompt = "Μια φορά κι έναν καιρό," input_ids = tokenizer(prompt, return_tensors="pt").input_ids output = model.generate(input_ids, max_new_tokens=200, do_sample=False, eos_token_id=tokenizer.eos_token_id) print(tokenizer.decode(output[0], skip_special_tokens=True)) ``` ## Limitations Trained on a single narrow domain (short, simple children's stories) for one epoch - not a general-purpose language model. Expect fluent, grammatical Modern Greek within that domain, and degraded coherence/factuality well outside of it.