Text Generation
Transformers
Safetensors
Polish
gpt2
polish
base-model
from-scratch
amd-rocm
continued-pretraining
Eval Results (legacy)
text-generation-inference
Instructions to use SlayerLab/GoLLeM-110M-PL-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/GoLLeM-110M-PL-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SlayerLab/GoLLeM-110M-PL-v3")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("SlayerLab/GoLLeM-110M-PL-v3") model = AutoModelForCausalLM.from_pretrained("SlayerLab/GoLLeM-110M-PL-v3", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SlayerLab/GoLLeM-110M-PL-v3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SlayerLab/GoLLeM-110M-PL-v3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SlayerLab/GoLLeM-110M-PL-v3", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/SlayerLab/GoLLeM-110M-PL-v3
- SGLang
How to use SlayerLab/GoLLeM-110M-PL-v3 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SlayerLab/GoLLeM-110M-PL-v3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SlayerLab/GoLLeM-110M-PL-v3", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SlayerLab/GoLLeM-110M-PL-v3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SlayerLab/GoLLeM-110M-PL-v3", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use SlayerLab/GoLLeM-110M-PL-v3 with Docker Model Runner:
docker model run hf.co/SlayerLab/GoLLeM-110M-PL-v3
File size: 6,957 Bytes
26baca5 ff16d04 26baca5 0009da0 2d69f4f 0009da0 2d69f4f 0009da0 2d69f4f 0009da0 2d69f4f 0009da0 2d69f4f 0009da0 2d69f4f 26baca5 ff16d04 26baca5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | ---
license: cc-by-sa-4.0
language:
- pl
library_name: transformers
pipeline_tag: text-generation
datasets:
- SlayerLab/gollem-corpus-2b-pl
- SlayerLab/polish-dynaword
tags:
- gpt2
- polish
- base-model
- from-scratch
- amd-rocm
- continued-pretraining
model-index:
- name: GoLLeM-110M-PL-v3
results:
- task: {type: text-classification, name: PolEmo2-IN sentiment}
dataset: {type: allegro/klej-polemo2-in, name: PolEmo2-IN}
metrics:
- {type: accuracy, value: 18.0, name: accuracy (0-shot; domain-PMI board-repro)}
- task: {type: text-classification, name: 8TAGS topic}
dataset: {type: sdadas/8tags, name: 8TAGS}
metrics:
- {type: accuracy, value: 41.2, name: accuracy (0-shot; domain-PMI board-repro)}
- task: {type: multiple-choice, name: Belebele PL reading}
dataset: {type: facebook/belebele, name: Belebele pol_Latn}
metrics:
- {type: accuracy, value: 24.0, name: accuracy (0-shot)}
- task: {type: text-classification, name: CBD cyberbullying}
dataset: {type: ptaszynski/PolishCyberbullyingDataset, name: CBD}
metrics:
- {type: f1, value: 12.6, name: macro-F1 (0-shot; domain-PMI board-repro)}
- task: {type: text-classification, name: DYK question-answer}
dataset: {type: allegro/klej-dyk, name: DYK}
metrics:
- {type: f1, value: 28.3, name: positive-F1 (0-shot; domain-PMI board-repro)}
- task: {type: token-classification, name: KLEJ NER}
dataset: {type: allegro/klej-nkjp-ner, name: KLEJ-NER}
metrics:
- {type: accuracy, value: 18.2, name: accuracy (0-shot; domain-PMI board-repro)}
- task: {type: text-classification, name: PSC summary}
dataset: {type: allegro/klej-psc, name: PSC}
metrics:
- {type: f1, value: 44.2, name: positive-F1 (0-shot; domain-PMI board-repro)}
---
# GoLLeM-110M-PL-v3
Polski model językowy **110M** (GPT-2-class), trenowany **od zera** na AMD Radeon RX 7900 XTX (ROCm/WSL2). Model **bazowy (completion)** — kontynuuje tekst, **nie jest chatbotem** (nie odpowiada na pytania; daj mu początek zdania).
Trzecia iteracja serii GoLLeM. **Główna zmiana vs v2: druga epoka na tym samym czystym korpusie ~2,0 mld** (podwojona ekspozycja, ~18 → ~36 tokenów/parametr) — korekta niedotrenowania v2.
> Model **completion**. Dobrze: `"Stolica Polski to"` · Źle: `"Jaka jest stolica Polski?"`
## Co nowego vs v2
- **Korekta niedotrenowania.** v2 widział korpus 1 raz (~18 tok/param). v3 to **kontynuacja pretreningu** przez 2. epokę (łącznie ~4 mld tokenów widzianych z tego samego 2,0 mld korpusu).
- **Zmierzony efekt:** wzrost na 7/9 zadań benchmarku (patrz Ewaluacja); największy na sentymencie, streszczeniach i NER.
- **Uczciwie:** to podwojona **ekspozycja na te same dane**, nie nowe dane.
## Trening
| | |
|---|---|
| Parametry | 110 025 216 (110M), weight-tied |
| Architektura | GPT-2 (decoder-only): 12 warstw / 12 głów / d_model 768 |
| Kontekst | 512 tokenów |
| Tokenizer | polski BPE (dynaword-32k), słownik 32 000, `<\|endoftext\|>`=0 |
| Dane | korpus v2 ~2,0 mld (58% curated: Wikipedia/Wikisource/Wolne Lektury/1000 Novels/Wiki\*/eltec + 42% HPLT v3 web clean; **zero legalese**), **2 epoki** (~4 mld tok widzianych) |
| Trening | kontynuacja z ckpt v2 (krok 60 733 → 121 466), bf16, AdamW, cosine LR + warmup, wd 0.1, batch 64 (grad-accum 4) |
| Sprzęt | 1× AMD Radeon RX 7900 XTX 24GB (gfx1100), ROCm/WSL2, ~9 h |
## Ewaluacja
Protokół: OpenPL (`polish4`, 0-shot) + **własna reprodukcja domain-PMI** OrisTeam (metoda KateMajzel: `ll(label|pełny) − ll(label|pusty-szablon)`). Nasza reprodukcja odtwarza tablicę [OrisTeam Polish-SLM-Benchmark](https://huggingface.co/spaces/OrisTeam/Polish-SLM-Benchmark) na **7/9 zadaniach w granicach ~1-2 pp** (belebele idealnie, tags8/cbd/dyk/klej_ner/polemo_in blisko). Ten sam scorer dla v2 i v3 → **delta jest wiarygodna**.
| zadanie | v2 | v3 | Δ |
|---|--:|--:|--:|
| PolEmo2-in | 16,2 | 18,0 | +1,8 |
| PolEmo2-out | 1,8\* | 10,7\* | +8,9 |
| 8tags | 41,3 | 41,2 | −0,1 |
| Belebele | 23,0 | 24,0 | +1,0 |
| CBD (hate) | 14,6 | 12,6 | **−2,0** |
| DYK | 23,4 | 28,3 | +4,9 |
| KLEJ-NER | 17,5 | 18,2 | +0,7 |
| PPC | 20,0\* | 25,7\* | +5,7 |
| PSC | 38,1 | 44,2 | +6,0 |
| **Sygnał-6** | **21,6** | **24,1** | **+2,55** |
| kompozyt-9 | 21,8 | 24,8 | +2,99 |
\* polemo_out i ppc: nasza reprodukcja domain-PMI odbiega od tablicy OrisTeam (polemo_out schodzi poniżej losowego — znany quirk scoringu, wg KateMajzel sygnał błędu bazy PMI); delty na tych zadaniach traktować ostrożnie.
**Wniosek:** v3 przewyższa v2 pod spójnym scoringiem (Sygnał-6 +2,55, kompozyt +2,99), rośnie na 7/9. **Regresja:** CBD (mowa nienawiści) −2,0. Pozycja na oficjalnej tablicy OrisTeam — do potwierdzenia przez zgłoszenie modelu do nich (autorytatywny scoring po ich stronie).
## Użycie
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
m = AutoModelForCausalLM.from_pretrained("Maggio33/GoLLeM-110M-PL-v3").eval()
t = AutoTokenizer.from_pretrained("Maggio33/GoLLeM-110M-PL-v3")
ids = t("Stolica Polski to", return_tensors="pt").input_ids
ids = torch.cat([torch.tensor([[0]]), ids], 1) # BOS = <|endoftext|>=0
out = m.generate(ids, max_new_tokens=80, do_sample=True, temperature=0.7,
top_k=40, repetition_penalty=1.3, pad_token_id=0)
print(t.decode(out[0].tolist()[1:], skip_special_tokens=True))
```
## Ograniczenia
- **110M = mały** → konfabuluje konkretne fakty; uczy się głównie płynności i formy polskiego.
- **Base/completion, nie chat** — do rozmowy potrzebny SFT/instruct-tuning.
- Kontekst 512 tokenów. Brak filtrów bezpieczeństwa na wyjściu.
- **PII** scrubowane w treningu (telefon/e-mail/PESEL/NIP → tagi); generowane imiona/adresy to konfabulacje.
- **CBD (mowa nienawiści) regresja vs v2** — do zastosowań wrażliwych na detekcję hate rozważ v2.
## Licencja i atrybucja
**Korpus = CC-BY-SA-4.0** (dominująca, share-alike): Wikipedia/Wikisource/Wiki\* — CC-BY-SA-3.0 (Wikimedia Foundation); Wolne Lektury — CC-BY-SA-4.0 / Wolna Sztuka 1.3; 1000 Novels, eltec_pol — CC-BY-4.0; HPLT v3 (web) — CC0-1.0. Użycie wymaga **ATTRIBUTION** (Wikimedia Foundation, Wolne Lektury, autorzy 1000 Novels, HPLT/CLARIN-PL) oraz **SHARE-ALIKE**. **Model:** CC-BY-SA-4.0.
## Podziękowania
Benchmark i protokół ewaluacyjny: **OrisTeam** ([Polish-SLM-Benchmark](https://huggingface.co/spaces/OrisTeam/Polish-SLM-Benchmark)); metoda kalibracji domain-PMI: **KateMajzel** ([gollem-pl](https://github.com/KateMajzel/gollem-pl)).
## Reprodukcja
Trening: `train_125m.py --run-id gollem_v3_e2b --data gollem_v2_train_32k.bin --epochs 2 --batch 64 --accum-steps 4 --lr 3e-4` (resume z ckpt v2). Dane treningowe jawnie: [`SlayerLab/gollem-corpus-2b-pl`](https://huggingface.co/datasets/SlayerLab/gollem-corpus-2b-pl) (dokładny korpus v2/v3). Ewaluacja: `board_eval.py` (domain-PMI). Ślad: repo `amd-torch`.
**Autor:** Arkadiusz Słota.
|