oza75 commited on
Commit
78e5c05
·
verified ·
1 Parent(s): ed1595b

docs: add model/dataset card

Browse files
Files changed (1) hide show
  1. README.md +477 -5
README.md CHANGED
@@ -5,11 +5,483 @@ language:
5
  base_model:
6
  - FacebookAI/xlm-roberta-base
7
  pipeline_tag: fill-mask
 
 
 
 
 
 
 
 
8
  ---
9
 
10
- # Model Card for Model ID
11
 
12
- This model is a fine-tuned version of [FacebookAI/xlm-roberta-base](https://huggingface.co/FacebookAI/xlm-roberta-base) on an 11M words Bambara dataset.
13
- It achieves the following results on the evaluation set:
14
- - Loss: 1.1160
15
- - Accuracy: 0.7599
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5
  base_model:
6
  - FacebookAI/xlm-roberta-base
7
  pipeline_tag: fill-mask
8
+ tags:
9
+ - xlm-roberta
10
+ - fill-mask
11
+ - masked-lm
12
+ - bambara
13
+ - bamanankan
14
+ - mali
15
+ - vocabulary-extension
16
  ---
17
 
18
+ # bm-xlm-roberta-base
19
 
20
+ A Bambara (Bamanankan, `bm`) adaptation of **XLM-RoBERTa base**, continued-pretrained with the
21
+ masked-language-modelling objective and shipped with a **vocabulary-extended tokenizer** (253,421
22
+ tokens vs. 250,002 in stock XLM-R).
23
+
24
+ > [!IMPORTANT]
25
+ > **This is a masked-LM checkpoint, not a task model.** The architecture is
26
+ > `XLMRobertaForMaskedLM`. It predicts `<mask>` and produces contextual representations. It does
27
+ > **not** classify, tag, translate, or embed sentences out of the box. Use it as an *initialisation*
28
+ > for fine-tuning on sequence classification, token classification (NER/POS), extractive QA, or as
29
+ > the encoder backbone of a retrieval model.
30
+
31
+ The headline reason to prefer this over stock `FacebookAI/xlm-roberta-base` for Bambara is
32
+ tokenization: on a 5,000-sentence Bambara sample it produces **30.0% fewer tokens** (measurement and
33
+ method below), which directly translates to shorter sequences, more real context inside the 512-token
34
+ window, and cheaper fine-tuning.
35
+
36
+ ## Model details
37
+
38
+ | | |
39
+ |---|---|
40
+ | Repository | `djelia/bm-xlm-roberta-base` (public) |
41
+ | Architecture | `XLMRobertaForMaskedLM` (encoder-only, absolute position embeddings) |
42
+ | Parameters | 280,924,397 (all `F32`) |
43
+ | Layers / hidden / heads / FFN | 12 / 768 / 12 / 3072 |
44
+ | Activation | `gelu` |
45
+ | `max_position_embeddings` | 514 (usable sequence length **512**; `model_max_length: 512`) |
46
+ | `vocab_size` | **253,421** |
47
+ | `type_vocab_size` | 1 |
48
+ | `tie_word_embeddings` | `true` (LM head decoder is tied to the input embedding matrix) |
49
+ | Special token ids | `<s>`=0, `<pad>`=1, `</s>`=2, `<unk>`=3, **`<mask>`=253420** |
50
+ | Files | `config.json`, `model.safetensors` (1,123,721,924 B), `tokenizer.json` (17,001,775 B), `tokenizer_config.json` |
51
+ | Total repo storage | 1,140,723,699 B (~1.14 GB) |
52
+ | Created / last modified | 2026-02-04 (both; all four commits authored by `oza75`) |
53
+ | `transformers_version` in config | `5.0.0` |
54
+
55
+ All of the above is read from the repo's `config.json`, `tokenizer_config.json`, and the Hub API
56
+ model record.
57
+
58
+ ### Lineage
59
+
60
+ The repo is tagged `base_model: FacebookAI/xlm-roberta-base`, which is the ultimate ancestor: the
61
+ first 250,001 vocabulary entries and their ids are inherited unchanged from stock XLM-R (verified by
62
+ an ordered vocabulary comparison, see below).
63
+
64
+ The immediate provenance is a training run in the author's personal namespace. The
65
+ `model.safetensors` in this repo is **byte-identical** (SHA-256
66
+ `025cd2563f9aa71c26999eed79c0115b400ce29bbe033a9026312032a2bbd9c2`) to the one in
67
+ [`oza75/xlm-roberta-bambara-01`](https://huggingface.co/oza75/xlm-roberta-bambara-01). I compared the
68
+ LFS SHA-256 of `model.safetensors` across every `oza75/*xlm-roberta*bambara*` repo; that one matched
69
+ and the other six did not. This repo was uploaded first (2026-02-04); the `oza75` copy was published
70
+ later (2026-04-23) *with* the Trainer-generated metadata that this repo omits.
71
+
72
+ The sibling repo declares its own base as `oza75/xlm-roberta-bambara-base-01`, which in turn declares
73
+ `oza75/xlm-roberta-bambara-initial` — so the real chain is a multi-stage adaptation
74
+ (`FacebookAI/xlm-roberta-base` → vocabulary extension + continued pretraining stages → this
75
+ checkpoint), not a single fine-tune. So this is **not** a re-publication of a different model: it is
76
+ the same artifact as `oza75/xlm-roberta-bambara-01`, published to the org namespace first.
77
+
78
+ ## Tokenizer: what the 3,419 added tokens actually are
79
+
80
+ This is the most substantive difference from stock XLM-R, so it is worth stating precisely. I
81
+ downloaded both `tokenizer.json` files and compared them directly.
82
+
83
+ **Structure of the change.** The stock 250,002-entry vocabulary is preserved *in its original order*
84
+ at ids 0–250,000. The 3,419 new tokens were **appended** at ids 250,001–253,419, and stock's `<mask>`
85
+ was **moved from id 250,001 to id 253,420** (the only ordered mismatch in the overlapping range). No
86
+ stock token was removed.
87
+
88
+ > [!WARNING]
89
+ > **Token ids are not interchangeable with stock XLM-R.** Ids 0–250,000 match, but `<mask>` moved
90
+ > from 250001 → 253420. Never mix this tokenizer with a stock XLM-R checkpoint or vice versa. Always
91
+ > load the tokenizer from this repo.
92
+
93
+ **Composition of the 3,419 added tokens** (my count over `tokenizer.json`):
94
+
95
+ | Property | Count | Share |
96
+ |---|---|---|
97
+ | Word-initial (begin with the `▁` metaspace marker) | 2,541 | 74.3% |
98
+ | Contain a Bambara-specific letter (`ɛ ɔ ɲ ŋ` / `Ɛ Ɔ Ɲ Ŋ`) | 1,333 | 39.0% |
99
+ | Pure ASCII-Latin (e.g. `▁taa`, `▁jama`, `▁kelen`, `baara`, `muso`) | 2,042 | 59.7% |
100
+ | Punctuation-bearing (`▁,`, `▁.`, `▁?..`, `▁....`, and the odd `##`) | 38 | 1.1% |
101
+ | Digit-bearing | 0 | 0% |
102
+ | Single-character (`ԑ`, `Ɛ`, `Ŋ`, `Ɲ`, and a stray `̸`) | 5 | 0.1% |
103
+
104
+ Token length (stripped of `▁`) peaks at 4–5 characters; the additions are overwhelmingly Bambara
105
+ *words and morphemes*, not characters. Representative examples: `▁bɛ`, `▁tɛ`, `▁kɛ`, `▁bɛɛ`, `▁fɔ`,
106
+ `▁mɔgɔ`, `▁jama`, `▁kelen`, `▁yɔrɔ`, `▁cogo`, `▁baara`, `▁muso`, plus productive suffix pieces
107
+ (`kɔrɔ`, `tigi`, `ya`, `nba`) and Malian proper nouns (`▁Ɲafunke`, `▁Ɔsitarali`, `▁Ɛsipaɲi`).
108
+
109
+ **So it is both.** The additions are dominated by Bambara whole words (which stock XLM-R has no
110
+ entries for at all), and they *also* close a genuine character-coverage hole:
111
+
112
+ > The capital letters **`Ɛ` (U+0190), `Ɲ` (U+01F8), and `Ŋ` (U+014A) are absent from the stock
113
+ > XLM-R vocabulary entirely** — I confirmed that stock encodes each of them to `<unk>` (id 3).
114
+ > Lowercase `ɛ ɔ ɲ ŋ` and capital `Ɔ` are present in stock, but only as isolated single characters,
115
+ > so stock shreds every Bambara word containing them.
116
+
117
+ A concrete side-by-side (`add_special_tokens=True`):
118
+
119
+ ```
120
+ Input: "Mɔgɔw bɛɛ bɛ bange hɔrɔnya ni danbe ni josira lakanani na."
121
+
122
+ this repo (20 tokens):
123
+ <s> ▁Mɔgɔ w ▁bɛɛ ▁bɛ ▁bange ▁hɔrɔn ya ▁ni ▁dan be ▁ni ▁jo sira ▁laka na ni ▁na . </s>
124
+
125
+ stock XLM-R (30 tokens):
126
+ <s> ▁M ɔ g ɔ w ▁b ɛ ɛ ▁b ɛ ▁bang e ▁h ɔ r ɔ nya ▁ni ▁dan be ▁ni ▁jos ira ▁la kan ani ▁na . </s>
127
+ ```
128
+
129
+ ### Measured tokenization benefit
130
+
131
+ **This is my own measurement, not a figure taken from any published card.** Method, stated so it is
132
+ reproducible:
133
+
134
+ - **Corpus:** [`djelia/bambara-texts`](https://huggingface.co/datasets/djelia/bambara-texts),
135
+ file `data/train-00000-of-00001.parquet`, column `text` (84,328 rows).
136
+ - **Sample:** rows with non-null, non-empty `text`, then
137
+ `pandas.Series.sample(n=5000, random_state=0)` — a fixed seed, so the exact sample is reproducible.
138
+ - **Tokenizers:** `tokenizers.Tokenizer.from_file()` on `tokenizer.json` from this repo and from
139
+ `FacebookAI/xlm-roberta-base`, `encode_batch(..., add_special_tokens=False)`.
140
+ - **Sample size:** 4,614,792 characters / 917,625 whitespace-delimited words.
141
+
142
+ | Metric | `djelia/bm-xlm-roberta-base` | stock `xlm-roberta-base` |
143
+ |---|---|---|
144
+ | Total tokens | **1,454,892** | 2,079,639 |
145
+ | Tokens per word | **1.585** | 2.266 |
146
+ | Characters per token | **3.172** | 2.219 |
147
+ | `<unk>` tokens (id 3) | **48** (0.0033%) | 1,382 (0.0665%) |
148
+ | Sentences containing ≥1 `<unk>` | **9** / 5,000 | 556 / 5,000 |
149
+
150
+ **Corpus-level reduction: 30.04% fewer tokens.** Per sentence, the reduction is a median of 30.00%
151
+ (mean 28.78%, p10 20.72%, p90 35.43%), and **this tokenizer was shorter-or-equal on 100.0% of the
152
+ 5,000 sentences** — there is no regression tail.
153
+
154
+ A useful sanity check on whether the new vocabulary is actually load-bearing: **29.38% of all tokens
155
+ produced on this Bambara corpus come from the 3,419 added ids.** 1.3% of the vocabulary carries
156
+ nearly 30% of the token mass.
157
+
158
+ > [!NOTE]
159
+ > The corpus is a mixed-provenance Bambara collection (largest sources: `saillab/alpaca-bambara-cleaned`,
160
+ > `bigscience/xP3all`, `cis-lmu/GlotCC-V1`, `facebook/belebele`). The 30% figure describes *this*
161
+ > distribution of Bambara text. Orthographically non-standard or heavily code-switched Bambara will
162
+ > land differently.
163
+
164
+ ### Tokenizer caveats
165
+
166
+ > [!WARNING]
167
+ > **The Unicode normalizer was replaced, and it no longer normalizes Unicode.** Stock XLM-R uses
168
+ > SentencePiece's `Precompiled` charsmap (NFKC-style folding). This tokenizer instead uses
169
+ > `Sequence[Strip(right), Replace(" {2,}" → "▁")]`. Two observed consequences:
170
+ >
171
+ > - **NFC vs. NFD input tokenizes differently.** `"ka nà"` in NFC → `▁n`, `à`; the same string in
172
+ > NFD → `▁na`, `̀`. Stock produces the identical result for both. **Normalize your text to NFC
173
+ > yourself** before tokenizing, at both fine-tuning and inference time, or you will get silent
174
+ > train/serve skew.
175
+ > - **Compatibility characters are not folded.** `①` and `fi` pass through as-is (stock maps them to
176
+ > `1` and `fi`).
177
+
178
+ > [!WARNING]
179
+ > `tokenizer_config.json` declares `"tokenizer_class": "XLMRobertaTokenizer"` and `"use_fast": false`,
180
+ > but the repo ships **only `tokenizer.json`** — there is no `sentencepiece.bpe.model`. I verified the
181
+ > tokenizer loads correctly on **`transformers` 5.14.1**. On older `transformers` 4.x, where
182
+ > `use_fast: false` would attempt to build the SentencePiece-backed slow tokenizer, loading may fail;
183
+ > pass `use_fast=True` explicitly if you hit that.
184
+
185
+ ## Are the new embeddings actually trained?
186
+
187
+ Yes. This matters — a vocabulary-extended checkpoint whose new rows were never trained is worthless —
188
+ so I checked it rather than assuming.
189
+
190
+ The embedding matrix **was resized correctly**: `roberta.embeddings.word_embeddings.weight` has shape
191
+ `[253421, 768]`, read from the `model.safetensors` header. Four independent pieces of evidence say the
192
+ added rows carry learned signal:
193
+
194
+ 1. **Magnitude is inconsistent with fresh initialisation.** Rows 250,001–253,420 have element-wise
195
+ σ = 0.1157 and mean L2 norm 3.204 (σ 0.273). `config.json` sets `initializer_range: 0.02`, so a
196
+ freshly-initialised row would have norm ≈ 0.02·√768 ≈ **0.55** — off by ~6×.
197
+ 2. **Norms track token frequency.** Correlation between a token's unigram score in `tokenizer.json`
198
+ and its embedding norm is **r = 0.818** over the added ids (vs. r = 0.277 over the stock ids).
199
+ Randomly initialised rows cannot correlate with corpus frequency.
200
+ 3. **The LM head bias for the new ids is non-zero and varied**: over ids 250,001+, mean −0.011,
201
+ σ 0.122, range [−2.644, 0.347], with **zero** exactly-zero entries.
202
+ 4. **Nearest neighbours are semantically coherent** (cosine similarity over the full L2-normalised
203
+ embedding matrix):
204
+
205
+ | Token | Top neighbours |
206
+ |---|---|
207
+ | `▁bɛ` ("is") | `▁tɛ` (0.75) — *its own negative form* — `▁Bɛ`, `bɛ`, `▁Nbɛ` |
208
+ | `▁mɔgɔ` ("person") | `▁Mogo`, `mɔgɔ`, `▁muso` ("woman"), `▁Krɛcɛn` |
209
+ | `▁yɔrɔ` ("place") | `▁Cɔgɔ`/`▁cogo` ("manner"), `▁waati` ("time"), `▁Yɔrɔ` |
210
+ | `▁kelen` ("one") | `kelen`, `▁Saaba` ("three"), `▁wɔɔrɔ` ("six"), `▁Kelen` |
211
+ | `▁Ɲafunke` (Malian town) | a tight cluster of other Malian toponyms at 0.79–0.83 — `▁Direyi`, `▁Yuwaru`, `▁Lere`, `▁Duwanza`, `▁Kemasina`, `▁Burɛmu` |
212
+
213
+ Numerals cluster with numerals, toponyms with toponyms, and a copula with its own negation. That
214
+ structure cannot arise from noise.
215
+
216
+ > [!NOTE]
217
+ > One observation worth recording without over-reading it: the added rows sit at a systematically
218
+ > **lower norm than the stock rows** (mean 3.204 vs. 5.505) and are much tighter (σ 0.273 vs. 0.432).
219
+ > This is the expected signature of embeddings that received fewer gradient updates than the
220
+ > pretrained ones — they are trained, but for less wall-clock than the original XLM-R vocabulary. No
221
+ > action is implied; it is simply the shape of the checkpoint.
222
+
223
+ ## Usage
224
+
225
+ ### Fill-mask
226
+
227
+ XLM-R's mask token is **`<mask>`** (not `[MASK]`). In this checkpoint its id is **253420**.
228
+
229
+ ```python
230
+ from transformers import pipeline
231
+
232
+ fill = pipeline("fill-mask", model="djelia/bm-xlm-roberta-base")
233
+
234
+ # "All people are born in freedom __ dignity." (expected fill: "ni" = "and")
235
+ for p in fill("Mɔgɔw bɛɛ bɛ bange hɔrɔnya <mask> danbe la."):
236
+ print(f"{p['score']:.4f} {p['token_str']!r} -> {p['sequence']}")
237
+ ```
238
+
239
+ Or explicitly, without the pipeline:
240
+
241
+ ```python
242
+ import torch
243
+ from transformers import AutoTokenizer, AutoModelForMaskedLM
244
+
245
+ model_id = "djelia/bm-xlm-roberta-base"
246
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
247
+ model = AutoModelForMaskedLM.from_pretrained(model_id).eval()
248
+
249
+ text = "Mɔgɔw bɛɛ bɛ bange hɔrɔnya <mask> danbe la."
250
+ inputs = tokenizer(text, return_tensors="pt")
251
+
252
+ with torch.no_grad():
253
+ logits = model(**inputs).logits
254
+
255
+ mask_index = (inputs.input_ids == tokenizer.mask_token_id).nonzero()[0, 1]
256
+ top = logits[0, mask_index].topk(5)
257
+
258
+ for score, token_id in zip(top.values.tolist(), top.indices.tolist()):
259
+ print(f"{score:7.3f} {tokenizer.decode([token_id])!r}")
260
+ ```
261
+
262
+ Both snippets were executed against this checkpoint (`transformers` 5.14.1, `torch` 2.13.0). The
263
+ actual output for that sentence — the opening of the UDHR Article 1 in Bambara, *"every person is
264
+ born in freedom \_\_ dignity"* — is:
265
+
266
+ | Rank | Token | Probability |
267
+ |---|---|---|
268
+ | 1 | `ni` ("and") | 0.875 |
269
+ | 2 | `ani` ("and", alt. form) | 0.097 |
270
+ | 3 | `,` | 0.004 |
271
+ | 4 | `kɔnɔ` ("in/inside") | 0.003 |
272
+ | 5 | `la` (locative postposition) | 0.003 |
273
+
274
+ The top two predictions are both correct Bambara conjunctions, and rank 1 is the word actually used
275
+ in the source text — a sanity check that the checkpoint is intact and the mask token is right, not a
276
+ benchmark.
277
+
278
+ ### Contextual embeddings
279
+
280
+ ```python
281
+ import torch
282
+ from transformers import AutoTokenizer, AutoModel
283
+
284
+ tokenizer = AutoTokenizer.from_pretrained("djelia/bm-xlm-roberta-base")
285
+ encoder = AutoModel.from_pretrained("djelia/bm-xlm-roberta-base").eval()
286
+
287
+ batch = tokenizer(["Aw ni ce.", "I ka kɛnɛ wa?"], padding=True, return_tensors="pt")
288
+ with torch.no_grad():
289
+ hidden = encoder(**batch).last_hidden_state # (batch, seq_len, 768)
290
+ ```
291
+
292
+ > [!WARNING]
293
+ > **Do not use `pooler_output`.** This checkpoint contains **no pooler weights**. I diffed its tensor
294
+ > key set against stock `FacebookAI/xlm-roberta-base` (after normalising the `gamma`/`beta` naming
295
+ > below) and the *only* difference is that `roberta.pooler.dense.weight` and
296
+ > `roberta.pooler.dense.bias` are **absent here**. Confirmed at runtime: `AutoModel.from_pretrained`
297
+ > reports `missing_keys = {'pooler.dense.weight', 'pooler.dense.bias'}`, so those two tensors are
298
+ > **randomly initialised** and `pooler_output` is meaningless noise. Use `last_hidden_state` and pool
299
+ > it yourself. (`AutoModel` also reports the five `lm_head.*` tensors as unexpected — that is normal
300
+ > and harmless when loading an MLM checkpoint into a bare encoder.)
301
+
302
+ > [!WARNING]
303
+ > `last_hidden_state` gives *token* representations. Mean-pooling them is **not** a sentence-embedding
304
+ > model — no contrastive objective was ever applied here. For retrieval, fine-tune with a
305
+ > sentence-similarity objective (e.g. `sentence-transformers`) on top of this encoder.
306
+
307
+ ### Fine-tuning starting point
308
+
309
+ Swap in whichever head your task needs; the encoder weights load, the head is initialised fresh.
310
+
311
+ ```python
312
+ from transformers import (
313
+ AutoModelForSequenceClassification, # or AutoModelForTokenClassification
314
+ AutoTokenizer,
315
+ DataCollatorWithPadding,
316
+ Trainer,
317
+ TrainingArguments,
318
+ )
319
+
320
+ MODEL_ID = "djelia/bm-xlm-roberta-base"
321
+
322
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
323
+ model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID, num_labels=3)
324
+
325
+ def preprocess(batch):
326
+ # Normalize to NFC first — this tokenizer does no Unicode normalization of its own.
327
+ import unicodedata
328
+ texts = [unicodedata.normalize("NFC", t) for t in batch["text"]]
329
+ return tokenizer(texts, truncation=True, max_length=512)
330
+
331
+ tokenized = your_dataset.map(preprocess, batched=True)
332
+
333
+ trainer = Trainer(
334
+ model=model,
335
+ args=TrainingArguments(
336
+ output_dir="bm-xlmr-finetuned",
337
+ learning_rate=2e-5, # standard XLM-R base range: 1e-5 .. 3e-5
338
+ per_device_train_batch_size=16,
339
+ num_train_epochs=3,
340
+ warmup_ratio=0.06,
341
+ weight_decay=0.01,
342
+ fp16=True,
343
+ ),
344
+ train_dataset=tokenized["train"],
345
+ eval_dataset=tokenized["validation"],
346
+ data_collator=DataCollatorWithPadding(tokenizer),
347
+ )
348
+ trainer.train()
349
+ ```
350
+
351
+ Two practical notes:
352
+
353
+ - **Do not resize the embeddings** unless you are adding your own special tokens. The matrix is
354
+ already 253,421 rows and matches the shipped tokenizer.
355
+ - Because the added embeddings are less-trained than the inherited ones (see above), leaving the
356
+ embedding matrix **unfrozen** during fine-tuning is the safer default.
357
+
358
+ ## Training
359
+
360
+ > [!NOTE]
361
+ > This repo ships **no** training artifacts (no `trainer_state.json`, `training_args.bin`, or
362
+ > `all_results.json`). Everything in this section is read from
363
+ > [`oza75/xlm-roberta-bambara-01`](https://huggingface.co/oza75/xlm-roberta-bambara-01), whose
364
+ > `model.safetensors` is byte-identical to this one (SHA-256 verified, see *Lineage*). It is therefore
365
+ > the same run — but it is second-hand for *this* repo, and is labelled as such.
366
+
367
+ The original card in this repo states the model was trained "on an 11M words Bambara dataset"
368
+ (author's claim, preserved verbatim; the dataset itself is not identified).
369
+
370
+ **Reported evaluation results** (MLM objective, held-out split of the same unnamed corpus):
371
+
372
+ | Metric | Value |
373
+ |---|---|
374
+ | Eval loss | 1.1160 |
375
+ | Eval masked-token accuracy | 0.7599 |
376
+ | Eval perplexity | 3.0525 |
377
+ | Eval samples | 3,861 |
378
+ | Train samples | 34,645 |
379
+
380
+ **Hyperparameters** (from the sibling's Trainer-generated card):
381
+
382
+ | | |
383
+ |---|---|
384
+ | Learning rate | 2e-05 |
385
+ | Train / eval batch size | 32 / 32 |
386
+ | Gradient accumulation | 2 (effective batch 64) |
387
+ | Optimizer | AdamW (torch), β=(0.9, 0.999), ε=1e-08 |
388
+ | LR schedule | linear, `warmup_steps: 0.06` |
389
+ | Epochs | 100 |
390
+ | Seed | 42 |
391
+ | Train runtime | 54,567 s (~15.2 h) |
392
+ | Total FLOPs | 2.99e18 |
393
+ | Frameworks | Transformers 5.0.1.dev0, PyTorch 2.4.1+cu124, Datasets 4.5.0, Tokenizers 0.22.2 |
394
+
395
+ The loss curve reported there runs from 1.4658 / 0.6942 accuracy at step 1,000 to 1.1143 / 0.7592 at
396
+ step 52,000 — a long, flattening tail, with most of the gain landing in the first ~20k steps.
397
+
398
+ > [!WARNING]
399
+ > These numbers describe **MLM performance on the authors' own held-out split of an unidentified
400
+ > corpus**. They are not a benchmark, they are not comparable to any published Bambara result, and
401
+ > they say nothing about downstream task quality. **No downstream evaluation of this checkpoint
402
+ > exists** — not on classification, NER, POS, QA, or retrieval. If you need to know whether it beats
403
+ > stock XLM-R on your task, you have to measure it.
404
+
405
+ ## Limitations and risks
406
+
407
+ - **Not a task model.** See the note at the top.
408
+ - **Bambara only.** The additions target Bambara; the model retains stock XLM-R's 100-language
409
+ vocabulary, but the continued pretraining was Bambara-focused, and the general risk of catastrophic
410
+ forgetting on other languages applies. It has not been evaluated on any language other than
411
+ Bambara — including French, which co-occurs heavily with Bambara in Malian text.
412
+ - **Orthography sensitivity.** Bambara orthography varies widely in practice (`ɛ` vs `è`, `ɔ` vs `o`,
413
+ tone marking, `nyɔgɔn` vs `ɲɔgɔn`). The added vocabulary was learned from one corpus distribution;
414
+ text written in a different convention will tokenize less efficiently. Note the tokenizer holds
415
+ *both* `▁Nyɔgon` and `▁Ŋɔgɔn`, which suggests the training corpus itself was orthographically mixed.
416
+ - **No Unicode normalization** — see the tokenizer warning. Normalize to NFC yourself.
417
+ - **512-token ceiling**, inherited from XLM-R.
418
+ - **Untested for bias, toxicity, or factuality.** No such evaluation was run, by the authors or by me.
419
+ Inherits whatever is in CommonCrawl-scale XLM-R pretraining plus an unidentified Bambara corpus.
420
+ - **F32 only.** No fp16/bf16 or quantized variant is published; the 1.12 GB `model.safetensors` is
421
+ full-precision.
422
+ - **Mixed LayerNorm parameter naming in the checkpoint.** All 50 LayerNorms under `roberta.*` are
423
+ stored with the *legacy* `.gamma` / `.beta` names (e.g. `roberta.embeddings.LayerNorm.gamma`),
424
+ while `lm_head.layer_norm` uses the modern `.weight` / `.bias`. Stock
425
+ `FacebookAI/xlm-roberta-base` uses `.weight` / `.bias` throughout, so this was introduced by the
426
+ adaptation run, not inherited. It is transparent under `transformers` — 5.14.1 still ships the
427
+ `LayerNorm.gamma`→`weight` / `LayerNorm.beta`→`bias` rename in `conversion_mapping.py`, and
428
+ `AutoModelForMaskedLM.from_pretrained` loads all 202 tensors with **0 missing, 0 unexpected and 0
429
+ mismatched keys** — but that rename lives in a rule set explicitly labelled **`"legacy"`**.
430
+ **Tooling that reads `model.safetensors` directly** (custom loaders, ONNX/GGUF converters,
431
+ `safetensors` inspection scripts) **must expect the old names**, or it will silently leave every
432
+ encoder LayerNorm uninitialised.
433
+ - **No pooler weights** — see the embeddings warning above. `pooler_output` is randomly initialised.
434
+
435
+ ## Undetermined
436
+
437
+ Things I could not establish from the repo or any linked artifact. They are open questions, not
438
+ omissions to be filled in with plausible guesses.
439
+
440
+ - **Licence.** **No licence is declared anywhere** — not in the repo's card metadata, not in the Hub
441
+ tags, and not in any of the byte-identical sibling repos. Stock `FacebookAI/xlm-roberta-base` is
442
+ MIT, but a derivative's terms are the publisher's to state, and this one has not stated them.
443
+ **Treat the licensing status as unresolved and seek clarification from Djelia before reuse.**
444
+ - **Training data identity.** "An 11M words Bambara dataset" is the entirety of what is documented.
445
+ The corpus, its sources, its licensing, its collection date, and its train/eval split are all
446
+ unknown. The 34,645 train / 3,861 eval sample counts are known; what a "sample" is (sentence?
447
+ document? 512-token block?) is not.
448
+ - **How the 3,419 tokens were selected.** Whether by training a new SentencePiece/Unigram model on
449
+ Bambara and merging the delta, or by a frequency-threshold extraction — not documented. The
450
+ presence of a `##` token (a WordPiece convention) in an otherwise Unigram/Metaspace vocabulary, and
451
+ a stray `̸` (U+0338 combining long solidus overlay), both hint at an automated merge with light
452
+ filtering.
453
+ - **How the new embeddings were initialised** before training (random, mean-of-subword, or otherwise).
454
+ I can show they are trained *now*; I cannot recover the starting point.
455
+ - **The intermediate stages.** `oza75/xlm-roberta-bambara-cpt`, `-base`, `-initial`, `-base-01`, `-02`
456
+ form a chain, but which stage did the vocabulary extension and which did the continued pretraining
457
+ is not documented in any of them.
458
+ - **Hardware.** Not recorded. Only the 54,567 s runtime and 2.99e18 FLOPs are known.
459
+ - **Whether it improves any downstream task.** No downstream evaluation exists. The 30% tokenization
460
+ win is real and measured; a corresponding accuracy win is *plausible* but entirely unverified.
461
+ - **Intended production role.** The repo carries no description of what Djelia built it for.
462
+
463
+ ## Provenance of this card
464
+
465
+ Written by inspecting the repository directly rather than by restating its existing README. Every
466
+ number above traces to one of:
467
+
468
+ - the Hub API model record for `djelia/bm-xlm-roberta-base` (file list, blob sizes, LFS SHA-256s,
469
+ safetensors parameter count, commit history, timestamps);
470
+ - `config.json` and `tokenizer_config.json` from this repo;
471
+ - a direct comparison of `tokenizer.json` against `FacebookAI/xlm-roberta-base`;
472
+ - the `model.safetensors` header and selected tensor byte-ranges (embedding rows and `lm_head.bias`),
473
+ fetched via HTTP range requests and analysed with NumPy;
474
+ - a tokenization measurement I ran on a seeded 5,000-sentence sample of `djelia/bambara-texts`
475
+ (method stated in full above);
476
+ - and, for the training section only, the Trainer-generated metadata in the byte-identical
477
+ `oza75/xlm-roberta-bambara-01`, explicitly labelled as second-hand.
478
+
479
+ Every code snippet above was executed against this checkpoint on `transformers` 5.14.1 / `torch`
480
+ 2.13.0 before being included: the tokenizer loads cleanly (`mask_token='<mask>'`,
481
+ `mask_token_id=253420`, `len(tokenizer)=253421`), `AutoModelForMaskedLM` loads all 202 tensors with
482
+ zero missing/unexpected/mismatched keys, and the `fill-mask` outputs shown are the real ones. The
483
+ missing pooler was found by diffing the tensor key set against stock XLM-R and then confirmed at
484
+ runtime.
485
+
486
+ Retrieval date: 2026-07-27. Nothing here is copied from an upstream card without verification, and
487
+ no benchmark number appears that I did not either measure or trace to a named source file.