Simpler model card; source1.py comment wording
Browse files- EVALUATION.md +2 -2
- README.md +171 -148
- source1.py +1 -1
EVALUATION.md
CHANGED
|
@@ -179,7 +179,7 @@ rubric text. The exams were graded with an older revision of the rubric, from be
|
|
| 179 |
ads.
|
| 180 |
|
| 181 |
The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
|
| 182 |
-
same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-
|
| 183 |
false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
|
| 184 |
drops" in this file are those chunks: spam, boilerplate or toxic text under the hard filters.
|
| 185 |
|
|
@@ -792,7 +792,7 @@ on a single field; 4 labels and 1 keep decision changed).
|
|
| 792 |
document for text-and-data-mining and AI-training reservations. This screening did not catch notices worded in ways
|
| 793 |
the patterns miss, or reservations made outside the text itself (on the terms pages of sites the rules do not list,
|
| 794 |
or in machine-readable opt-out signals such as robots.txt). If you find such a document, tell us (see the contact
|
| 795 |
-
section of [README.md](README.md#contact
|
| 796 |
- **No reproduction kit.** See [Reproducing the evaluation](#reproducing-the-evaluation).
|
| 797 |
|
| 798 |
</details>
|
|
|
|
| 179 |
ads.
|
| 180 |
|
| 181 |
The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
|
| 182 |
+
same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-you-get)), and keep =
|
| 183 |
false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
|
| 184 |
drops" in this file are those chunks: spam, boilerplate or toxic text under the hard filters.
|
| 185 |
|
|
|
|
| 792 |
document for text-and-data-mining and AI-training reservations. This screening did not catch notices worded in ways
|
| 793 |
the patterns miss, or reservations made outside the text itself (on the terms pages of sites the rules do not list,
|
| 794 |
or in machine-readable opt-out signals such as robots.txt). If you find such a document, tell us (see the contact
|
| 795 |
+
section of [README.md](README.md#contact)).
|
| 796 |
- **No reproduction kit.** See [Reproducing the evaluation](#reproducing-the-evaluation).
|
| 797 |
|
| 798 |
</details>
|
README.md
CHANGED
|
@@ -6,6 +6,7 @@ base_model_relation: finetune
|
|
| 6 |
# "library_name: transformers" entry would offer AutoModel and pipeline snippets that load the backbone without
|
| 7 |
# the trained heads and return meaningless scores.
|
| 8 |
pipeline_tag: text-classification
|
|
|
|
| 9 |
language:
|
| 10 |
- en
|
| 11 |
- ar
|
|
@@ -136,33 +137,30 @@ datasets:
|
|
| 136 |
|
| 137 |
# Source-1
|
| 138 |
|
| 139 |
-
Source-1 scores
|
| 140 |
-
|
| 141 |
-
|
| 142 |
|
| 143 |

|
| 144 |
|
| 145 |
## Highlights
|
| 146 |
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
English chunks, 0.86 against 0.53 for the FineWeb-Edu classifier.
|
| 150 |
-
- **About 13x fewer parameters than propella-1 4B, and still closer to the grader.** Source-1 has 307M parameters,
|
| 151 |
-
propella-1 4B about 4.0B, and Source-1 agrees with the grader more closely on all three sets.
|
| 152 |
-
- **Within 0.012 of its 27B teacher at about 1/88 the size.** On the held-out set Source-1 scores 0.900; the
|
| 153 |
-
open-weight 27B LLM teacher it learned from scores 0.912.
|
| 154 |
-
- **Ahead on educational value alone, too.** On the English exam its `educational_value` reaches 0.90 rank agreement
|
| 155 |
-
with the grader's, against 0.61 for the FineWeb-Edu classifier and 0.88 for propella-1 4B.
|
| 156 |
|
| 157 |
-
*
|
| 158 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 159 |
|
| 160 |
## Quick start
|
| 161 |
|
| 162 |
-
Source-1 runs through the included `source1.py`
|
| 163 |
-
|
| 164 |
-
the backbone without the trained heads in `heads.safetensors` and return meaningless scores. The model is published on
|
| 165 |
-
the Hugging Face Hub as `msmth/Source-1`.
|
| 166 |
|
| 167 |
```bash
|
| 168 |
pip install -U huggingface_hub # provides the hf command
|
|
@@ -175,6 +173,20 @@ python source1.py --model . --input examples/sample.jsonl --output scores.jsonl
|
|
| 175 |
```python
|
| 176 |
from source1 import Source1
|
| 177 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 178 |
model = Source1.from_pretrained(".") # a local directory or a Hub repo id; bfloat16 weights by default
|
| 179 |
|
| 180 |
doc = model.score(open("article.txt", encoding="utf-8").read(), title="Optional title")
|
|
@@ -193,40 +205,82 @@ python source1.py --model . --input docs.jsonl --text-field text --output scores
|
|
| 193 |
python source1.py --model . --input page.txt --device cpu --precision fp32 --dtype fp32
|
| 194 |
```
|
| 195 |
|
| 196 |
-
- **Output.** One flat dict per document
|
| 197 |
-
length in tokens, whether a chunk was cut, and per-chunk results (`chunks`)
|
| 198 |
-
- **Precision.**
|
| 199 |
-
|
| 200 |
-
|
| 201 |
-
|
| 202 |
-
|
| 203 |
-
|
| 204 |
-
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
|
| 211 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 212 |
|---|---|---|
|
| 213 |
-
| `format` | label |
|
| 214 |
-
| `topic` | label |
|
| 215 |
-
| `content_type` | label |
|
| 216 |
-
| `educational_value` | quality, 0-5
|
| 217 |
-
| `reasoning_depth` | quality, 0-5 |
|
| 218 |
-
| `writing_quality` | quality, 0-5 |
|
| 219 |
-
| `information_density` | quality, 0-5 |
|
| 220 |
-
| `reliability` | quality, 0-5 |
|
| 221 |
-
| `spam_seo` | red flag, 0-5
|
| 222 |
-
| `boilerplate` | red flag, 0-5 |
|
| 223 |
-
| `toxicity` | red flag, 0-5 |
|
| 224 |
-
| `code_quality` |
|
| 225 |
-
| `math_quality` |
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 230 |
|
| 231 |
```
|
| 232 |
quality = 0.30*educational_value + 0.20*reasoning_depth + 0.15*writing_quality
|
|
@@ -239,34 +293,63 @@ overall = clip(quality - penalty, 0, 5)
|
|
| 239 |
drop if toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 (keep = false)
|
| 240 |
```
|
| 241 |
|
| 242 |
-
The drop line in `calibration.json` is
|
| 243 |
-
(`spam_seo >= 4`)
|
| 244 |
-
teacher's
|
| 245 |
-
[EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking
|
| 246 |
-
single fields, may
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 247 |
|
| 248 |
-
|
| 249 |
-
after a one-line header, as in training (source type, title if given, "Part i of n"). The document then takes the
|
| 250 |
-
label that covers the most tokens and the token-weighted mean of each score (the maximum for `toxicity`; gated scores
|
| 251 |
-
over the chunks where they apply), and `overall` and `keep` are recomputed from those.
|
| 252 |
|
| 253 |
-
|
|
|
|
|
|
|
|
|
|
| 254 |
|
| 255 |
-
|
| 256 |
-
-
|
| 257 |
-
-
|
| 258 |
-
-
|
|
|
|
| 259 |
|
| 260 |
-
|
| 261 |
|
| 262 |
-
-
|
| 263 |
-
|
| 264 |
-
|
| 265 |
-
-
|
| 266 |
-
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 267 |
|
| 268 |
## Training
|
| 269 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 270 |
- **Base model.** [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) (ModernBERT architecture, 8,192-token
|
| 271 |
context), all weights fine-tuned, with 13 linear heads on the mean-pooled hidden states: 307M parameters in total.
|
| 272 |
- **Data.** 172,895 chunks (350M tokens, 151,281 documents) in 53 languages, 38.9% English: filtered and unfiltered
|
|
@@ -277,88 +360,22 @@ Out of scope:
|
|
| 277 |
- **Recipe.** 2 epochs (5,320 steps of 131,072 tokens), AdamW, learning rate 5e-5 (heads 10x), bfloat16 mixed
|
| 278 |
precision, about 1.9 hours on one H100 NVL. The learning rate and language mix came from a short sweep read on
|
| 279 |
validation agreement with the teacher; the checkpoint is the final step, which also had the best validation score.
|
| 280 |
-
- **Development note.** AI assistants helped draft the rubric and write the code. Grades from proprietary LLM graders,
|
| 281 |
-
among them models of the evaluation grader's family, were used to choose between rubric revisions, to vet the
|
| 282 |
-
teacher and choose its prompt setup (on the samples that became the exam sets), and to set the candidate drop lines
|
| 283 |
-
and the 0.95 floor of the drop-line rule; LLM reviews also informed the license filtering. None of these grades or
|
| 284 |
-
reviews was used as a label or training example.
|
| 285 |
|
| 286 |
Data sources, filtering, recipe and development details:
|
| 287 |
[EVALUATION.md](EVALUATION.md#appendix-b-training-in-detail) and
|
| 288 |
[Independence from development](EVALUATION.md#independence-from-development).
|
| 289 |
|
| 290 |
-
|
| 291 |
-
|
| 292 |
-
Source-1 was compared with 16 public quality scorers on three test sets, each graded by an independent proprietary
|
| 293 |
-
LLM grader applying Source-1's rubric; the grader's labels were never trained on. Each cell is the rank agreement
|
| 294 |
-
(Spearman) between each model's main score and the grader's overall score, on the chunks that model scored.
|
| 295 |
|
| 296 |
-
|
| 297 |
-
|---|---|---|---|---|
|
| 298 |
-
| held-out set, 53 languages (main result) | 495 | **0.900** | 0.756 | 0.529 (159 English chunks; Source-1 0.864) |
|
| 299 |
-
| English exam | 412 to 413 | **0.921** | 0.820 | 0.453 |
|
| 300 |
-
| 12-language exam (web text) | 350 to 352 | **0.895** | 0.637 | English only |
|
| 301 |
-
|
| 302 |
-
The open-weight 27B LLM teacher scores 0.912, 0.912 and 0.880 on the same sets. All 39 public-scorer comparisons
|
| 303 |
-
favor Source-1 with 95% intervals clear of zero. The held-out set is the main result because, unlike the exams, it
|
| 304 |
-
played no part in choosing the teacher. Scoring each held-out chunk in full rather than its first 512 tokens raises
|
| 305 |
-
agreement from 0.85 to 0.90 ([details](EVALUATION.md#reading-the-whole-chunk)).
|
| 306 |
-
|
| 307 |
-
propella-1 has no single score: we read it through a composite of four of its quality ratings. Counting its own
|
| 308 |
-
commercial-bias, content-ratio, integrity and safety ratings as well narrows its held-out gap from 0.14 to 0.06-0.08,
|
| 309 |
-
depending on the weighting and languages, still in Source-1's favor
|
| 310 |
-
([details](EVALUATION.md#how-propella-1-is-read)). The public scorers also read the text without Source-1's one-line
|
| 311 |
-
header ([details](EVALUATION.md#the-one-line-header)).
|
| 312 |
-
|
| 313 |
-
Every scorer, interval, drop-line result and caveat is in [EVALUATION.md](EVALUATION.md).
|
| 314 |
|
| 315 |
-
|
|
|
|
|
|
|
| 316 |
|
| 317 |
-
|
| 318 |
-
|
| 319 |
-
of quality. Whether filtering with Source-1 trains better language models has not been tested.
|
| 320 |
-
- **The exams helped choose the teacher.** The two exam sets are the samples on which the teacher and its prompt setup
|
| 321 |
-
were chosen, against the exam grader's labels, so they are not independent of the grader.
|
| 322 |
-
- **It rates some qualities higher than the grader.** On the English exam's random sample, `educational_value` is 0.28
|
| 323 |
-
levels above the grader on average, and `reasoning_depth`, `writing_quality` and `reliability` 0.26 to 0.39 (the
|
| 324 |
-
teacher: 0.31, and 0.26 to 0.42).
|
| 325 |
-
- **It misses a third of the grader's drops at the shipped line.** It catches 43 of 64 (0.67) on the held-out set and
|
| 326 |
-
the teacher 46; 17 of Source-1's 21 misses are also missed by the teacher. A stricter `drop_line` catches more.
|
| 327 |
-
- **Weaker on some kinds of text.** Held-out rank is 0.68 for books and 0.73 for conversations, code and synthetic text,
|
| 328 |
-
against about 0.90 for web text; `code_quality` and `math_quality` are the least reliable fields.
|
| 329 |
-
- **Not a fact checker or a moderation tool.** `reliability` is a surface judgment; `toxicity` was trained on data where
|
| 330 |
-
toxic text is rare.
|
| 331 |
-
- **One chunk at a time, and less data for some languages.** Nothing outside a chunk of up to 8,192 tokens is visible to
|
| 332 |
-
it; the 16 languages with the fewest training chunks have 1,249 to 1,470 each.
|
| 333 |
-
- **License screening has limits.** Notices worded in ways the patterns miss, and opt-outs outside the text (such as
|
| 334 |
-
robots.txt), were not caught. If you find such a document, tell us (see below).
|
| 335 |
-
- **No reproduction kit.** The evaluation chunks, the grader's labels, per-chunk scores, the metrics script and the
|
| 336 |
-
teacher's prompt are not included, so the numbers cannot be recomputed from this repository.
|
| 337 |
-
|
| 338 |
-
More: [Limitations in detail](EVALUATION.md#limitations-in-detail).
|
| 339 |
-
|
| 340 |
-
## Files
|
| 341 |
|
| 342 |
-
| file | contents |
|
| 343 |
-
|---|---|
|
| 344 |
-
| `model.safetensors` | the fine-tuned mmBERT-base backbone in bfloat16 (default) |
|
| 345 |
-
| `model.fp32.safetensors` | the same backbone in float32 (load with `precision="fp32"`) |
|
| 346 |
-
| `config.json` | backbone configuration (ModernBERT) |
|
| 347 |
-
| `heads.safetensors` | the 13 scoring heads |
|
| 348 |
-
| `source1.json` | rubric, head layout, pooling, maximum length and text normalization |
|
| 349 |
-
| `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
|
| 350 |
-
| `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
|
| 351 |
-
| `source1.py` | standalone loader, Python API and command line |
|
| 352 |
-
| `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers` |
|
| 353 |
-
| `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
|
| 354 |
-
| `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
|
| 355 |
-
| `images/` | the benchmark chart above |
|
| 356 |
-
| `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
|
| 357 |
-
| `CREDITS_BOOKS.tsv` | per-work credits for the training books that are not public domain (part of `NOTICE`) |
|
| 358 |
-
|
| 359 |
-
## License and credits
|
| 360 |
-
|
| 361 |
-
Source-1 is released under the [Apache License 2.0](LICENSE). Copyright 2026 The Source-1 Authors (see `AUTHORS`).
|
| 362 |
[`NOTICE`](NOTICE) holds the full third-party notices and data credits. In short:
|
| 363 |
|
| 364 |
- **Base model.** Fine-tuned from [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) by the mmBERT authors at
|
|
@@ -377,14 +394,18 @@ Source-1 is released under the [Apache License 2.0](LICENSE). Copyright 2026 The
|
|
| 377 |
text it was trained on. Copyleft code was removed anyway.
|
| 378 |
- Upstream license metadata can be wrong. If you find a source that should not be here, please tell us (below).
|
| 379 |
|
| 380 |
-
|
| 381 |
|
| 382 |
-
|
| 383 |
-
|
| 384 |
-
the
|
|
|
|
| 385 |
|
| 386 |
## Citation
|
| 387 |
|
|
|
|
|
|
|
|
|
|
| 388 |
```bibtex
|
| 389 |
@misc{source1_2026,
|
| 390 |
title = {Source-1: a multilingual 13-field scorer for pretraining data},
|
|
@@ -407,3 +428,5 @@ Please also cite mmBERT:
|
|
| 407 |
url = {https://arxiv.org/abs/2509.06888}
|
| 408 |
}
|
| 409 |
```
|
|
|
|
|
|
|
|
|
| 6 |
# "library_name: transformers" entry would offer AutoModel and pipeline snippets that load the backbone without
|
| 7 |
# the trained heads and return meaningless scores.
|
| 8 |
pipeline_tag: text-classification
|
| 9 |
+
inference: false
|
| 10 |
language:
|
| 11 |
- en
|
| 12 |
- ar
|
|
|
|
| 137 |
|
| 138 |
# Source-1
|
| 139 |
|
| 140 |
+
Source-1 scores how useful a text is for training language models. It returns 13 scores and labels, such as
|
| 141 |
+
educational value, spam, toxicity and topic, plus one overall score and a keep-or-drop decision. It reads 53
|
| 142 |
+
languages, up to 8,192 tokens at a time. It has 307M parameters and is free to use under Apache-2.0.
|
| 143 |
|
| 144 |

|
| 145 |
|
| 146 |
## Highlights
|
| 147 |
|
| 148 |
+
The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same
|
| 149 |
+
order as the grader does (1.0 means the same order, 0 means no link).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 150 |
|
| 151 |
+
- **Beats all 16 public quality scorers we tested, including FineWeb-Edu and propella-1, on all three test sets.** On the main test set (495 held-out
|
| 152 |
+
texts in 53 languages): 0.90 vs 0.76 for the best of them, propella-1 4B, a model 13x larger.
|
| 153 |
+
- **Nearly matches its teacher**, the open-weight 27B LLM that labeled its training data, at 1/88 the size: 0.90 vs 0.91
|
| 154 |
+
on the main test set.
|
| 155 |
+
- **Also ahead when judged on educational value alone**, the thing most public scorers were built for.
|
| 156 |
+
|
| 157 |
+
*Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's
|
| 158 |
+
strengths. See [Limitations](#limitations) and [EVALUATION.md](EVALUATION.md).*
|
| 159 |
|
| 160 |
## Quick start
|
| 161 |
|
| 162 |
+
Source-1 runs through the included `source1.py`. Do not load it with transformers' `pipeline` or `AutoModel`: they
|
| 163 |
+
skip the trained scoring heads and give meaningless scores.
|
|
|
|
|
|
|
| 164 |
|
| 165 |
```bash
|
| 166 |
pip install -U huggingface_hub # provides the hf command
|
|
|
|
| 173 |
```python
|
| 174 |
from source1 import Source1
|
| 175 |
|
| 176 |
+
model = Source1.from_pretrained(".")
|
| 177 |
+
doc = model.score("Photosynthesis is how plants turn light, water and carbon dioxide into sugar and oxygen.")
|
| 178 |
+
print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision
|
| 179 |
+
```
|
| 180 |
+
|
| 181 |
+
<details>
|
| 182 |
+
<summary>More usage: many texts, precision, options, speed, files</summary>
|
| 183 |
+
|
| 184 |
+
`source1.py` needs only `torch`, `transformers`, `safetensors` and `tokenizers` (no `trust_remote_code`). The model
|
| 185 |
+
is published on the Hugging Face Hub as `msmth/Source-1`.
|
| 186 |
+
|
| 187 |
+
```python
|
| 188 |
+
from source1 import Source1
|
| 189 |
+
|
| 190 |
model = Source1.from_pretrained(".") # a local directory or a Hub repo id; bfloat16 weights by default
|
| 191 |
|
| 192 |
doc = model.score(open("article.txt", encoding="utf-8").read(), title="Optional title")
|
|
|
|
| 205 |
python source1.py --model . --input page.txt --device cpu --precision fp32 --dtype fp32
|
| 206 |
```
|
| 207 |
|
| 208 |
+
- **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the
|
| 209 |
+
length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`).
|
| 210 |
+
- **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB),
|
| 211 |
+
remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer
|
| 212 |
+
GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04
|
| 213 |
+
depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that.
|
| 214 |
+
- **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`.
|
| 215 |
+
`max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version. The command
|
| 216 |
+
line reads JSONL, JSON or plain text, and `python source1.py --help` lists every flag.
|
| 217 |
+
- **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few
|
| 218 |
+
hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16.
|
| 219 |
+
|
| 220 |
+
**Files**
|
| 221 |
+
|
| 222 |
+
| file | contents |
|
| 223 |
+
|---|---|
|
| 224 |
+
| `model.safetensors` | the fine-tuned mmBERT-base backbone in bfloat16 (default) |
|
| 225 |
+
| `model.fp32.safetensors` | the same backbone in float32 (load with `precision="fp32"`) |
|
| 226 |
+
| `config.json` | backbone configuration (ModernBERT) |
|
| 227 |
+
| `heads.safetensors` | the 13 scoring heads |
|
| 228 |
+
| `source1.json` | rubric, head layout, pooling, maximum length and text normalization |
|
| 229 |
+
| `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
|
| 230 |
+
| `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
|
| 231 |
+
| `source1.py` | standalone loader, Python API and command line |
|
| 232 |
+
| `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers` |
|
| 233 |
+
| `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
|
| 234 |
+
| `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
|
| 235 |
+
| `images/` | the benchmark chart above |
|
| 236 |
+
| `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
|
| 237 |
+
| `CREDITS_BOOKS.tsv` | per-work credits for the training books that are not public domain (part of `NOTICE`) |
|
| 238 |
+
|
| 239 |
+
</details>
|
| 240 |
+
|
| 241 |
+
## What you get
|
| 242 |
+
|
| 243 |
+
| field | kind | what it measures |
|
| 244 |
|---|---|---|
|
| 245 |
+
| `format` | label | kind of text (10 types) |
|
| 246 |
+
| `topic` | label | subject (15 topics) |
|
| 247 |
+
| `content_type` | label | plain text, code or math |
|
| 248 |
+
| `educational_value` | quality, 0-5 | teaches something useful |
|
| 249 |
+
| `reasoning_depth` | quality, 0-5 | explains why, step by step |
|
| 250 |
+
| `writing_quality` | quality, 0-5 | clear and well organized |
|
| 251 |
+
| `information_density` | quality, 0-5 | real content, not filler |
|
| 252 |
+
| `reliability` | quality, 0-5 | careful and trustworthy |
|
| 253 |
+
| `spam_seo` | red flag, 0-5 | ads, SEO spam and scams |
|
| 254 |
+
| `boilerplate` | red flag, 0-5 | menus, templates, link lists |
|
| 255 |
+
| `toxicity` | red flag, 0-5 | hate, harassment, explicit content |
|
| 256 |
+
| `code_quality` | code only, 0-5 | quality of the code |
|
| 257 |
+
| `math_quality` | math only, 0-5 | quality of the math |
|
| 258 |
+
|
| 259 |
+
Source-1 returns these 13 fields. Higher is better for quality scores and worse for red flags; the code and math
|
| 260 |
+
scores are `null` when they do not apply. You also get:
|
| 261 |
+
|
| 262 |
+
- `overall`: one 0-5 score, the quality scores minus penalties for red flags.
|
| 263 |
+
- `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`.
|
| 264 |
+
|
| 265 |
+
Long documents are split into chunks. Each chunk is scored, and the results are combined into one.
|
| 266 |
+
|
| 267 |
+
<details>
|
| 268 |
+
<summary>All fields in detail, the scoring formula, the drop line and long documents</summary>
|
| 269 |
+
|
| 270 |
+
**Label values.**
|
| 271 |
+
|
| 272 |
+
- `format`: tutorial, reference, news, forum_qa, academic, fiction, code_file, product_page, blog_opinion, other.
|
| 273 |
+
- `topic`: science, technology, programming, math, health, finance, history, politics_law, society,
|
| 274 |
+
philosophy_religion, arts_entertainment, literature, sports, lifestyle, other.
|
| 275 |
+
- `content_type`: plain_text, text_with_code, code_only, math_heavy.
|
| 276 |
+
|
| 277 |
+
**Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and
|
| 278 |
+
the score is the average level weighted by those chances. `code_quality` applies when `content_type` is
|
| 279 |
+
text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or
|
| 280 |
+
`topic` is math. Otherwise they are `null`. What each level means is in `source1.json` and
|
| 281 |
+
[EVALUATION.md Appendix A](EVALUATION.md#appendix-a-rubric-anchors), with the extra rules the teacher was given.
|
| 282 |
+
|
| 283 |
+
**Formula.**
|
| 284 |
|
| 285 |
```
|
| 286 |
quality = 0.30*educational_value + 0.20*reasoning_depth + 0.15*writing_quality
|
|
|
|
| 293 |
drop if toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 (keep = false)
|
| 294 |
```
|
| 295 |
|
| 296 |
+
**The drop line.** The line in `calibration.json` is a bit stricter on spam than the rubric's own rule
|
| 297 |
+
(`spam_seo >= 4`). Both keep ads and promotional pages, which the rubric scores `spam_seo` 3. The line was chosen by a
|
| 298 |
+
fixed rule on the teacher's labels. Stricter lines and what they cost are in
|
| 299 |
+
[EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules
|
| 300 |
+
on single fields, may work better for you.
|
| 301 |
+
|
| 302 |
+
**Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each
|
| 303 |
+
chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets
|
| 304 |
+
the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and
|
| 305 |
+
the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed.
|
| 306 |
+
|
| 307 |
+
</details>
|
| 308 |
|
| 309 |
+
## Good for / Not for
|
|
|
|
|
|
|
|
|
|
| 310 |
|
| 311 |
+
Good for:
|
| 312 |
+
- Filtering, ranking, weighting and mixing pretraining text in its 53 languages, using any of its fields.
|
| 313 |
+
- Checking a corpus: how much of it is spam, boilerplate, code, math or fiction.
|
| 314 |
+
- Research on data quality and on teaching small models to copy LLM judgments.
|
| 315 |
|
| 316 |
+
Not for:
|
| 317 |
+
- Judging people, job applications or student work.
|
| 318 |
+
- Fact-checking or content moderation: `reliability` judges care, not truth, and `toxicity` is a rough data filter.
|
| 319 |
+
- Deciding whether text is licensed or legal to use.
|
| 320 |
+
- Other languages, non-text input, or generating text.
|
| 321 |
|
| 322 |
+
## Limitations
|
| 323 |
|
| 324 |
+
- **Home ground.** The grader used Source-1's own rubric. Models from the grader's family also helped shape that
|
| 325 |
+
rubric, pick the teacher and tune the drop rule. The test texts come from the same kinds of sources as the training
|
| 326 |
+
data. Whether filtering with Source-1 trains better models is untested.
|
| 327 |
+
- **Two of the three test sets also helped pick the teacher**, by how well it agreed with the grader on them, so they
|
| 328 |
+
may flatter Source-1. The main test set was built after that and was not used to pick the model.
|
| 329 |
+
- **propella-1 has no single score.** We combine four of its ratings. If we also count its red-flag ratings,
|
| 330 |
+
Source-1's lead shrinks by about half but remains ([details](EVALUATION.md#how-propella-1-is-read)).
|
| 331 |
+
- **Scores run a bit high.** On English text it rates educational value, reasoning, writing and reliability about
|
| 332 |
+
0.3 points above the grader, as its teacher does.
|
| 333 |
+
- **Its keep/drop rule catches about 2 of every 3 texts the grader drops.** A stricter rule catches more but also
|
| 334 |
+
drops more good text.
|
| 335 |
+
- **Weaker outside web text** (books, chats, code, synthetic text). The code and math scores are the least reliable.
|
| 336 |
+
- **No reproduction kit.** The test data, the grader's labels and the metrics script are not included.
|
| 337 |
+
|
| 338 |
+
More in [EVALUATION.md](EVALUATION.md#limitations).
|
| 339 |
|
| 340 |
## Training
|
| 341 |
|
| 342 |
+
- Fine-tuned from [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base), with 13 small scoring heads added.
|
| 343 |
+
- About 173,000 text chunks in 53 languages, each labeled by an open-weight 27B LLM teacher (no human labels).
|
| 344 |
+
- License and safety filters were applied: see [NOTICE](NOTICE) and [EVALUATION.md](EVALUATION.md#appendix-b-training-in-detail).
|
| 345 |
+
- **Development note.** AI assistants helped write the rubric and the code. Grades from proprietary LLMs, including
|
| 346 |
+
the grader's model family, helped choose the rubric, the teacher and its setup, and the drop rule. LLM reviews also
|
| 347 |
+
helped shape the license filters. None of these grades or reviews was ever used as a label or training example
|
| 348 |
+
([full account](EVALUATION.md#independence-from-development)).
|
| 349 |
+
|
| 350 |
+
<details>
|
| 351 |
+
<summary>Training details</summary>
|
| 352 |
+
|
| 353 |
- **Base model.** [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) (ModernBERT architecture, 8,192-token
|
| 354 |
context), all weights fine-tuned, with 13 linear heads on the mean-pooled hidden states: 307M parameters in total.
|
| 355 |
- **Data.** 172,895 chunks (350M tokens, 151,281 documents) in 53 languages, 38.9% English: filtered and unfiltered
|
|
|
|
| 360 |
- **Recipe.** 2 epochs (5,320 steps of 131,072 tokens), AdamW, learning rate 5e-5 (heads 10x), bfloat16 mixed
|
| 361 |
precision, about 1.9 hours on one H100 NVL. The learning rate and language mix came from a short sweep read on
|
| 362 |
validation agreement with the teacher; the checkpoint is the final step, which also had the best validation score.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 363 |
|
| 364 |
Data sources, filtering, recipe and development details:
|
| 365 |
[EVALUATION.md](EVALUATION.md#appendix-b-training-in-detail) and
|
| 366 |
[Independence from development](EVALUATION.md#independence-from-development).
|
| 367 |
|
| 368 |
+
</details>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 369 |
|
| 370 |
+
## License and credits
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 371 |
|
| 372 |
+
- Source-1: [Apache License 2.0](LICENSE). Copyright 2026 The Source-1 Authors (see `AUTHORS`).
|
| 373 |
+
- Base model: [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) by the mmBERT authors, MIT License.
|
| 374 |
+
- Training data credits and third-party notices: [NOTICE](NOTICE) and `CREDITS_BOOKS.tsv`.
|
| 375 |
|
| 376 |
+
<details>
|
| 377 |
+
<summary>License details</summary>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 378 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 379 |
[`NOTICE`](NOTICE) holds the full third-party notices and data credits. In short:
|
| 380 |
|
| 381 |
- **Base model.** Fine-tuned from [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) by the mmBERT authors at
|
|
|
|
| 394 |
text it was trained on. Copyleft code was removed anyway.
|
| 395 |
- Upstream license metadata can be wrong. If you find a source that should not be here, please tell us (below).
|
| 396 |
|
| 397 |
+
</details>
|
| 398 |
|
| 399 |
+
## Contact
|
| 400 |
+
|
| 401 |
+
Questions, corrections and removal requests: the [Community tab](https://huggingface.co/msmth/Source-1/discussions).
|
| 402 |
+
Send the URL or dataset of your content and we will remove it from the next release's training data.
|
| 403 |
|
| 404 |
## Citation
|
| 405 |
|
| 406 |
+
<details>
|
| 407 |
+
<summary>BibTeX</summary>
|
| 408 |
+
|
| 409 |
```bibtex
|
| 410 |
@misc{source1_2026,
|
| 411 |
title = {Source-1: a multilingual 13-field scorer for pretraining data},
|
|
|
|
| 428 |
url = {https://arxiv.org/abs/2509.06888}
|
| 429 |
}
|
| 430 |
```
|
| 431 |
+
|
| 432 |
+
</details>
|
source1.py
CHANGED
|
@@ -51,7 +51,7 @@ How a document is scored (the same steps the model was trained and evaluated wit
|
|
| 51 |
Every training input had a Source line, almost always "dataset record", so that is the default; code files had
|
| 52 |
"<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
|
| 53 |
more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
|
| 54 |
-
one. On 1,261 held-out benchmark chunks (495 graded
|
| 55 |
dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
|
| 56 |
that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
|
| 57 |
graders of the model card's evaluation.
|
|
|
|
| 51 |
Every training input had a Source line, almost always "dataset record", so that is the default; code files had
|
| 52 |
"<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
|
| 53 |
more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
|
| 54 |
+
one. On 1,261 held-out benchmark chunks (495 graded held-out chunks from the test split from the test split and 766 exam chunks),
|
| 55 |
dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
|
| 56 |
that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
|
| 57 |
graders of the model card's evaluation.
|