--- title: README emoji: ๐Ÿงญ colorFrom: green colorTo: indigo sdk: static pinned: false --- # Synthetic Persona Pretraining (SPP) **Alignment from token zero.** Research artifacts from [EPFL DLAB](https://huggingface.co/epfl-dlab). Alignment โ€” and the assistant identity itself โ€” is normally introduced only *after* pretraining, once behavioral priors are already set. That makes aligned behavior easy to circumvent, and it misses pretraining's broad contextual coverage. **SPP** installs the desired assistant persona from the first token. We define the persona through normative values in a **constitution**, generate first-person moral reflections grounded in it, and insert them throughout the pretraining corpus behind an `` token. Post-training then binds the chat assistant identity to that installed persona โ€” a step we call **persona binding**. Pretraining models up to **3B on 500B tokens**, we find that SPP improves constitution following and jailbreak robustness while preserving capabilities. *When* the data arrives matters: models trained with reflections from token zero prioritize values differently and choose less risky actions in out-of-distribution moral dilemmas than models given the exact same reflection data only at the end of pretraining โ€” and that advantage **grows with scale**. --- ## The five recipes Every model below is one of five data-matched recipes, released at two scales (**1.7B**/100B tokens and **3B**/500B tokens), as both `base` and SP-SFT-post-trained `instruct`: | Recipe | Intervention | | --- | --- | | **Vanilla** | Standard next-token pretraining. The control. | | **Filtered** | Harmful documents (safety โ‰ฅ 3) loss-masked but retained, so batches match Vanilla exactly. | | **SPP(T0)** | Reflections distributed uniformly from the **first batch** onward. | | **SPP(MT)** | The same reflections, introduced **only at midtraining**. Token-matched to SPP(T0) โ€” isolates *when* the data is seen. | | **SPP(T0,MT)** | Both: SPP(T0) pretraining plus the reflection-focused midtraining stage. | The SPP(T0) vs SPP(MT) contrast is the core comparison: identical data, different timing. ## Artifacts ๐Ÿ“ฆ **[Pretraining Datasets](https://huggingface.co/collections/dlab-spp/pretraining-datasets-6a6b3b62c372179e342a1de6)** โ€” `reflection-50m` is the main artifact: 51.4M documents with the reflections the released models were trained on, and the exact token offset each was inserted at. Alongside it, a corpus manifest, SafeLM safety scores for 391M documents, and verification data. ๐Ÿค– **[Models โ€” 3B](https://huggingface.co/collections/dlab-spp/models-3b-6a6b3cbc2bf4be9aba1c4a16)** ยท **[Models โ€” 1.7B](https://huggingface.co/collections/dlab-spp/models-17b-6a6b3cbf48d091f505d412aa)** โ€” all five recipes, base and instruct, at both scales. ๐Ÿ’ฌ **[Post-training Dataset](https://huggingface.co/collections/dlab-spp/post-training-dataset-6a6c801cd9d46e4ff5d2ac25)** โ€” **SP-SFT**, the mixture that performs persona binding. Each prompt is paired with a constitution-aware response, a citation-free rendering of it, and the original source answer. ๐Ÿ“Š **[Evals](https://huggingface.co/collections/dlab-spp/evals-6a6c810619bb9009709286da)** โ€” **ConstitutionEval** for in-domain constitution following, and an **audited AIRiskDilemmas** for out-of-distribution value prioritisation and risk-taking. ## Reproducibility The pretraining corpus is a seeded subsample of [Dolma 3](https://huggingface.co/datasets/allenai/dolma3_mix-6T). Rather than redistribute ~2.6 TB of already-public text, we publish the **selection decisions** keyed by upstream document id, plus the safety-classifier outputs โ€” the one stage of the pipeline that is not bit-reproducible. Together with the reflections and the `.idx` verification data, the corpus can be rebuilt from public data and checked byte for byte. Corpus-derived datasets are released under **ODC-BY 1.0**, inherited from Dolma 3. > Contains information from [`allenai/dolma3_mix-6T`](https://huggingface.co/datasets/allenai/dolma3_mix-6T), > made available under the Open Data Commons Attribution License (ODC-BY 1.0). ## โš ๏ธ Content warning The reflection datasets contain real web documents, roughly half of which are flagged harmful, together with model-generated commentary on them. This is research data on alignment training โ€” treat it accordingly.