---
title: README
emoji: ๐งญ
colorFrom: green
colorTo: indigo
sdk: static
pinned: false
---
Alignment โ and the assistant identity itself โ is normally introduced only *after*
pretraining, once behavioral priors are already set. **SPP** installs the desired persona from
token zero instead: we define it through normative values in a constitution, generate
first-person moral reflections grounded in that constitution, and insert them throughout the
pretraining corpus behind an `` token. Post-training then binds the chat assistant
identity to the installed persona. Pretraining up to **3B on 500B tokens**, SPP improves
constitution following and jailbreak robustness while preserving capabilities โ and *when* the
data arrives matters: models trained with reflections from token zero prioritize values
differently and take fewer risky actions in out-of-distribution moral dilemmas than models
given the exact same data only at the end of pretraining, an advantage that **grows with scale**.
๐ **Paper:** [Synthetic Persona Pretraining: Alignment from Token Zero](https://arxiv.org/abs/2608.13482)
## Collections
๐ฆ **[Pretraining Datasets](https://huggingface.co/collections/dlab-spp/pretraining-datasets-6a6b3b62c372179e342a1de6)**
โ the reflection data, the corpus selection manifest, safety scores, and verification files.
๐ค **[Models โ 3B](https://huggingface.co/collections/dlab-spp/models-3b-6a6b3cbc2bf4be9aba1c4a16)** ยท
**[Models โ 1.7B](https://huggingface.co/collections/dlab-spp/models-17b-6a6b3cbf48d091f505d412aa)**
โ all five recipes, at both scales. We release all pretraining checkpoints, base, and instruct models, at both scales
๐ฌ **[Post-training Dataset](https://huggingface.co/collections/dlab-spp/post-training-dataset-6a6c801cd9d46e4ff5d2ac25)**
โ SP-SFT, the mixture that performs persona binding.
๐ **[Evals](https://huggingface.co/collections/dlab-spp/evals-6a6c810619bb9009709286da)**
โ ConstitutionEval and an audited AIRiskDilemmas.
From [EPFL DLAB](https://huggingface.co/epfl-dlab).
## Citation
```bibtex
@misc{minder2026syntheticpersonapretrainingalignment,
title={Synthetic Persona Pretraining: Alignment from Token Zero},
author={Julian Minder and Viktor Moskvoretskii and Raghav Singhal and Difan Jiao and Andy Arditi and Shaobo Cui and Yiderigun Borjigin and Kartik Bali and Stefan Krsteski and Harsh Raj and Huu Nguyen and Jannik Brinkmann and Ashton Anderson and Roland Aydin and Robert West},
year={2026},
eprint={2608.13482},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.13482},
}
```