Long-Horizon Historical BERT
Long-Horizon Historical BERT is a French historical language model based on dbmdz/bert-base-french-europeana-cased, adapted with temporal adapters trained on French public-domain historical newspapers.
The model is designed for experiments on historical NLP, especially where language use changes over time.
Model description
This model uses:
- base encoder:
dbmdz/bert-base-french-europeana-cased - task: masked language modeling
- adaptation method: temporal adapter bank
- training data:
PleIAs/French-PD-Newspapers - temporal periods:
PERIODS = [
"pre_1850",
"1850_1899",
"1900_1938",
"1939_1945",
"post_1945",
]
Each input document is assigned to a temporal period based on its publication year. The corresponding adapter is applied after the BERT encoder.
How to load the model
This repository contains a custom PyTorch model. The file modeling_temporal_bert.py must be available locally.
import torch
from transformers import AutoTokenizer
from modeling_temporal_bert import HistoricalTemporalBertForMLM, year_to_period_id
model_id = "EmanuelaBoros/long-horizon-historical-bert"
base_model = "dbmdz/bert-base-french-europeana-cased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = HistoricalTemporalBertForMLM(
base_model_name=base_model,
adapter_bottleneck_size=64,
adapter_dropout=0.1,
)
state = torch.load("pytorch_model.bin", map_location="cpu")
model.load_state_dict(state, strict=False)
model.eval()
text = "Munich est mentionné dans le journal."
year = 1888
inputs = tokenizer(text, return_tensors="pt")
inputs["period_ids"] = torch.tensor([year_to_period_id(year)])
with torch.no_grad():
outputs = model(**inputs)
Downstream evaluation: Named Entity Recognition
The model was evaluated on French HIPE-style historical newspaper NER data.
Dataset format
The NER data follows the HIPE TSV format:
TOKEN NE-COARSE-LIT NE-COARSE-METO NE-FINE-LIT NE-FINE-METO NE-FINE-COMP NE-NESTED NEL-LIT NEL-METO MISC
For this evaluation, the label column was:
NE-COARSE-LIT
The data was split into sentences using MISC = EndOfSentence.
Data statistics
| Split | Sentences |
|---|---|
| Train | 5,743 |
| Dev | 1,244 |
| Test | 1,462 |
Labels:
{
"O": 0,
"B-loc": 1,
"B-org": 2,
"B-pers": 3,
"B-prod": 4,
"B-time": 5,
"I-loc": 6,
"I-org": 7,
"I-pers": 8,
"I-prod": 9,
"I-time": 10
}
NER results
Three models were compared:
dbmdz/bert-base-french-europeana-caseddbmdz/bert-base-french-europeana-casedwith the year concatenated as a special prefix tokenEmanuelaBoros/long-horizon-historical-bert
All models were fine-tuned for token classification on the same HIPE French NER data.
Overall test results
| Model | Precision | Recall | F1 | Test loss |
|---|---|---|---|---|
dbmdz/bert-base-french-europeana-cased |
0.7657 | 0.8309 | 0.7970 | 0.0998 |
dbmdz/bert-base-french-europeana-cased + year prefix |
0.7544 | 0.8256 | 0.7884 | 0.1009 |
EmanuelaBoros/long-horizon-historical-bert |
0.7621 | 0.8210 | 0.7905 | 0.0991 |
Development results
| Model | Precision | Recall | F1 | Dev loss |
|---|---|---|---|---|
EmanuelaBoros/long-horizon-historical-bert |
0.8200 | 0.8570 | 0.8381 | 0.0748 |
Test results by temporal period
| Model | Period | Sentences | Entities | Precision | Recall | F1 |
|---|---|---|---|---|---|---|
dbmdz/bert-base-french-europeana-cased |
pre_1850 | 472 | 404 | 0.6713 | 0.7902 | 0.7259 |
dbmdz/bert-base-french-europeana-cased |
1850_1899 | 297 | 328 | 0.8022 | 0.8649 | 0.8324 |
dbmdz/bert-base-french-europeana-cased |
1900_1938 | 432 | 394 | 0.8456 | 0.8922 | 0.8683 |
dbmdz/bert-base-french-europeana-cased |
post_1945 | 261 | 346 | 0.7682 | 0.7790 | 0.7736 |
dbmdz/bert-base-french-europeana-cased + year prefix |
pre_1850 | 472 | 404 | 0.6517 | 0.7762 | 0.7085 |
dbmdz/bert-base-french-europeana-cased + year prefix |
1850_1899 | 297 | 328 | 0.7910 | 0.8408 | 0.8151 |
dbmdz/bert-base-french-europeana-cased + year prefix |
1900_1938 | 432 | 394 | 0.8551 | 0.9023 | 0.8780 |
dbmdz/bert-base-french-europeana-cased + year prefix |
post_1945 | 261 | 346 | 0.7466 | 0.7847 | 0.7652 |
EmanuelaBoros/long-horizon-historical-bert |
pre_1850 | 472 | 404 | 0.6673 | 0.7809 | 0.7197 |
EmanuelaBoros/long-horizon-historical-bert |
1850_1899 | 297 | 328 | 0.8145 | 0.8438 | 0.8289 |
EmanuelaBoros/long-horizon-historical-bert |
1900_1938 | 432 | 394 | 0.8504 | 0.8972 | 0.8732 |
EmanuelaBoros/long-horizon-historical-bert |
post_1945 | 261 | 346 | 0.7410 | 0.7620 | 0.7514 |
Per-period F1 comparison
| Period | Baseline BERT F1 | Year-prefix BERT F1 | Temporal BERT F1 |
|---|---|---|---|
| pre_1850 | 0.7259 | 0.7085 | 0.7197 |
| 1850_1899 | 0.8324 | 0.8151 | 0.8289 |
| 1900_1938 | 0.8683 | 0.8780 | 0.8732 |
| post_1945 | 0.7736 | 0.7652 | 0.7514 |
Per-class test results: baseline BERT
| Entity type | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| loc | 0.85 | 0.90 | 0.87 | 759 |
| org | 0.55 | 0.57 | 0.56 | 122 |
| pers | 0.73 | 0.83 | 0.78 | 527 |
| prod | 0.70 | 0.75 | 0.73 | 53 |
| time | 0.49 | 0.60 | 0.54 | 53 |
| micro avg | 0.77 | 0.83 | 0.80 | 1514 |
| macro avg | 0.67 | 0.73 | 0.70 | 1514 |
| weighted avg | 0.77 | 0.83 | 0.80 | 1514 |
Per-class test results: year-prefix BERT
| Entity type | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| loc | 0.84 | 0.89 | 0.86 | 759 |
| org | 0.52 | 0.57 | 0.54 | 122 |
| pers | 0.72 | 0.82 | 0.77 | 527 |
| prod | 0.73 | 0.72 | 0.72 | 53 |
| time | 0.56 | 0.72 | 0.63 | 53 |
| micro avg | 0.75 | 0.83 | 0.79 | 1514 |
| macro avg | 0.67 | 0.74 | 0.70 | 1514 |
| weighted avg | 0.76 | 0.83 | 0.79 | 1514 |
Per-class test results: long-horizon temporal BERT
| Entity type | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| loc | 0.86 | 0.88 | 0.87 | 759 |
| org | 0.50 | 0.52 | 0.51 | 122 |
| pers | 0.71 | 0.82 | 0.76 | 527 |
| prod | 0.69 | 0.75 | 0.72 | 53 |
| time | 0.61 | 0.70 | 0.65 | 53 |
| micro avg | 0.76 | 0.82 | 0.79 | 1514 |
| macro avg | 0.67 | 0.73 | 0.70 | 1514 |
| weighted avg | 0.76 | 0.82 | 0.79 | 1514 |
Summary
On this HIPE French NER evaluation, the base Europeana BERT model obtains the best overall test F1:
0.7970 baseline vs 0.7884 year-prefix vs 0.7905 temporal
The year-prefix baseline does not improve overall performance, despite explicitly giving the publication year as input. The temporal model performs slightly better than the year-prefix baseline, but remains slightly below the base BERT model in this single run.
The per-period results show that temporal information has mixed effects. The year-prefix model performs best on 1900_1938, while the base BERT model is strongest on pre_1850, 1850_1899, and post_1945. The temporal adapter model is competitive across periods but does not clearly outperform the baseline in this run.
These results suggest that temporal adaptation preserves downstream NER performance, but more evaluation is needed to establish whether it improves robustness across periods. Future work should include multiple random seeds, additional HIPE label settings, and more historical NER datasets.
Citation
If you use this model, please cite the model repository and the underlying base model:
@misc{boros_long_horizon_historical_bert,
title = {Long-Horizon Historical BERT},
author = {Boros, Emanuela},
year = {2026},
howpublished = {https://huggingface.co/EmanuelaBoros/long-horizon-historical-bert}
}
- Downloads last month
- 13
Model tree for emanuelaboros/long-horizon-historical-bert
Base model
dbmdz/bert-base-french-europeana-cased