Long-Horizon Historical BERT

Long-Horizon Historical BERT is a French historical language model based on dbmdz/bert-base-french-europeana-cased, adapted with temporal adapters trained on French public-domain historical newspapers.

The model is designed for experiments on historical NLP, especially where language use changes over time.

Model description

This model uses:

  • base encoder: dbmdz/bert-base-french-europeana-cased
  • task: masked language modeling
  • adaptation method: temporal adapter bank
  • training data: PleIAs/French-PD-Newspapers
  • temporal periods:
PERIODS = [
    "pre_1850",
    "1850_1899",
    "1900_1938",
    "1939_1945",
    "post_1945",
]

Each input document is assigned to a temporal period based on its publication year. The corresponding adapter is applied after the BERT encoder.

How to load the model

This repository contains a custom PyTorch model. The file modeling_temporal_bert.py must be available locally.

import torch
from transformers import AutoTokenizer
from modeling_temporal_bert import HistoricalTemporalBertForMLM, year_to_period_id

model_id = "EmanuelaBoros/long-horizon-historical-bert"
base_model = "dbmdz/bert-base-french-europeana-cased"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = HistoricalTemporalBertForMLM(
    base_model_name=base_model,
    adapter_bottleneck_size=64,
    adapter_dropout=0.1,
)

state = torch.load("pytorch_model.bin", map_location="cpu")
model.load_state_dict(state, strict=False)
model.eval()

text = "Munich est mentionné dans le journal."
year = 1888

inputs = tokenizer(text, return_tensors="pt")
inputs["period_ids"] = torch.tensor([year_to_period_id(year)])

with torch.no_grad():
    outputs = model(**inputs)

Downstream evaluation: Named Entity Recognition

The model was evaluated on French HIPE-style historical newspaper NER data.

Dataset format

The NER data follows the HIPE TSV format:

TOKEN   NE-COARSE-LIT   NE-COARSE-METO   NE-FINE-LIT   NE-FINE-METO   NE-FINE-COMP   NE-NESTED   NEL-LIT   NEL-METO   MISC

For this evaluation, the label column was:

NE-COARSE-LIT

The data was split into sentences using MISC = EndOfSentence.

Data statistics

Split Sentences
Train 5,743
Dev 1,244
Test 1,462

Labels:

{
  "O": 0,
  "B-loc": 1,
  "B-org": 2,
  "B-pers": 3,
  "B-prod": 4,
  "B-time": 5,
  "I-loc": 6,
  "I-org": 7,
  "I-pers": 8,
  "I-prod": 9,
  "I-time": 10
}

NER results

Three models were compared:

  1. dbmdz/bert-base-french-europeana-cased
  2. dbmdz/bert-base-french-europeana-cased with the year concatenated as a special prefix token
  3. EmanuelaBoros/long-horizon-historical-bert

All models were fine-tuned for token classification on the same HIPE French NER data.

Overall test results

Model Precision Recall F1 Test loss
dbmdz/bert-base-french-europeana-cased 0.7657 0.8309 0.7970 0.0998
dbmdz/bert-base-french-europeana-cased + year prefix 0.7544 0.8256 0.7884 0.1009
EmanuelaBoros/long-horizon-historical-bert 0.7621 0.8210 0.7905 0.0991

Development results

Model Precision Recall F1 Dev loss
EmanuelaBoros/long-horizon-historical-bert 0.8200 0.8570 0.8381 0.0748

Test results by temporal period

Model Period Sentences Entities Precision Recall F1
dbmdz/bert-base-french-europeana-cased pre_1850 472 404 0.6713 0.7902 0.7259
dbmdz/bert-base-french-europeana-cased 1850_1899 297 328 0.8022 0.8649 0.8324
dbmdz/bert-base-french-europeana-cased 1900_1938 432 394 0.8456 0.8922 0.8683
dbmdz/bert-base-french-europeana-cased post_1945 261 346 0.7682 0.7790 0.7736
dbmdz/bert-base-french-europeana-cased + year prefix pre_1850 472 404 0.6517 0.7762 0.7085
dbmdz/bert-base-french-europeana-cased + year prefix 1850_1899 297 328 0.7910 0.8408 0.8151
dbmdz/bert-base-french-europeana-cased + year prefix 1900_1938 432 394 0.8551 0.9023 0.8780
dbmdz/bert-base-french-europeana-cased + year prefix post_1945 261 346 0.7466 0.7847 0.7652
EmanuelaBoros/long-horizon-historical-bert pre_1850 472 404 0.6673 0.7809 0.7197
EmanuelaBoros/long-horizon-historical-bert 1850_1899 297 328 0.8145 0.8438 0.8289
EmanuelaBoros/long-horizon-historical-bert 1900_1938 432 394 0.8504 0.8972 0.8732
EmanuelaBoros/long-horizon-historical-bert post_1945 261 346 0.7410 0.7620 0.7514

Per-period F1 comparison

Period Baseline BERT F1 Year-prefix BERT F1 Temporal BERT F1
pre_1850 0.7259 0.7085 0.7197
1850_1899 0.8324 0.8151 0.8289
1900_1938 0.8683 0.8780 0.8732
post_1945 0.7736 0.7652 0.7514

Per-class test results: baseline BERT

Entity type Precision Recall F1 Support
loc 0.85 0.90 0.87 759
org 0.55 0.57 0.56 122
pers 0.73 0.83 0.78 527
prod 0.70 0.75 0.73 53
time 0.49 0.60 0.54 53
micro avg 0.77 0.83 0.80 1514
macro avg 0.67 0.73 0.70 1514
weighted avg 0.77 0.83 0.80 1514

Per-class test results: year-prefix BERT

Entity type Precision Recall F1 Support
loc 0.84 0.89 0.86 759
org 0.52 0.57 0.54 122
pers 0.72 0.82 0.77 527
prod 0.73 0.72 0.72 53
time 0.56 0.72 0.63 53
micro avg 0.75 0.83 0.79 1514
macro avg 0.67 0.74 0.70 1514
weighted avg 0.76 0.83 0.79 1514

Per-class test results: long-horizon temporal BERT

Entity type Precision Recall F1 Support
loc 0.86 0.88 0.87 759
org 0.50 0.52 0.51 122
pers 0.71 0.82 0.76 527
prod 0.69 0.75 0.72 53
time 0.61 0.70 0.65 53
micro avg 0.76 0.82 0.79 1514
macro avg 0.67 0.73 0.70 1514
weighted avg 0.76 0.82 0.79 1514

Summary

On this HIPE French NER evaluation, the base Europeana BERT model obtains the best overall test F1:

0.7970 baseline vs 0.7884 year-prefix vs 0.7905 temporal

The year-prefix baseline does not improve overall performance, despite explicitly giving the publication year as input. The temporal model performs slightly better than the year-prefix baseline, but remains slightly below the base BERT model in this single run.

The per-period results show that temporal information has mixed effects. The year-prefix model performs best on 1900_1938, while the base BERT model is strongest on pre_1850, 1850_1899, and post_1945. The temporal adapter model is competitive across periods but does not clearly outperform the baseline in this run.

These results suggest that temporal adaptation preserves downstream NER performance, but more evaluation is needed to establish whether it improves robustness across periods. Future work should include multiple random seeds, additional HIPE label settings, and more historical NER datasets.

Citation

If you use this model, please cite the model repository and the underlying base model:

@misc{boros_long_horizon_historical_bert,
  title = {Long-Horizon Historical BERT},
  author = {Boros, Emanuela},
  year = {2026},
  howpublished = {https://huggingface.co/EmanuelaBoros/long-horizon-historical-bert}
}
Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emanuelaboros/long-horizon-historical-bert

Finetuned
(2)
this model

Dataset used to train emanuelaboros/long-horizon-historical-bert

Collection including emanuelaboros/long-horizon-historical-bert