Multicentury TrOCR — Barbados handwriting, 3-member ensemble (line level)
Fine-tunes of Kansallisarkisto/multicentury-htr-model on single handwritten lines of 17th-century Barbados records.
Three members trained with different augmentation mixes (elastic / rotation + underline / mixed), each with weight EMA, early stopping and its
own decoding chosen on dev150; the submitted system is their line-level vote (majority, else medoid).
| member | checkpoint | weights | decoding | dev150 |
|---|---|---|---|---|
| elastic | epoch_15_ema_publisher | ema | publisher | 0.9017 |
| rotate_underline | epoch_13_ema_beam8 | ema | beam8 | 0.9007 |
| mixed | epoch_14_ema_publisher | ema | publisher | 0.9003 |
| vote of 3 | 0.9042 |
original100 (scored once, never used for any choice): 0.8864. Metric: 1 − WER_w/24 − CER_w/110 (√length-weighted). audit350 is training data (train3497 + audit350 + pseudo-labelled test lines), so dev150 / original100 are the fair local comparisons.
Use
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Abdoul27/multicentury-trocr-barbados-v3", token="hf_...") # gated: your access is approved manually
sys.path.insert(0, path)
from barbados_inference import load_members, transcribe_ensemble
members = load_members(path)
final, per_member = transcribe_ensemble(members, ["line_001.jpg", "line_002.jpg"], return_members=True)
Requirements: transformers==4.49.0, torch (CUDA), Pillow, rapidfuzz, safetensors. Preprocessing: TrOCR slow processor at 192×1024 with the
publisher's rectangular-input ViT patch (interpolation mode legacy); decodings: publisher = the base model's own generation config
(beam 4, length penalty 2.0, 3-gram blocking), beamN / barbados = N beams, no n-gram blocking, length penalty 1.0, cap 51 tokens.
barbados_training_worker.py is the complete training / evaluation code.
Model tree for Abdoul27/multicentury-trocr-barbados-v3
Base model
microsoft/trocr-large-handwritten