You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This model is a research checkpoint shared on request. Access is granted manually by the authors after reviewing your request. Please tell us who you are and what you want to use it for.

Log in or Sign Up to review the conditions and access this model content.

LightOnOCR-2-1B — fine-tuned for historical Barbados handwriting (line level)

Full fine-tune of lightonai/LightOnOCR-2-1B-base (LightOnOCR-2, ~1B parameters: Pixtral-style native-resolution vision encoder + Qwen3 decoder) for transcribing single handwritten text lines from historical Barbados records (17th–19th-century English legal and administrative documents). The released LightOnOCR-2 models do not target handwriting; this checkpoint does. Access is reviewed manually.

Local scores

Metric (competition metric): score = 1 − WER_w/24 − CER_w/110, where word and character edit distances are weighted per line by √(reference length). WER/CER columns are plain micro-averaged percentages. Greedy decoding at image scale 2.0 unless noted.

split lines score WER % CER % exact lines role
dev minus original100 ("dev150") 150 0.8915 16.03 5.44 34 selection set: checkpoint and image scale chosen here
dev250 (dev150 + original100) 250 0.8903 16.48 5.24 52
original100 100 0.8884 17.15 4.95 18 reported once, never used for any choice
audit350 350 0.8912 16.31 5.26 81 holdout, never used for any choice
dev150, zero-shot base model 150 0.5972 49.88 27.65 0 scale 2.0, same usage

Decoding comparison (same checkpoint; beam = 5 beams; "+ char LM" = beam candidates re-ranked with a character 6-gram trained on the training transcriptions, weights λ=0.2, μ=0.5 chosen on dev150 — the LM is not included in this repo):

split greedy beam-5 top-1 beam-5 re-ranked by model log-prob beam-5 + char LM best candidate in the 5-best (oracle)
dev-minus-o100 (tuning set) 0.8915 0.8933 0.8946 0.9020 0.9198
audit350 (holdout) 0.8912 0.8949 0.8927 0.8987 0.9240
original100 (once) 0.8884 0.8905 0.8891 0.8929 0.9245

Paired bootstrap on audit350: beam-5 vs greedy +0.0014 (95% CI -0.0018 to +0.0045); beam-5 + char LM vs greedy +0.0075 (95% CI +0.0031 to +0.0123).

Data splits (frozen, SHA-256 verified images)

split lines use
train3497 3,497 training only
dev150 = dev250 minus original100 150 every choice: checkpoint, image scale, decoding settings
original100 100 reported once at the end; never used for any choice
audit350 350 holdout for comparisons; never tuned on
test 1,374 competition test lines (no labels)

The splits are disjoint by line ID and by exact image hash.

Requirements

  • transformers==5.2.0 (LightOnOCR is native in transformers 5.x: LightOnOcrForConditionalGeneration, LightOnOcrProcessor)
  • torch>=2.4 with torchvision, Pillow, huggingface_hub
  • a CUDA GPU with bf16 (≈ 5 GB for inference; the weights are FP32, 4 GB); CPU works but is slow
pip install "transformers==5.2.0" torch torchvision pillow huggingface_hub
huggingface-cli login        # after your access request is approved

How to use

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Abdoul27/lightonocr-2-1b-barbados")
sys.path.insert(0, path)
from barbados_inference import load, transcribe
model, processor = load(path)
print(transcribe(model, processor, ["line_001.jpg", "line_002.jpg"]))            # greedy (scores above)
print(transcribe(model, processor, ["line_001.jpg"], num_beams=5))               # beam-5 (see decoding comparison)

What barbados_inference.py does (reproduce this exactly to get the scores above):

  1. Input: one image per handwritten line (not a full page), converted to RGB and resized by 2.0× (bicubic). The processor then fits the longest edge into 1,540 px (the model's trained maximum) and normalises it.
  2. Prompt: the image only, with no text, through the model's own chat template (official LightOnOCR usage).
  3. Decoding: greedy (or beam search), at most 51 new tokens, stop at <|im_end|> or <|endoftext|>; FP32 weights with BF16 autocast on GPU; one image per call.
  4. Output: multi-line outputs are joined with single spaces; repetition loops are trimmed only when the length cap is hit.

Transcription conventions learned from the training labels: original spelling and abbreviations are kept (pnts, Xpian, w^th); ^ marks raised (superscript) letters, e.g. w^th, Exec:^rs; & is kept.

Training

  • Full fine-tune of all weights of lightonai/LightOnOCR-2-1B-base (the pre-RL checkpoint the authors recommend for fine-tuning).
  • FP32 weights with BF16 autocast, AdamW (no weight decay), lr 2e-05 (vision encoder 4e-06), 5 % warm-up, cosine to 10 %, effective batch 16, 6 epochs, evaluation twice per epoch; best checkpoint epoch_06_end, selected on dev150.
  • Loss on the answer tokens only (incl. <|im_end|>); mild augmentation (±10 % scale, ±1.5° rotation, brightness/contrast, light blur); gradient checkpointing. barbados_training_worker.py is the complete training/evaluation code.

Limitations

  • Line-level model: it expects one cropped text line per image; full pages need line segmentation first.
  • Trained on 3497 lines from one archive and period; other hands, languages or scripts may degrade.

Licence

Apache-2.0, as the base model lightonai/LightOnOCR-2-1B-base by LightOn.

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Abdoul27/lightonocr-2-1b-barbados

Finetuned
(25)
this model