You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This model is a research checkpoint shared on request. Access is granted manually by the authors after reviewing your request. Please tell us who you are and what you want to use it for. Use must respect the base model licence (modified OpenRAIL-M).

Log in or Sign Up to review the conditions and access this model content.

Chandra OCR 2 — fine-tuned for historical Barbados handwriting (line level)

Full fine-tune of datalab-to/chandra-ocr-2 (Qwen3.5 architecture, ~5B parameters: hybrid Gated-DeltaNet / full-attention decoder + vision encoder) for transcribing single handwritten text lines from historical Barbados records (17th–19th-century English legal and administrative documents). Access is reviewed manually.

Local scores

Metric (competition metric): score = 1 − WER_w/24 − CER_w/110, where word and character edit distances are weighted per line by √(reference length). WER/CER columns are plain micro-averaged percentages. Greedy decoding at image scale 1.0.

split lines score WER % CER % exact lines role
dev minus original100 ("dev150") 150 0.8893 16.38 5.43 31 selection set: checkpoint and image scale chosen here
dev250 (dev150 + original100) 250 0.8883 16.66 5.24 47
original100 100 0.8867 17.06 4.95 16 reported once, never used for any choice
dev150, zero-shot base model 150 0.6849 40.74 20.51 2 scale 1.0, same prompt

Data splits (frozen, SHA-256 verified images)

split lines use
train3497 3,497 training only
dev150 = dev250 minus original100 150 every choice: checkpoint, image scale, decoding
original100 100 reported once at the end; never used for any choice
audit350 350 holdout for comparisons; never tuned on
test 1,374 competition test lines (no labels)

The splits are disjoint by line ID and by exact image hash.

How to use

Requirements: transformers==5.2.0, torch>=2.4 (bf16 GPU, ~12 GB for inference), Pillow. Optional speed-up: flash-linear-attention (needs Triton ≥ 3.3, i.e. PyTorch ≥ 2.7); without it transformers uses its PyTorch Gated-DeltaNet path (same results, slower).

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Abdoul27/chandra-ocr-2-barbados")          # after your access request is approved; log in with `huggingface-cli login`
sys.path.insert(0, path)
from barbados_inference import load, transcribe
model, processor = load(path)
print(transcribe(model, processor, ["line_001.jpg", "line_002.jpg"], batch_size=8))

What barbados_inference.py does (reproduce this exactly to get the scores above):

  1. Input: one image per handwritten line (not a full page). The image is resized by 1.0× (LANCZOS) and capped at 6,291,456 pixels; the processor then snaps it to multiples of 32 px (16-px patches, 2×2 merge).
  2. Prompt: the official Chandra "ocr" prompt (OCR_PROMPT, image first, then the text), through the model's own chat template (its generation prompt ends with an empty <think></think> block).
  3. Decoding: greedy, at most 60 new tokens, stop at <|im_end|> or <|endoftext|>; batched with left padding.
  4. Output: the model answers <p>…</p> with HTML escaping; html_to_text strips tags, unescapes entities (&amp; → &), turns <sup>x</sup> into ^x and normalises spaces. Repetition loops are trimmed only when the length cap is hit.

Transcription conventions learned from the training labels: original spelling and abbreviations are kept (pnts, Xpian, w^th); ^ marks raised (superscript) letters, e.g. w^th, Exec:^rs; & is kept; single spaces between words.

Training

  • Full fine-tune of all decoder and vision-encoder weights (embeddings frozen), FP32 master weights with BF16 autocast, 8-bit AdamW, lr 1e-05 (vision encoder 3e-06), 5 % warm-up, cosine to 10 %, effective batch 16, 4 epochs + a 2-epoch warm-restart extension (lr 3e-6 -> 1e-6), best checkpoint epoch_05_end. Evaluation twice per epoch; selection on dev150.
  • Targets in Chandra's native output format (<p> + HTML-escaped line + </p>); loss on the answer tokens only (incl. the end token).
  • Mild augmentation (±10 % scale, ±1.5° rotation, brightness/contrast, light blur); gradient checkpointing; flash-linear-attention kernels.
  • barbados_training_worker.py is the complete training/evaluation code.

Limitations

  • Line-level model: it expects one cropped text line per image; full pages need line segmentation first.
  • Trained on 3497 lines from one archive and period; other hands, languages or scripts may degrade.
  • Rare characters absent from the training labels cannot be expected in the output.

Licence

The base model weights are released under a modified OpenRAIL-M licence (see LICENSE); its use restrictions apply to this fine-tuned checkpoint as well. Base model: datalab-to/chandra-ocr-2 by Datalab.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Abdoul27/chandra-ocr-2-barbados

Finetuned
(7)
this model