You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This model is a research checkpoint shared on request. Access is granted manually by the authors after reviewing your request. Please tell us who you are and what you want to use it for. Use must respect the base model licence.

Log in or Sign Up to review the conditions and access this model content.

GLM-OCR — fine-tuned for historical Barbados handwriting (line level)

Full fine-tune of zai-org/GLM-OCR (~1B parameters: CogViT-style vision encoder + GLM decoder, native in transformers as GlmOcrForConditionalGeneration) for transcribing single handwritten text lines from historical Barbados records (17th–19th-century English legal and administrative documents). Access is reviewed manually.

Local scores

Metric (competition metric): score = 1 − WER_w/24 − CER_w/110, where word and character edit distances are weighted per line by √(reference length). WER/CER columns are plain micro-averaged percentages. Greedy decoding at image scale 1.0.

split lines score WER % CER % exact lines role
dev minus original100 ("dev150") 150 0.8925 15.86 5.29 32 selection set: checkpoint and image scale chosen here
dev250 (dev150 + original100) 250 0.8872 16.93 5.21 48
original100 100 0.8792 18.51 5.09 16 reported once, never used for any choice
dev150, zero-shot base model 150 0.7175 39.29 16.09 1 scale 1.0, same prompt

Data splits (frozen, SHA-256 verified images)

split lines use
train3497 3,497 training only
dev150 = dev250 minus original100 150 every choice: checkpoint, image scale, decoding
original100 100 reported once at the end; never used for any choice
audit350 350 holdout for comparisons; never tuned on
test 1,374 competition test lines (no labels)

The splits are disjoint by line ID and by exact image hash.

Requirements

  • transformers==5.2.0 (GLM-OCR is native: GlmOcrForConditionalGeneration; no remote code), torch>=2.4 with torchvision, Pillow, huggingface_hub. Tested with PyTorch 2.x + CUDA, BF16.
  • GPU: ~4 GB for inference in BF16 (batch 16 of line images); CPU works but is slow.

How to use

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Abdoul27/glm-ocr-barbados")          # after your access request is approved; log in with `huggingface-cli login`
sys.path.insert(0, path)
from barbados_inference import load, transcribe
model, processor = load(path)
print(transcribe(model, processor, ["line_001.jpg", "line_002.jpg"], batch_size=16))

What barbados_inference.py does (reproduce this exactly to get the scores above):

  1. Input: one image per handwritten line (not a full page), resized by 1.0× (LANCZOS) and capped at 6,291,456 pixels; the processor snaps it to multiples of 28 px (14-px patches, 2×2 merge).
  2. Prompt: the official GLM-OCR prompt Text Recognition: (image first, then this text) through the model's chat template; answer prefix '' (the empty <think></think> block is not used).
  3. Decoding: greedy, at most 61 new tokens, stop at token ids [59246, 59253] (the model was trained to end with 59253); batched with left padding.
  4. Output: glm_to_text removes any <think> block, turns <sup>x</sup> / $^{x}$ superscripts into ^x, strips LaTeX $ and HTML tags, unescapes entities and normalises spaces. Repetition loops are trimmed only when the length cap is hit.

barbados_inference.json holds the same settings in machine-readable form.

Transcription conventions learned from the training labels: original spelling and abbreviations are kept (pnts, Xpian, w^th); ^ marks raised (superscript) letters, e.g. w^th, Exec:^rs; & is kept; single spaces between words.

Training

  • Full fine-tune of every weight (vision encoder, projector, decoder, embeddings), FP32 master weights with BF16 autocast, AdamW, lr 2e-05 (vision encoder 4e-06), 5 % warm-up, cosine to 10 %, effective batch 16; 6 epochs, best checkpoint epoch_06_end (selected on dev150). Evaluation twice per epoch.
  • Targets: the plain line text + the end token the model itself uses; loss on the answer tokens only.
  • Mild augmentation (±10 % scale, ±1.5° rotation, brightness/contrast, light blur); gradient checkpointing.
  • barbados_training_worker.py is the complete training/evaluation code.

Limitations

  • Line-level model: it expects one cropped text line per image; full pages need line segmentation first.
  • Trained on 3497 lines from one archive and period; other hands, languages or scripts may degrade.
  • Rare characters absent from the training labels cannot be expected in the output.

Licence

This fine-tune inherits the licence of the base model zai-org/GLM-OCR by Z.ai (mit; see LICENSE).

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Abdoul27/glm-ocr-barbados

Base model

zai-org/GLM-OCR
Finetuned
(31)
this model