Instructions to use Abdoul27/glm-ocr-barbados with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Abdoul27/glm-ocr-barbados with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Abdoul27/glm-ocr-barbados") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMultimodalLM tokenizer = AutoTokenizer.from_pretrained("Abdoul27/glm-ocr-barbados") model = AutoModelForMultimodalLM.from_pretrained("Abdoul27/glm-ocr-barbados", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Abdoul27/glm-ocr-barbados with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Abdoul27/glm-ocr-barbados" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/glm-ocr-barbados", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Abdoul27/glm-ocr-barbados
- SGLang
How to use Abdoul27/glm-ocr-barbados with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Abdoul27/glm-ocr-barbados" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/glm-ocr-barbados", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Abdoul27/glm-ocr-barbados" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/glm-ocr-barbados", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Abdoul27/glm-ocr-barbados with Docker Model Runner:
docker model run hf.co/Abdoul27/glm-ocr-barbados
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This model is a research checkpoint shared on request. Access is granted manually by the authors after reviewing your request. Please tell us who you are and what you want to use it for. Use must respect the base model licence.
Log in or Sign Up to review the conditions and access this model content.
GLM-OCR — fine-tuned for historical Barbados handwriting (line level)
Full fine-tune of zai-org/GLM-OCR (~1B parameters: CogViT-style vision encoder + GLM decoder,
native in transformers as GlmOcrForConditionalGeneration) for transcribing single handwritten text lines from historical Barbados
records (17th–19th-century English legal and administrative documents). Access is reviewed manually.
Local scores
Metric (competition metric): score = 1 − WER_w/24 − CER_w/110, where word and character edit distances are weighted per line by √(reference length). WER/CER columns are plain micro-averaged percentages. Greedy decoding at image scale 1.0.
| split | lines | score | WER % | CER % | exact lines | role |
|---|---|---|---|---|---|---|
| dev minus original100 ("dev150") | 150 | 0.8925 | 15.86 | 5.29 | 32 | selection set: checkpoint and image scale chosen here |
| dev250 (dev150 + original100) | 250 | 0.8872 | 16.93 | 5.21 | 48 | |
| original100 | 100 | 0.8792 | 18.51 | 5.09 | 16 | reported once, never used for any choice |
| dev150, zero-shot base model | 150 | 0.7175 | 39.29 | 16.09 | 1 | scale 1.0, same prompt |
Data splits (frozen, SHA-256 verified images)
| split | lines | use |
|---|---|---|
| train3497 | 3,497 | training only |
| dev150 = dev250 minus original100 | 150 | every choice: checkpoint, image scale, decoding |
| original100 | 100 | reported once at the end; never used for any choice |
| audit350 | 350 | holdout for comparisons; never tuned on |
| test | 1,374 | competition test lines (no labels) |
The splits are disjoint by line ID and by exact image hash.
Requirements
transformers==5.2.0(GLM-OCR is native:GlmOcrForConditionalGeneration; no remote code),torch>=2.4withtorchvision,Pillow,huggingface_hub. Tested with PyTorch 2.x + CUDA, BF16.- GPU: ~4 GB for inference in BF16 (batch 16 of line images); CPU works but is slow.
How to use
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Abdoul27/glm-ocr-barbados") # after your access request is approved; log in with `huggingface-cli login`
sys.path.insert(0, path)
from barbados_inference import load, transcribe
model, processor = load(path)
print(transcribe(model, processor, ["line_001.jpg", "line_002.jpg"], batch_size=16))
What barbados_inference.py does (reproduce this exactly to get the scores above):
- Input: one image per handwritten line (not a full page), resized by 1.0× (LANCZOS) and capped at 6,291,456 pixels; the processor snaps it to multiples of 28 px (14-px patches, 2×2 merge).
- Prompt: the official GLM-OCR prompt
Text Recognition:(image first, then this text) through the model's chat template; answer prefix''(the empty<think></think>block is not used). - Decoding: greedy, at most 61 new tokens, stop at token ids [59246, 59253] (the model was trained to end with 59253); batched with left padding.
- Output:
glm_to_textremoves any<think>block, turns<sup>x</sup>/$^{x}$superscripts into^x, strips LaTeX$and HTML tags, unescapes entities and normalises spaces. Repetition loops are trimmed only when the length cap is hit.
barbados_inference.json holds the same settings in machine-readable form.
Transcription conventions learned from the training labels: original spelling and abbreviations are kept (pnts, Xpian, w^th);
^ marks raised (superscript) letters, e.g. w^th, Exec:^rs; & is kept; single spaces between words.
Training
- Full fine-tune of every weight (vision encoder, projector, decoder, embeddings), FP32 master weights with BF16 autocast, AdamW, lr 2e-05 (vision encoder 4e-06), 5 % warm-up, cosine to 10 %, effective batch 16; 6 epochs, best checkpoint epoch_06_end (selected on dev150). Evaluation twice per epoch.
- Targets: the plain line text + the end token the model itself uses; loss on the answer tokens only.
- Mild augmentation (±10 % scale, ±1.5° rotation, brightness/contrast, light blur); gradient checkpointing.
barbados_training_worker.pyis the complete training/evaluation code.
Limitations
- Line-level model: it expects one cropped text line per image; full pages need line segmentation first.
- Trained on 3497 lines from one archive and period; other hands, languages or scripts may degrade.
- Rare characters absent from the training labels cannot be expected in the output.
Licence
This fine-tune inherits the licence of the base model zai-org/GLM-OCR by Z.ai (mit; see LICENSE).
- Downloads last month
- -
Model tree for Abdoul27/glm-ocr-barbados
Base model
zai-org/GLM-OCR