Instructions to use Abdoul27/dots-mocr-barbados-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Abdoul27/dots-mocr-barbados-v4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Abdoul27/dots-mocr-barbados-v4", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Abdoul27/dots-mocr-barbados-v4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Abdoul27/dots-mocr-barbados-v4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Abdoul27/dots-mocr-barbados-v4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/dots-mocr-barbados-v4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Abdoul27/dots-mocr-barbados-v4
- SGLang
How to use Abdoul27/dots-mocr-barbados-v4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Abdoul27/dots-mocr-barbados-v4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/dots-mocr-barbados-v4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Abdoul27/dots-mocr-barbados-v4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdoul27/dots-mocr-barbados-v4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Abdoul27/dots-mocr-barbados-v4 with Docker Model Runner:
docker model run hf.co/Abdoul27/dots-mocr-barbados-v4
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This model is a research checkpoint shared on request. Access is granted manually by the authors after reviewing your request. Please tell us who you are and what you want to use it for. Use must respect the base model licence.
Log in or Sign Up to review the conditions and access this model content.
dots.mocr (3B) — fine-tuned for historical Barbados handwriting, line level (v4)
Full fine-tune of dots-studio/dots.mocr (rednote-hilab: 1.2B NaViT vision encoder at native resolution +
Qwen2-style 1.5B language model; remote code) for transcribing single handwritten text lines from historical Barbados records
(17th–19th-century English legal and administrative documents). Access is reviewed manually.
Local scores
Metric (competition metric): score = 1 − WER_w/24 − CER_w/110, where word and character edit distances are weighted per line by √(reference length). WER/CER columns are plain micro-averaged percentages. Greedy decoding, image scale 1.5.
| split | lines | score | WER % | CER % | exact lines | role |
|---|---|---|---|---|---|---|
| dev150 (dev250 minus original100) | 150 | 0.9129 | 13.19 | 4.00 | 37 | selection set: checkpoint chosen here |
| dev250 (dev150 + original100) | 250 | 0.9073 | 14.11 | 4.14 | 55 | |
| original100 | 100 | 0.8991 | 15.45 | 4.33 | 18 | reported once, never used for any choice |
Compared with the previous version (same model, same recipe, trained on train3497 only):
| version | dev150 | original100 |
|---|---|---|
| previous version (train3497 only) | 0.9013 | 0.8949 |
| this version (v4) | 0.9129 | 0.8991 |
audit350 is training data for this version, so it cannot serve as a holdout; dev150 and original100 are the only fair local comparisons. original100 has 100 lines, so differences below ~0.005 are within noise.
Training data
| source | lines | note |
|---|---|---|
| train3497 | 3497 | labelled training lines |
| audit350 | 350 | labelled lines that earlier versions kept as a holdout |
| pseudo-labelled test lines | 650 | competition test lines transcribed by our best ensemble (semi-supervised); never overlapping a labelled line (ID and SHA-256 checked) |
Selection used dev150 only (dev250 minus original100); original100 was scored once at the end. All splits are disjoint by line ID and by exact image hash.
Requirements
transformers==4.57.6(the version pinned by dots.mocr; the model uses remote code:trust_remote_code=True),torch>=2.4withtorchvision,Pillow,huggingface_hub.flash-attnis optional (without it,barbados_inference.pyregisters a stub because the vision code imports it unconditionally, and uses SDPA).- Weights are stored in FP32 (as trained and evaluated; ~12 GB); inference runs under BF16 autocast. GPU with ≥16 GB recommended.
How to use
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Abdoul27/dots-mocr-barbados-v4") # after your access request is approved; log in with `huggingface-cli login`
sys.path.insert(0, path)
from barbados_inference import load, transcribe
model, processor = load(path)
print(transcribe(model, processor, ["line_001.jpg", "line_002.jpg"]))
What barbados_inference.py does (reproduce this exactly to get the scores above):
- Input: one image per handwritten line (not a full page), resized by 1.5× (bicubic) and capped at 1,605,632 pixels (2,048 visual tokens; one token per 28×28 px after the processor's own rounding).
- Prompt: the official dots.mocr OCR prompt
Extract the text content from this image., no system prompt, image before text, through the model's chat template (<|user|><|img|>…<|endofimg|>{prompt}<|endofuser|><|assistant|>). - Decoding: greedy, one line at a time, at most 51 new tokens, stop at <|endofassistant|>, <|endoftext|>, <|assistant|> (the model was trained to end its
answer with
<|endofassistant|>). - Output: special tokens removed, line breaks joined with a space; repetition loops are trimmed only when the length cap is hit.
barbados_inference.json holds the same settings in machine-readable form.
Transcription conventions learned from the labels: original spelling and abbreviations kept (pnts, Xpian), ^ before raised
(superscript) letters (w^th, Exec^rs, S^t), & kept, single spaces between words.
Training
- Full fine-tune of every weight, FP32 master weights with BF16 autocast, AdamW (β (0.9, 0.95)), lr 1e-05 (vision encoder 2e-06), weight decay 0.1 on matrices, 5 % warm-up + cosine to 10 %, effective batch 16; 5 epochs; selected checkpoint epoch_05_end (best on dev150). Answer-only loss including the end token; gradient checkpointing; micro-batches with right padding.
- Augmentation: scale ±10 %, rotation ±1.5°, brightness/contrast, light blur, mild elastic distortion (25 %), slant/shear ±4.6° (30 %), light Gaussian noise (20 %). No dilation/erosion, perspective, cut-out or StackMix.
barbados_training_worker.pyis the complete training/evaluation code.
Limitations
- Line-level model: it expects one cropped text line per image; full pages need line segmentation first.
- Trained on one archive and period; other hands, languages or scripts may degrade. Part of the training data are machine-made pseudo-labels.
- Rare characters absent from the training labels cannot be expected in the output.
Licence
This fine-tune inherits the licence of the base model dots-studio/dots.mocr (mit; see LICENSE if present).
- Downloads last month
- 2
Model tree for Abdoul27/dots-mocr-barbados-v4
Base model
dots-studio/dots.mocr