Instructions to use decosaai/decosa-oneline-reader-paddleocr-vl with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use decosaai/decosa-oneline-reader-paddleocr-vl with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="decosaai/decosa-oneline-reader-paddleocr-vl", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("decosaai/decosa-oneline-reader-paddleocr-vl", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("decosaai/decosa-oneline-reader-paddleocr-vl", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use decosaai/decosa-oneline-reader-paddleocr-vl with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "decosaai/decosa-oneline-reader-paddleocr-vl" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "decosaai/decosa-oneline-reader-paddleocr-vl", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/decosaai/decosa-oneline-reader-paddleocr-vl
- SGLang
How to use decosaai/decosa-oneline-reader-paddleocr-vl with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "decosaai/decosa-oneline-reader-paddleocr-vl" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "decosaai/decosa-oneline-reader-paddleocr-vl", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "decosaai/decosa-oneline-reader-paddleocr-vl" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "decosaai/decosa-oneline-reader-paddleocr-vl", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use decosaai/decosa-oneline-reader-paddleocr-vl with Docker Model Runner:
docker model run hf.co/decosaai/decosa-oneline-reader-paddleocr-vl
decosa-oneline-reader-paddleocr-vl
A 0.96B vision-language model that reads an electrical one-line (single-line) diagram for a solar or storage interconnection request, including poor fax-like scans, and returns one JSON object: the point-of-interconnection voltage, the plant output limit, every inverter group (quantity, model, rating), battery storage, the main transformer, and whether a disconnect switch, meter and main breaker are drawn.
It is a full fine-tune of PaddleOCR-VL-1.6. It was trained for one job: turning the drawing in an interconnection packet into fields that can be checked against the application form. It is not a general OCR or document parser any more.
Results
Four test sets of 60 drawings each, drawn by the same procedural generator as the training data but with seeds that were never used for training. "All fields right" means every field of the drawing is right (9 fields: POI kV, plant limit, inverter models, counts and kVA, storage MW and MWh, transformer MVA, disconnect). The baselines read the same images with the same scorer.
| Test set (60 drawings each) | Base PaddleOCR-VL-1.6 | Qwen3.8-27B (vision) | This model |
|---|---|---|---|
| Poor scans (scaled down, blurred, skewed, speckled, JPEG) | 0 of 12 tried (no usable JSON) | 25 | 59 |
| The clean set below, degraded the same way | — | 16 | 59 |
| Clean drawings (regression check) | 0 of 12 tried | 59 | 60 |
| Poor scans where every inverter model name and rating is random (never seen in training) | — | — | 50 |
Per field on the poor scans, this model vs Qwen3.8-27B: inverter models 60 vs 38 of 60, inverter counts 91 vs 57 of 91, inverter kVA 91 vs 67 of 91; every other field 98-100% for both. On the random-name set the misses are one wrong character in a model name (for example 8908 read as 6908), a hyphen read as a space, or a count off by one.
Speed: about 1.2 s per drawing in bf16 on one workstation GPU, batch 6 (2-3 s on the larger clean drawings).
All numbers are in eval_summary.json.
Intended use
- First-pass reading of one-line diagrams in interconnection packets, so a reviewer can compare the drawing with the application form (capacity, inverter count, POI voltage, storage) without retyping it.
- A second reader next to a general vision model on scans that model struggles with.
Not for: approving or rejecting an interconnection request, engineering or protection studies, or any decision without a person checking the drawing. Treat every value as a reading to verify, not a fact.
Limitations (read these)
- One drawing family. Training and every test set come from one procedural drawing template (four label styles, three font sets, many degradations). The results show it reads that family well, even when badly scanned. No real utility one-line has been scored. Real drawings (CAD title blocks, dense protection schemes, several sheets, hand mark-ups) are out of distribution; expect missed or invented fields until it is tested and re-trained on them.
- Invented model names: about 1 drawing in 8 has a one-character slip in a model name. Check model strings against the application or a datasheet list.
- It reports what it read; it does not know whether a value is plausible. Pair it with arithmetic checks (inverter count × rating vs plant limit) and a person.
- English labels, US-style units (kV, kVA, MW, MWh) only.
Training
| Base | PaddlePaddle/PaddleOCR-VL-1.6 @ c5630abae1d9 (ERNIE-4.5-0.3B language model + NaViT vision encoder), Apache-2.0 |
| Data | 16,000 synthetic one-line diagrams drawn by a Decosa script: fictional plants, places and inverter models; 35% clean, 15% one fixed poor-scan degradation, 50% random degradations (scale 0.42-0.9, blur, skew up to 2.5°, fax threshold, grey paper, JPEG 20-75); half the drawings use random fictional model names and ratings. Fonts: Liberation (SIL OFL 1.1) and DejaVu. No real drawings, no personal data, nothing scraped. The generator and data are not released. |
| Target | The compact JSON below, one object per drawing; every training target was parsed back and checked against the drawing's ground truth before use |
| Method | Full fine-tune, AdamW (β 0.9/0.95), lr 3e-5 cosine with 40 warm-up steps, vision tower at 0.3× lr, bf16 autocast with fp32 master weights, loss on answer tokens only, token-budgeted batches (24k tokens), 1,707 steps (58 min on one H200), seed 20260928. Validation loss 2.55 → 0.0041. |
| Selection | No choice was made on the test sets; the final checkpoint is the end of the time-boxed run. |
Usage
pip install "transformers>=5.17" torch pillow
python usage.py drawing.png
import json, torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
repo = "decosaai/decosa-oneline-reader-paddleocr-vl"
proc = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16).to("cuda").eval()
im = Image.open("drawing.png").convert("RGB") # see usage.py: small scans are enlarged to ~1 MP first
msgs = [{"role": "user", "content": [{"type": "image", "image": im},
{"type": "text", "text": "One-line Diagram Recognition:"}]}]
x = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True,
return_tensors="pt").to("cuda")
out = model.generate(**x, max_new_tokens=900, do_sample=False, use_cache=True)
reading = json.loads(proc.decode(out[0][x["input_ids"].shape[-1]:], skip_special_tokens=True))
Pass use_cache=True to generate(): the base config ships with the cache off in the text config, which makes
generation many times slower. Tested with transformers 5.17, which loads the model natively (no trust_remote_code);
the base repo's code files are included only because config.json still names them in auto_map.
Output (field order as trained; values here are placeholders):
{"is_one_line": true, "poi_kv": 0, "poi_label": "...",
"plant_limit": {"value": 0, "unit": "MW", "text": "..."},
"inverter_groups": [{"text": "...", "qty": 0, "model": "...", "rating": 0, "rating_unit": "kVA", "type": "pv"}],
"total_inverter_note": {"value": null, "unit": "kVA", "text": ""},
"storage": {"mw": 0, "mwh": 0, "text": "..."},
"transformer": {"text": "...", "max_rating": 0, "rating_unit": "MVA", "hv_kv": 0, "lv_kv": 0},
"disconnect": {"present": true, "lockable": true, "text": "..."},
"meter": {"present": true, "text": "..."},
"main_breaker": {"present": true, "text": "..."}}
One inverter_groups entry per feeder branch, even when the same model repeats; quantities are copied as drawn, not
summed.
Files: model.safetensors (bf16), config.json, generation_config.json, processor_config.json, tokenizer files,
chat_template.jinja, the base model's *_paddleocr_vl.py code files (unchanged, Apache-2.0, PaddlePaddle Authors),
usage.py, eval_summary.json, SHA256SUMS, LICENSE, NOTICE.
Licence
Apache-2.0 for the weights and usage.py. The base model and its code files are Apache-2.0, Copyright PaddlePaddle
Authors; see NOTICE.
- Downloads last month
- 25
Model tree for decosaai/decosa-oneline-reader-paddleocr-vl
Base model
PaddlePaddle/PaddleOCR-VL-1.6