decosa-oneline-reader-paddleocr-vl

A 0.96B vision-language model that reads an electrical one-line (single-line) diagram for a solar or storage interconnection request, including poor fax-like scans, and returns one JSON object: the point-of-interconnection voltage, the plant output limit, every inverter group (quantity, model, rating), battery storage, the main transformer, and whether a disconnect switch, meter and main breaker are drawn.

It is a full fine-tune of PaddleOCR-VL-1.6. It was trained for one job: turning the drawing in an interconnection packet into fields that can be checked against the application form. It is not a general OCR or document parser any more.

Results

Four test sets of 60 drawings each, drawn by the same procedural generator as the training data but with seeds that were never used for training. "All fields right" means every field of the drawing is right (9 fields: POI kV, plant limit, inverter models, counts and kVA, storage MW and MWh, transformer MVA, disconnect). The baselines read the same images with the same scorer.

Test set (60 drawings each) Base PaddleOCR-VL-1.6 Qwen3.8-27B (vision) This model
Poor scans (scaled down, blurred, skewed, speckled, JPEG) 0 of 12 tried (no usable JSON) 25 59
The clean set below, degraded the same way — 16 59
Clean drawings (regression check) 0 of 12 tried 59 60
Poor scans where every inverter model name and rating is random (never seen in training) — — 50

Per field on the poor scans, this model vs Qwen3.8-27B: inverter models 60 vs 38 of 60, inverter counts 91 vs 57 of 91, inverter kVA 91 vs 67 of 91; every other field 98-100% for both. On the random-name set the misses are one wrong character in a model name (for example 8908 read as 6908), a hyphen read as a space, or a count off by one.

Speed: about 1.2 s per drawing in bf16 on one workstation GPU, batch 6 (2-3 s on the larger clean drawings).

All numbers are in eval_summary.json.

Intended use

  • First-pass reading of one-line diagrams in interconnection packets, so a reviewer can compare the drawing with the application form (capacity, inverter count, POI voltage, storage) without retyping it.
  • A second reader next to a general vision model on scans that model struggles with.

Not for: approving or rejecting an interconnection request, engineering or protection studies, or any decision without a person checking the drawing. Treat every value as a reading to verify, not a fact.

Limitations (read these)

  • One drawing family. Training and every test set come from one procedural drawing template (four label styles, three font sets, many degradations). The results show it reads that family well, even when badly scanned. No real utility one-line has been scored. Real drawings (CAD title blocks, dense protection schemes, several sheets, hand mark-ups) are out of distribution; expect missed or invented fields until it is tested and re-trained on them.
  • Invented model names: about 1 drawing in 8 has a one-character slip in a model name. Check model strings against the application or a datasheet list.
  • It reports what it read; it does not know whether a value is plausible. Pair it with arithmetic checks (inverter count × rating vs plant limit) and a person.
  • English labels, US-style units (kV, kVA, MW, MWh) only.

Training

Base PaddlePaddle/PaddleOCR-VL-1.6 @ c5630abae1d9 (ERNIE-4.5-0.3B language model + NaViT vision encoder), Apache-2.0
Data 16,000 synthetic one-line diagrams drawn by a Decosa script: fictional plants, places and inverter models; 35% clean, 15% one fixed poor-scan degradation, 50% random degradations (scale 0.42-0.9, blur, skew up to 2.5°, fax threshold, grey paper, JPEG 20-75); half the drawings use random fictional model names and ratings. Fonts: Liberation (SIL OFL 1.1) and DejaVu. No real drawings, no personal data, nothing scraped. The generator and data are not released.
Target The compact JSON below, one object per drawing; every training target was parsed back and checked against the drawing's ground truth before use
Method Full fine-tune, AdamW (β 0.9/0.95), lr 3e-5 cosine with 40 warm-up steps, vision tower at 0.3× lr, bf16 autocast with fp32 master weights, loss on answer tokens only, token-budgeted batches (24k tokens), 1,707 steps (58 min on one H200), seed 20260928. Validation loss 2.55 → 0.0041.
Selection No choice was made on the test sets; the final checkpoint is the end of the time-boxed run.

Usage

pip install "transformers>=5.17" torch pillow
python usage.py drawing.png
import json, torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor

repo = "decosaai/decosa-oneline-reader-paddleocr-vl"
proc = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16).to("cuda").eval()

im = Image.open("drawing.png").convert("RGB")      # see usage.py: small scans are enlarged to ~1 MP first
msgs = [{"role": "user", "content": [{"type": "image", "image": im},
                                     {"type": "text", "text": "One-line Diagram Recognition:"}]}]
x = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True,
                             return_tensors="pt").to("cuda")
out = model.generate(**x, max_new_tokens=900, do_sample=False, use_cache=True)
reading = json.loads(proc.decode(out[0][x["input_ids"].shape[-1]:], skip_special_tokens=True))

Pass use_cache=True to generate(): the base config ships with the cache off in the text config, which makes generation many times slower. Tested with transformers 5.17, which loads the model natively (no trust_remote_code); the base repo's code files are included only because config.json still names them in auto_map.

Output (field order as trained; values here are placeholders):

{"is_one_line": true, "poi_kv": 0, "poi_label": "...",
 "plant_limit": {"value": 0, "unit": "MW", "text": "..."},
 "inverter_groups": [{"text": "...", "qty": 0, "model": "...", "rating": 0, "rating_unit": "kVA", "type": "pv"}],
 "total_inverter_note": {"value": null, "unit": "kVA", "text": ""},
 "storage": {"mw": 0, "mwh": 0, "text": "..."},
 "transformer": {"text": "...", "max_rating": 0, "rating_unit": "MVA", "hv_kv": 0, "lv_kv": 0},
 "disconnect": {"present": true, "lockable": true, "text": "..."},
 "meter": {"present": true, "text": "..."},
 "main_breaker": {"present": true, "text": "..."}}

One inverter_groups entry per feeder branch, even when the same model repeats; quantities are copied as drawn, not summed.

Files: model.safetensors (bf16), config.json, generation_config.json, processor_config.json, tokenizer files, chat_template.jinja, the base model's *_paddleocr_vl.py code files (unchanged, Apache-2.0, PaddlePaddle Authors), usage.py, eval_summary.json, SHA256SUMS, LICENSE, NOTICE.

Licence

Apache-2.0 for the weights and usage.py. The base model and its code files are Apache-2.0, Copyright PaddlePaddle Authors; see NOTICE.

Downloads last month
25
Safetensors
Model size
0.9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for decosaai/decosa-oneline-reader-paddleocr-vl

Finetuned
(13)
this model