--- license: apache-2.0 language: [en] pipeline_tag: image-text-to-text library_name: transformers base_model: PaddlePaddle/PaddleOCR-VL-1.6 tags: [document-ai, ocr, engineering-drawings, one-line-diagram, single-line-diagram, interconnection, energy] --- # decosa-oneline-reader-paddleocr-vl A 0.96B vision-language model that reads an electrical one-line (single-line) diagram for a solar or storage interconnection request, including poor fax-like scans, and returns one JSON object: the point-of-interconnection voltage, the plant output limit, every inverter group (quantity, model, rating), battery storage, the main transformer, and whether a disconnect switch, meter and main breaker are drawn. It is a full fine-tune of [PaddleOCR-VL-1.6](https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6). It was trained for one job: turning the drawing in an interconnection packet into fields that can be checked against the application form. It is not a general OCR or document parser any more. ## Results Four test sets of 60 drawings each, drawn by the same procedural generator as the training data but with seeds that were never used for training. "All fields right" means every field of the drawing is right (9 fields: POI kV, plant limit, inverter models, counts and kVA, storage MW and MWh, transformer MVA, disconnect). The baselines read the same images with the same scorer. | Test set (60 drawings each) | Base PaddleOCR-VL-1.6 | Qwen3.8-27B (vision) | **This model** | |---|---|---|---| | Poor scans (scaled down, blurred, skewed, speckled, JPEG) | 0 of 12 tried (no usable JSON) | 25 | **59** | | The clean set below, degraded the same way | — | 16 | **59** | | Clean drawings (regression check) | 0 of 12 tried | 59 | **60** | | Poor scans where every inverter model name and rating is random (never seen in training) | — | — | **50** | Per field on the poor scans, this model vs Qwen3.8-27B: inverter models 60 vs 38 of 60, inverter counts 91 vs 57 of 91, inverter kVA 91 vs 67 of 91; every other field 98-100% for both. On the random-name set the misses are one wrong character in a model name (for example 8908 read as 6908), a hyphen read as a space, or a count off by one. Speed: about 1.2 s per drawing in bf16 on one workstation GPU, batch 6 (2-3 s on the larger clean drawings). All numbers are in `eval_summary.json`. ## Intended use - First-pass reading of one-line diagrams in interconnection packets, so a reviewer can compare the drawing with the application form (capacity, inverter count, POI voltage, storage) without retyping it. - A second reader next to a general vision model on scans that model struggles with. Not for: approving or rejecting an interconnection request, engineering or protection studies, or any decision without a person checking the drawing. Treat every value as a reading to verify, not a fact. ## Limitations (read these) - **One drawing family.** Training and every test set come from one procedural drawing template (four label styles, three font sets, many degradations). The results show it reads that family well, even when badly scanned. **No real utility one-line has been scored.** Real drawings (CAD title blocks, dense protection schemes, several sheets, hand mark-ups) are out of distribution; expect missed or invented fields until it is tested and re-trained on them. - Invented model names: about 1 drawing in 8 has a one-character slip in a model name. Check model strings against the application or a datasheet list. - It reports what it read; it does not know whether a value is plausible. Pair it with arithmetic checks (inverter count × rating vs plant limit) and a person. - English labels, US-style units (kV, kVA, MW, MWh) only. ## Training | | | |---|---| | Base | PaddlePaddle/PaddleOCR-VL-1.6 @ c5630abae1d9 (ERNIE-4.5-0.3B language model + NaViT vision encoder), Apache-2.0 | | Data | 16,000 synthetic one-line diagrams drawn by a Decosa script: fictional plants, places and inverter models; 35% clean, 15% one fixed poor-scan degradation, 50% random degradations (scale 0.42-0.9, blur, skew up to 2.5°, fax threshold, grey paper, JPEG 20-75); half the drawings use random fictional model names and ratings. Fonts: Liberation (SIL OFL 1.1) and DejaVu. No real drawings, no personal data, nothing scraped. The generator and data are not released. | | Target | The compact JSON below, one object per drawing; every training target was parsed back and checked against the drawing's ground truth before use | | Method | Full fine-tune, AdamW (β 0.9/0.95), lr 3e-5 cosine with 40 warm-up steps, vision tower at 0.3× lr, bf16 autocast with fp32 master weights, loss on answer tokens only, token-budgeted batches (24k tokens), 1,707 steps (58 min on one H200), seed 20260928. Validation loss 2.55 → 0.0041. | | Selection | No choice was made on the test sets; the final checkpoint is the end of the time-boxed run. | ## Usage ```bash pip install "transformers>=5.17" torch pillow python usage.py drawing.png ``` ```python import json, torch from PIL import Image from transformers import AutoModelForImageTextToText, AutoProcessor repo = "decosaai/decosa-oneline-reader-paddleocr-vl" proc = AutoProcessor.from_pretrained(repo) model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16).to("cuda").eval() im = Image.open("drawing.png").convert("RGB") # see usage.py: small scans are enlarged to ~1 MP first msgs = [{"role": "user", "content": [{"type": "image", "image": im}, {"type": "text", "text": "One-line Diagram Recognition:"}]}] x = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to("cuda") out = model.generate(**x, max_new_tokens=900, do_sample=False, use_cache=True) reading = json.loads(proc.decode(out[0][x["input_ids"].shape[-1]:], skip_special_tokens=True)) ``` Pass `use_cache=True` to `generate()`: the base config ships with the cache off in the text config, which makes generation many times slower. Tested with transformers 5.17, which loads the model natively (no `trust_remote_code`); the base repo's code files are included only because `config.json` still names them in `auto_map`. Output (field order as trained; values here are placeholders): ```json {"is_one_line": true, "poi_kv": 0, "poi_label": "...", "plant_limit": {"value": 0, "unit": "MW", "text": "..."}, "inverter_groups": [{"text": "...", "qty": 0, "model": "...", "rating": 0, "rating_unit": "kVA", "type": "pv"}], "total_inverter_note": {"value": null, "unit": "kVA", "text": ""}, "storage": {"mw": 0, "mwh": 0, "text": "..."}, "transformer": {"text": "...", "max_rating": 0, "rating_unit": "MVA", "hv_kv": 0, "lv_kv": 0}, "disconnect": {"present": true, "lockable": true, "text": "..."}, "meter": {"present": true, "text": "..."}, "main_breaker": {"present": true, "text": "..."}} ``` One `inverter_groups` entry per feeder branch, even when the same model repeats; quantities are copied as drawn, not summed. Files: `model.safetensors` (bf16), `config.json`, `generation_config.json`, `processor_config.json`, tokenizer files, `chat_template.jinja`, the base model's `*_paddleocr_vl.py` code files (unchanged, Apache-2.0, PaddlePaddle Authors), `usage.py`, `eval_summary.json`, `SHA256SUMS`, `LICENSE`, `NOTICE`. ## Licence Apache-2.0 for the weights and `usage.py`. The base model and its code files are Apache-2.0, Copyright PaddlePaddle Authors; see NOTICE.