sipratt commited on
Commit
a13acce
·
verified ·
1 Parent(s): 65e5a5e

decosa-oneline-reader-paddleocr-vl v1: PaddleOCR-VL-1.6 fine-tune for one-line diagrams (Apache-2.0)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright (c) 2025 PaddlePaddle Authors. All Rights Reserved.
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
NOTICE ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ decosa-oneline-reader-paddleocr-vl
2
+ Copyright 2026 Decosa. Licensed under the Apache License, Version 2.0 (see LICENSE).
3
+
4
+ Base model: PaddleOCR-VL-1.6 (PaddlePaddle/PaddleOCR-VL-1.6, revision c5630abae1d940eafe0697512a0325494b02ab42),
5
+ Copyright (c) 2025 PaddlePaddle Authors, Apache License 2.0. This repository contains a full fine-tune of its weights.
6
+ The files configuration_paddleocr_vl.py, modeling_paddleocr_vl.py, image_processing_paddleocr_vl.py and
7
+ processing_paddleocr_vl.py are copied unchanged from the base repository (Apache License 2.0, PaddlePaddle Authors).
8
+ The tokenizer and chat template come from the base model.
9
+
10
+ Training data: synthetic one-line diagrams drawn by a Decosa script (fictional plants, places and equipment). Text was
11
+ rendered with the Liberation fonts (SIL Open Font License 1.1, Red Hat / Google) and DejaVu fonts (Bitstream Vera and
12
+ DejaVu licence). No font files, drawings or training data are redistributed here.
README.md ADDED
@@ -0,0 +1,127 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language: [en]
4
+ pipeline_tag: image-text-to-text
5
+ library_name: transformers
6
+ base_model: PaddlePaddle/PaddleOCR-VL-1.6
7
+ tags: [document-ai, ocr, engineering-drawings, one-line-diagram, single-line-diagram, interconnection, energy]
8
+ ---
9
+
10
+ # decosa-oneline-reader-paddleocr-vl
11
+
12
+ A 0.96B vision-language model that reads an electrical one-line (single-line) diagram for a solar or storage
13
+ interconnection request, including poor fax-like scans, and returns one JSON object: the point-of-interconnection
14
+ voltage, the plant output limit, every inverter group (quantity, model, rating), battery storage, the main transformer,
15
+ and whether a disconnect switch, meter and main breaker are drawn.
16
+
17
+ It is a full fine-tune of [PaddleOCR-VL-1.6](https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6). It was trained for
18
+ one job: turning the drawing in an interconnection packet into fields that can be checked against the application form.
19
+ It is not a general OCR or document parser any more.
20
+
21
+ ## Results
22
+
23
+ Four test sets of 60 drawings each, drawn by the same procedural generator as the training data but with seeds that
24
+ were never used for training. "All fields right" means every field of the drawing is right (9 fields: POI kV, plant
25
+ limit, inverter models, counts and kVA, storage MW and MWh, transformer MVA, disconnect). The baselines read the same
26
+ images with the same scorer.
27
+
28
+ | Test set (60 drawings each) | Base PaddleOCR-VL-1.6 | Qwen3.8-27B (vision) | **This model** |
29
+ |---|---|---|---|
30
+ | Poor scans (scaled down, blurred, skewed, speckled, JPEG) | 0 of 12 tried (no usable JSON) | 25 | **59** |
31
+ | The clean set below, degraded the same way | — | 16 | **59** |
32
+ | Clean drawings (regression check) | 0 of 12 tried | 59 | **60** |
33
+ | Poor scans where every inverter model name and rating is random (never seen in training) | — | — | **50** |
34
+
35
+ Per field on the poor scans, this model vs Qwen3.8-27B: inverter models 60 vs 38 of 60, inverter counts 91 vs 57 of 91,
36
+ inverter kVA 91 vs 67 of 91; every other field 98-100% for both. On the random-name set the misses are one wrong
37
+ character in a model name (for example 8908 read as 6908), a hyphen read as a space, or a count off by one.
38
+
39
+ Speed: about 1.2 s per drawing in bf16 on one workstation GPU, batch 6 (2-3 s on the larger clean drawings).
40
+
41
+ All numbers are in `eval_summary.json`.
42
+
43
+ ## Intended use
44
+
45
+ - First-pass reading of one-line diagrams in interconnection packets, so a reviewer can compare the drawing with the
46
+ application form (capacity, inverter count, POI voltage, storage) without retyping it.
47
+ - A second reader next to a general vision model on scans that model struggles with.
48
+
49
+ Not for: approving or rejecting an interconnection request, engineering or protection studies, or any decision without
50
+ a person checking the drawing. Treat every value as a reading to verify, not a fact.
51
+
52
+ ## Limitations (read these)
53
+
54
+ - **One drawing family.** Training and every test set come from one procedural drawing template (four label styles,
55
+ three font sets, many degradations). The results show it reads that family well, even when badly scanned. **No real
56
+ utility one-line has been scored.** Real drawings (CAD title blocks, dense protection schemes, several sheets,
57
+ hand mark-ups) are out of distribution; expect missed or invented fields until it is tested and re-trained on them.
58
+ - Invented model names: about 1 drawing in 8 has a one-character slip in a model name. Check model strings against the
59
+ application or a datasheet list.
60
+ - It reports what it read; it does not know whether a value is plausible. Pair it with arithmetic checks (inverter
61
+ count × rating vs plant limit) and a person.
62
+ - English labels, US-style units (kV, kVA, MW, MWh) only.
63
+
64
+ ## Training
65
+
66
+ | | |
67
+ |---|---|
68
+ | Base | PaddlePaddle/PaddleOCR-VL-1.6 @ c5630abae1d9 (ERNIE-4.5-0.3B language model + NaViT vision encoder), Apache-2.0 |
69
+ | Data | 16,000 synthetic one-line diagrams drawn by a Decosa script: fictional plants, places and inverter models; 35% clean, 15% one fixed poor-scan degradation, 50% random degradations (scale 0.42-0.9, blur, skew up to 2.5°, fax threshold, grey paper, JPEG 20-75); half the drawings use random fictional model names and ratings. Fonts: Liberation (SIL OFL 1.1) and DejaVu. No real drawings, no personal data, nothing scraped. The generator and data are not released. |
70
+ | Target | The compact JSON below, one object per drawing; every training target was parsed back and checked against the drawing's ground truth before use |
71
+ | Method | Full fine-tune, AdamW (β 0.9/0.95), lr 3e-5 cosine with 40 warm-up steps, vision tower at 0.3× lr, bf16 autocast with fp32 master weights, loss on answer tokens only, token-budgeted batches (24k tokens), 1,707 steps (58 min on one H200), seed 20260928. Validation loss 2.55 → 0.0041. |
72
+ | Selection | No choice was made on the test sets; the final checkpoint is the end of the time-boxed run. |
73
+
74
+ ## Usage
75
+
76
+ ```bash
77
+ pip install "transformers>=5.17" torch pillow
78
+ python usage.py drawing.png
79
+ ```
80
+
81
+ ```python
82
+ import json, torch
83
+ from PIL import Image
84
+ from transformers import AutoModelForImageTextToText, AutoProcessor
85
+
86
+ repo = "decosaai/decosa-oneline-reader-paddleocr-vl"
87
+ proc = AutoProcessor.from_pretrained(repo)
88
+ model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16).to("cuda").eval()
89
+
90
+ im = Image.open("drawing.png").convert("RGB") # see usage.py: small scans are enlarged to ~1 MP first
91
+ msgs = [{"role": "user", "content": [{"type": "image", "image": im},
92
+ {"type": "text", "text": "One-line Diagram Recognition:"}]}]
93
+ x = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True,
94
+ return_tensors="pt").to("cuda")
95
+ out = model.generate(**x, max_new_tokens=900, do_sample=False, use_cache=True)
96
+ reading = json.loads(proc.decode(out[0][x["input_ids"].shape[-1]:], skip_special_tokens=True))
97
+ ```
98
+
99
+ Pass `use_cache=True` to `generate()`: the base config ships with the cache off in the text config, which makes
100
+ generation many times slower. Tested with transformers 5.17, which loads the model natively (no `trust_remote_code`);
101
+ the base repo's code files are included only because `config.json` still names them in `auto_map`.
102
+
103
+ Output (field order as trained; values here are placeholders):
104
+
105
+ ```json
106
+ {"is_one_line": true, "poi_kv": 0, "poi_label": "...",
107
+ "plant_limit": {"value": 0, "unit": "MW", "text": "..."},
108
+ "inverter_groups": [{"text": "...", "qty": 0, "model": "...", "rating": 0, "rating_unit": "kVA", "type": "pv"}],
109
+ "total_inverter_note": {"value": null, "unit": "kVA", "text": ""},
110
+ "storage": {"mw": 0, "mwh": 0, "text": "..."},
111
+ "transformer": {"text": "...", "max_rating": 0, "rating_unit": "MVA", "hv_kv": 0, "lv_kv": 0},
112
+ "disconnect": {"present": true, "lockable": true, "text": "..."},
113
+ "meter": {"present": true, "text": "..."},
114
+ "main_breaker": {"present": true, "text": "..."}}
115
+ ```
116
+
117
+ One `inverter_groups` entry per feeder branch, even when the same model repeats; quantities are copied as drawn, not
118
+ summed.
119
+
120
+ Files: `model.safetensors` (bf16), `config.json`, `generation_config.json`, `processor_config.json`, tokenizer files,
121
+ `chat_template.jinja`, the base model's `*_paddleocr_vl.py` code files (unchanged, Apache-2.0, PaddlePaddle Authors),
122
+ `usage.py`, `eval_summary.json`, `SHA256SUMS`, `LICENSE`, `NOTICE`.
123
+
124
+ ## Licence
125
+
126
+ Apache-2.0 for the weights and `usage.py`. The base model and its code files are Apache-2.0, Copyright PaddlePaddle
127
+ Authors; see NOTICE.
SHA256SUMS ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 2f27812dab7f333e471884e0c803d807f11953d5453140dfb1aaba234f872bc8 chat_template.jinja
2
+ 9c0379f6cc68c8ad950bd76716c40e3cbb4e502f96fe7dcc591e147173c92262 config.json
3
+ 753dd93654c3a9c8c85a3eaee1e3092dd12591b0f2dce0305e1abfb7a41ff160 configuration_paddleocr_vl.py
4
+ 7f0b73aa1bbffa1f9572ca7499aaec6bfc96f6b5beef680a472b4f4a00d0adc8 eval_summary.json
5
+ d34a3c0ceac7541c8d89c3eafb685adf5e3f5c63f27fbe8f7bb03bf1cfe1fa6e generation_config.json
6
+ a4fa521b9cb16e207f94b7f2d16427771776dfc634420d319fc4916ee58049ec image_processing_paddleocr_vl.py
7
+ b8c4d7deccd236af023af1c88c4d4e8f0fc2f41914e0fb23a3ec9678fb5a8456 LICENSE
8
+ c5013dff57ca8b87dc1de64d0fd839a44313de09d230a4fb2d08289d2cad5111 modeling_paddleocr_vl.py
9
+ d1b2613b26009260e1f98cf691f14da9f9051ad3880a172cc428d1f7f05934e7 model.safetensors
10
+ cafb13672ad15bcadb55b77f76a12b73c5842fd6f985e6d6e9525ab9a06d9041 NOTICE
11
+ e29cb1e5f275f2bd3ce051bd5c9983a33894e693b2823a0e13d4c07c8c4f9e13 processing_paddleocr_vl.py
12
+ fcf182714867cf2a401d644037a8b24678ff0780da5e14783bcc1f1b1c0ed21e processor_config.json
13
+ 8361f04a4e022dd4fa232b796fa40ccb7c81627df7c9d06c130b5b9dfe7a0b7f README.md
14
+ 0daa4d084c1b7a7343d1c5fbad2e933ca1f63c6858d8db0d72955e9e0bb93356 tokenizer_config.json
15
+ c8a215a59183d0d0781adc33bacd3ce6162716f7fd568fb30234a74d69803a7d tokenizer.json
16
+ f5ad20e35f9a323aacc3a76409f346770e96aa3cd9f7cb702066fd38e436e5f9 usage.py
chat_template.jinja ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if not add_generation_prompt is defined -%}
2
+ {%- set add_generation_prompt = true -%}
3
+ {%- endif -%}
4
+ {%- if not cls_token is defined -%}
5
+ {%- set cls_token = "<|begin_of_sentence|>" -%}
6
+ {%- endif -%}
7
+ {%- if not eos_token is defined -%}
8
+ {%- set eos_token = "</s>" -%}
9
+ {%- endif -%}
10
+ {{- cls_token -}}
11
+ {%- for message in messages -%}
12
+ {%- if message["role"] == "user" -%}
13
+ {{- "User: " -}}
14
+ {%- for content in message["content"] -%}
15
+ {%- if content["type"] == "image" -%}
16
+ {{ "<|IMAGE_START|><|IMAGE_PLACEHOLDER|><|IMAGE_END|>" }}
17
+ {%- endif -%}
18
+ {%- endfor -%}
19
+ {%- for content in message["content"] -%}
20
+ {%- if content["type"] == "text" -%}
21
+ {{ content["text"] }}
22
+ {%- endif -%}
23
+ {%- endfor -%}
24
+ {{ "\n" -}}
25
+ {%- elif message["role"] == "assistant" -%}
26
+ {{- "Assistant:\n" -}}
27
+ {%- for content in message["content"] -%}
28
+ {%- if content["type"] == "text" -%}
29
+ {{ content["text"] }}
30
+ {%- endif -%}
31
+ {%- endfor -%}
32
+ {{ eos_token -}}
33
+ {%- elif message["role"] == "system" -%}
34
+ {%- for content in message["content"] -%}
35
+ {%- if content["type"] == "text" -%}
36
+ {{ content["text"] + "\n" }}
37
+ {%- endif -%}
38
+ {%- endfor -%}
39
+ {%- endif -%}
40
+ {%- endfor -%}
41
+ {%- if add_generation_prompt -%}
42
+ {{- "Assistant:\n" -}}
43
+ {%- endif -%}
config.json ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "PaddleOCRVLForConditionalGeneration"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.0,
6
+ "auto_map": {
7
+ "AutoConfig": "configuration_paddleocr_vl.PaddleOCRVLConfig",
8
+ "AutoModel": "modeling_paddleocr_vl.PaddleOCRVLForConditionalGeneration",
9
+ "AutoModelForCausalLM": "modeling_paddleocr_vl.PaddleOCRVLForConditionalGeneration"
10
+ },
11
+ "compression_ratio": 1.0,
12
+ "dtype": "float32",
13
+ "hidden_dropout_prob": 0.0,
14
+ "ignored_index": -100,
15
+ "image_token_id": 100295,
16
+ "max_sequence_length": null,
17
+ "model_type": "paddleocr_vl",
18
+ "rope_is_neox_style": true,
19
+ "sliding_window": null,
20
+ "text_config": {
21
+ "bos_token_id": 1,
22
+ "dtype": "float32",
23
+ "eos_token_id": 2,
24
+ "head_dim": 128,
25
+ "hidden_act": "silu",
26
+ "hidden_size": 1024,
27
+ "initializer_range": 0.02,
28
+ "intermediate_size": 3072,
29
+ "max_position_embeddings": 131072,
30
+ "model_type": "paddleocr_vl_text",
31
+ "num_attention_heads": 16,
32
+ "num_hidden_layers": 18,
33
+ "num_key_value_heads": 2,
34
+ "pad_token_id": 0,
35
+ "rms_norm_eps": 1e-05,
36
+ "rope_parameters": {
37
+ "mrope_section": [
38
+ 16,
39
+ 24,
40
+ 24
41
+ ],
42
+ "rope_theta": 500000,
43
+ "rope_type": "default",
44
+ "type": "default"
45
+ },
46
+ "tie_word_embeddings": true,
47
+ "use_bias": false,
48
+ "use_cache": false,
49
+ "vocab_size": 103424
50
+ },
51
+ "tie_word_embeddings": true,
52
+ "transformers_version": "5.17.0",
53
+ "use_3d_rope": true,
54
+ "use_cache": true,
55
+ "use_flash_attention": false,
56
+ "video_token_id": 101307,
57
+ "vision_config": {
58
+ "architectures": [
59
+ "PaddleOCRVisionModel"
60
+ ],
61
+ "attention_dropout": 0.0,
62
+ "auto_map": {
63
+ "AutoConfig": "configuration_paddleocr_vl.PaddleOCRVLConfig",
64
+ "AutoModel": "modeling_paddleocr_vl.PaddleOCRVisionModel"
65
+ },
66
+ "dtype": "float32",
67
+ "hidden_act": "gelu_pytorch_tanh",
68
+ "hidden_size": 1152,
69
+ "image_size": 384,
70
+ "intermediate_size": 4304,
71
+ "layer_norm_eps": 1e-06,
72
+ "model_type": "paddleocr_vl_vision",
73
+ "num_attention_heads": 16,
74
+ "num_channels": 3,
75
+ "num_hidden_layers": 27,
76
+ "pad_token_id": 0,
77
+ "patch_size": 14,
78
+ "rope_parameters": {
79
+ "rope_theta": 10000.0,
80
+ "rope_type": "axial"
81
+ },
82
+ "spatial_merge_size": 2,
83
+ "temporal_patch_size": 2,
84
+ "tokens_per_second": 2
85
+ },
86
+ "vision_end_token_id": 101306,
87
+ "vision_start_token_id": 101305,
88
+ "weight_share_add_bias": true
89
+ }
configuration_paddleocr_vl.py ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright (c) 2025 PaddlePaddle Authors. All Rights Reserved.
2
+ #
3
+ # Licensed under the Apache License, Version 2.0 (the "License");
4
+ # you may not use this file except in compliance with the License.
5
+ # You may obtain a copy of the License at
6
+ #
7
+ # http://www.apache.org/licenses/LICENSE-2.0
8
+ #
9
+ # Unless required by applicable law or agreed to in writing, software
10
+ # distributed under the License is distributed on an "AS IS" BASIS,
11
+ # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12
+ # See the License for the specific language governing permissions and
13
+ # limitations under the License.
14
+
15
+ from transformers.configuration_utils import PretrainedConfig
16
+ from transformers.modeling_rope_utils import rope_config_validation
17
+
18
+ class PaddleOCRVisionConfig(PretrainedConfig):
19
+ model_type = "paddleocr_vl"
20
+ base_config_key = "vision_config"
21
+
22
+ def __init__(
23
+ self,
24
+ hidden_size=768,
25
+ intermediate_size=3072,
26
+ num_hidden_layers=12,
27
+ num_attention_heads=12,
28
+ num_channels=3,
29
+ image_size=224,
30
+ patch_size=14,
31
+ hidden_act="gelu_pytorch_tanh",
32
+ layer_norm_eps=1e-6,
33
+ attention_dropout=0.0,
34
+ spatial_merge_size=2,
35
+ temporal_patch_size=2,
36
+ tokens_per_second=2,
37
+ **kwargs,
38
+ ):
39
+ super().__init__(**kwargs)
40
+
41
+ self.hidden_size = hidden_size
42
+ self.intermediate_size = intermediate_size
43
+ self.num_hidden_layers = num_hidden_layers
44
+ self.num_attention_heads = num_attention_heads
45
+ self.num_channels = num_channels
46
+ self.patch_size = patch_size
47
+ self.image_size = image_size
48
+ self.attention_dropout = attention_dropout
49
+ self.layer_norm_eps = layer_norm_eps
50
+ self.hidden_act = hidden_act
51
+ self.spatial_merge_size = spatial_merge_size
52
+ self.temporal_patch_size = temporal_patch_size
53
+ self.tokens_per_second = tokens_per_second
54
+
55
+
56
+
57
+ class PaddleOCRVLConfig(PretrainedConfig):
58
+ """
59
+ Configuration class.
60
+
61
+ This class stores the configuration of an Ernie model, defining the model architecture.
62
+ It inherits from PretrainedConfig and can be used to control model outputs.
63
+ """
64
+
65
+ model_type = "paddleocr_vl"
66
+ keys_to_ignore_at_inference = ["past_key_values"]
67
+ sub_configs = {"vision_config": PaddleOCRVisionConfig}
68
+
69
+ # Default tensor parallel plan for base model `Qwen3`
70
+ base_model_tp_plan = {
71
+ "layers.*.self_attn.q_proj": "colwise",
72
+ "layers.*.self_attn.k_proj": "colwise",
73
+ "layers.*.self_attn.v_proj": "colwise",
74
+ "layers.*.self_attn.o_proj": "rowwise",
75
+ "layers.*.mlp.gate_proj": "colwise",
76
+ "layers.*.mlp.up_proj": "colwise",
77
+ "layers.*.mlp.down_proj": "rowwise",
78
+ }
79
+ base_model_pp_plan = {
80
+ "embed_tokens": (["input_ids"], ["inputs_embeds"]),
81
+ "layers": (["hidden_states", "attention_mask"], ["hidden_states"]),
82
+ "norm": (["hidden_states"], ["hidden_states"]),
83
+ }
84
+
85
+ def __init__(
86
+ self,
87
+ vocab_size=32000,
88
+ hidden_size=768,
89
+ intermediate_size=11008,
90
+ max_position_embeddings=32768,
91
+ num_hidden_layers=2,
92
+ num_attention_heads=2,
93
+ image_token_id=101304,
94
+ video_token_id=101305,
95
+ vision_start_token_id=101306,
96
+ rms_norm_eps=1e-6,
97
+ use_cache=False,
98
+ use_flash_attention=False,
99
+ pad_token_id=0,
100
+ bos_token_id=1,
101
+ eos_token_id=2,
102
+ head_dim=128,
103
+ hidden_act="silu",
104
+ use_bias=False,
105
+ rope_theta=10000,
106
+ weight_share_add_bias=True,
107
+ ignored_index=-100,
108
+ attention_probs_dropout_prob=0.0,
109
+ hidden_dropout_prob=0.0,
110
+ compression_ratio: float = 1.0,
111
+ num_key_value_heads=None,
112
+ max_sequence_length=None,
113
+ tie_word_embeddings=False,
114
+ vision_config=None,
115
+ rope_scaling=None,
116
+ **kwargs,
117
+ ):
118
+ """
119
+ Initialize configuration with default or specified parameters.
120
+
121
+ Args:
122
+ vocab_size (int): Size of the vocabulary (number of unique tokens)
123
+ hidden_size (int): Dimensionality of the encoder layers and the pooler layer
124
+ intermediate_size (int): Dimensionality of the "intermediate" (feed-forward) layer
125
+ max_position_embeddings (int): Maximum sequence length the model can handle
126
+ num_hidden_layers (int): Number of hidden layers in the Transformer encoder
127
+ num_attention_heads (int): Number of attention heads for each attention layer
128
+ rms_norm_eps (float): The epsilon used by the RMS normalization layers
129
+ use_cache (bool): Whether to use caching for faster generation (decoding)
130
+ use_flash_attention (bool): Whether to use FlashAttention for optimized attention computation
131
+ pad_token_id (int): Token ID used for padding sequences
132
+ bos_token_id (int): Token ID used for beginning-of-sequence
133
+ eos_token_id (int): Token ID used for end-of-sequence
134
+ use_bias (bool): Whether to use bias terms in linear layers
135
+ rope_theta (float): The base period of the RoPE embeddings
136
+ weight_share_add_bias (bool): Whether to share bias weights in certain layers
137
+ ignored_index (int): Target value that is ignored during loss computation
138
+ attention_probs_dropout_prob (float): Dropout probability for attention weights
139
+ hidden_dropout_prob (float): Dropout probability for hidden layers
140
+ compression_ratio (float): Ratio for KV cache compression (1.0 = no compression)
141
+ num_key_value_heads (int): Number of key/value heads (for Grouped Query Attention)
142
+ max_sequence_length (int): Maximum sequence length for positional embeddings
143
+ **kwargs: Additional keyword arguments passed to parent class
144
+ """
145
+
146
+ # Set default for tied embeddings if not specified.
147
+ super().__init__(
148
+ pad_token_id=pad_token_id,
149
+ bos_token_id=bos_token_id,
150
+ eos_token_id=eos_token_id,
151
+ **kwargs,
152
+ )
153
+ if isinstance(vision_config, dict):
154
+ self.vision_config = self.sub_configs["vision_config"](**vision_config)
155
+ elif vision_config is None:
156
+ self.vision_config = self.sub_configs["vision_config"]()
157
+ self.vocab_size = vocab_size
158
+ self.hidden_size = hidden_size
159
+ self.intermediate_size = intermediate_size
160
+ self.max_position_embeddings = max_position_embeddings
161
+ self.num_hidden_layers = num_hidden_layers
162
+ self.num_attention_heads = num_attention_heads
163
+ self.rms_norm_eps = rms_norm_eps
164
+ self.use_cache = use_cache
165
+ self.use_flash_attention = use_flash_attention
166
+ self.pad_token_id = pad_token_id
167
+ self.bos_token_id = bos_token_id
168
+ self.eos_token_id = eos_token_id
169
+ self.image_token_id = image_token_id
170
+ self.video_token_id = video_token_id
171
+ self.vision_start_token_id = vision_start_token_id
172
+ self.head_dim = head_dim
173
+ self.hidden_act=hidden_act
174
+ self.sliding_window = None
175
+ self.hidden_size = hidden_size
176
+ self.use_bias = use_bias
177
+ self.weight_share_add_bias = weight_share_add_bias
178
+ self.rope_theta = rope_theta
179
+ self.ignored_index = ignored_index
180
+ self.attention_probs_dropout_prob = attention_probs_dropout_prob
181
+ self.hidden_dropout_prob = hidden_dropout_prob
182
+ self.compression_ratio = compression_ratio
183
+ self.num_key_value_heads = num_key_value_heads
184
+ self.max_sequence_length = max_sequence_length
185
+ self.rope_scaling = rope_scaling
186
+ if self.rope_scaling is not None and "type" in self.rope_scaling:
187
+ if self.rope_scaling["type"] == "mrope":
188
+ self.rope_scaling["type"] = "default"
189
+ self.rope_scaling["rope_type"] = self.rope_scaling["type"]
190
+ rope_config_validation(self, ignore_keys={"mrope_section"})
191
+ super().__init__(tie_word_embeddings=tie_word_embeddings, **kwargs)
eval_summary.json ADDED
@@ -0,0 +1,330 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "decosa-oneline-reader-paddleocr-vl",
3
+ "date": "2026-09-28",
4
+ "base": "PaddlePaddle/PaddleOCR-VL-1.6@c5630abae1d940eafe0697512a0325494b02ab42",
5
+ "metric": "all_fields_right per drawing (9 fields: poi_kv, plant_limit_mw, models, qty, kva, bess_mw, bess_mwh, transformer_mva, has_disconnect); per-field right/n",
6
+ "setup": "test sets of 60 synthetic drawings each from the training generator family, seeds disjoint from training; bf16, greedy, max_new_tokens 900, batch 6; a re-run on a second machine gave identical results",
7
+ "caveat": "no real utility one-line diagram has been scored",
8
+ "sets": {
9
+ "hard2": {
10
+ "desc": "poor scans (fresh), the headline set",
11
+ "this_model": {
12
+ "all_fields_right": 59,
13
+ "cases": 60,
14
+ "fields": {
15
+ "poi_kv": [
16
+ 60,
17
+ 60
18
+ ],
19
+ "plant_limit_mw": [
20
+ 60,
21
+ 60
22
+ ],
23
+ "models": [
24
+ 60,
25
+ 60
26
+ ],
27
+ "qty": [
28
+ 91,
29
+ 91
30
+ ],
31
+ "kva": [
32
+ 91,
33
+ 91
34
+ ],
35
+ "bess_mw": [
36
+ 49,
37
+ 49
38
+ ],
39
+ "bess_mwh": [
40
+ 48,
41
+ 49
42
+ ],
43
+ "transformer_mva": [
44
+ 60,
45
+ 60
46
+ ],
47
+ "has_disconnect": [
48
+ 60,
49
+ 60
50
+ ]
51
+ }
52
+ },
53
+ "seconds_total": 69.3,
54
+ "qwen3.8-27b_vision": {
55
+ "all_fields_right": 25,
56
+ "cases": 60,
57
+ "fields": {
58
+ "poi_kv": [
59
+ 60,
60
+ 60
61
+ ],
62
+ "plant_limit_mw": [
63
+ 59,
64
+ 60
65
+ ],
66
+ "models": [
67
+ 38,
68
+ 60
69
+ ],
70
+ "qty": [
71
+ 57,
72
+ 91
73
+ ],
74
+ "kva": [
75
+ 67,
76
+ 91
77
+ ],
78
+ "bess_mw": [
79
+ 48,
80
+ 49
81
+ ],
82
+ "bess_mwh": [
83
+ 48,
84
+ 49
85
+ ],
86
+ "transformer_mva": [
87
+ 56,
88
+ 60
89
+ ],
90
+ "has_disconnect": [
91
+ 60,
92
+ 60
93
+ ]
94
+ }
95
+ },
96
+ "base_first12": {
97
+ "all_fields_right": 0,
98
+ "cases": 12,
99
+ "parsed": 0
100
+ }
101
+ },
102
+ "hard": {
103
+ "desc": "the clean set degraded",
104
+ "this_model": {
105
+ "all_fields_right": 59,
106
+ "cases": 60,
107
+ "fields": {
108
+ "poi_kv": [
109
+ 60,
110
+ 60
111
+ ],
112
+ "plant_limit_mw": [
113
+ 60,
114
+ 60
115
+ ],
116
+ "models": [
117
+ 60,
118
+ 60
119
+ ],
120
+ "qty": [
121
+ 88,
122
+ 88
123
+ ],
124
+ "kva": [
125
+ 88,
126
+ 88
127
+ ],
128
+ "bess_mw": [
129
+ 48,
130
+ 48
131
+ ],
132
+ "bess_mwh": [
133
+ 47,
134
+ 48
135
+ ],
136
+ "transformer_mva": [
137
+ 60,
138
+ 60
139
+ ],
140
+ "has_disconnect": [
141
+ 60,
142
+ 60
143
+ ]
144
+ }
145
+ },
146
+ "seconds_total": 137.8,
147
+ "qwen3.8-27b_vision": {
148
+ "all_fields_right": 16,
149
+ "cases": 60,
150
+ "fields": {
151
+ "poi_kv": [
152
+ 60,
153
+ 60
154
+ ],
155
+ "plant_limit_mw": [
156
+ 53,
157
+ 60
158
+ ],
159
+ "models": [
160
+ 34,
161
+ 60
162
+ ],
163
+ "qty": [
164
+ 43,
165
+ 88
166
+ ],
167
+ "kva": [
168
+ 57,
169
+ 88
170
+ ],
171
+ "bess_mw": [
172
+ 34,
173
+ 48
174
+ ],
175
+ "bess_mwh": [
176
+ 35,
177
+ 48
178
+ ],
179
+ "transformer_mva": [
180
+ 46,
181
+ 60
182
+ ],
183
+ "has_disconnect": [
184
+ 60,
185
+ 60
186
+ ]
187
+ }
188
+ }
189
+ },
190
+ "crisp": {
191
+ "desc": "clean drawings (regression check)",
192
+ "this_model": {
193
+ "all_fields_right": 60,
194
+ "cases": 60,
195
+ "fields": {
196
+ "poi_kv": [
197
+ 60,
198
+ 60
199
+ ],
200
+ "plant_limit_mw": [
201
+ 60,
202
+ 60
203
+ ],
204
+ "models": [
205
+ 60,
206
+ 60
207
+ ],
208
+ "qty": [
209
+ 88,
210
+ 88
211
+ ],
212
+ "kva": [
213
+ 88,
214
+ 88
215
+ ],
216
+ "bess_mw": [
217
+ 48,
218
+ 48
219
+ ],
220
+ "bess_mwh": [
221
+ 48,
222
+ 48
223
+ ],
224
+ "transformer_mva": [
225
+ 60,
226
+ 60
227
+ ],
228
+ "has_disconnect": [
229
+ 60,
230
+ 60
231
+ ]
232
+ }
233
+ },
234
+ "seconds_total": 120.3,
235
+ "qwen3.8-27b_vision": {
236
+ "all_fields_right": 59,
237
+ "cases": 60,
238
+ "fields": {
239
+ "poi_kv": [
240
+ 60,
241
+ 60
242
+ ],
243
+ "plant_limit_mw": [
244
+ 60,
245
+ 60
246
+ ],
247
+ "models": [
248
+ 60,
249
+ 60
250
+ ],
251
+ "qty": [
252
+ 87,
253
+ 88
254
+ ],
255
+ "kva": [
256
+ 88,
257
+ 88
258
+ ],
259
+ "bess_mw": [
260
+ 48,
261
+ 48
262
+ ],
263
+ "bess_mwh": [
264
+ 48,
265
+ 48
266
+ ],
267
+ "transformer_mva": [
268
+ 60,
269
+ 60
270
+ ],
271
+ "has_disconnect": [
272
+ 60,
273
+ 60
274
+ ]
275
+ }
276
+ },
277
+ "base_first12": {
278
+ "all_fields_right": 0,
279
+ "cases": 12,
280
+ "parsed": 0
281
+ }
282
+ },
283
+ "novel": {
284
+ "desc": "poor scans with random inverter model names and ratings never used in training",
285
+ "this_model": {
286
+ "all_fields_right": 50,
287
+ "cases": 60,
288
+ "fields": {
289
+ "poi_kv": [
290
+ 60,
291
+ 60
292
+ ],
293
+ "plant_limit_mw": [
294
+ 60,
295
+ 60
296
+ ],
297
+ "models": [
298
+ 53,
299
+ 60
300
+ ],
301
+ "qty": [
302
+ 89,
303
+ 98
304
+ ],
305
+ "kva": [
306
+ 91,
307
+ 98
308
+ ],
309
+ "bess_mw": [
310
+ 46,
311
+ 47
312
+ ],
313
+ "bess_mwh": [
314
+ 47,
315
+ 47
316
+ ],
317
+ "transformer_mva": [
318
+ 60,
319
+ 60
320
+ ],
321
+ "has_disconnect": [
322
+ 60,
323
+ 60
324
+ ]
325
+ }
326
+ },
327
+ "seconds_total": 105.8
328
+ }
329
+ }
330
+ }
generation_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "eos_token_id": 2,
4
+ "pad_token_id": 0,
5
+ "transformers_version": "5.17.0",
6
+ "use_cache": true
7
+ }
image_processing_paddleocr_vl.py ADDED
@@ -0,0 +1,570 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright (c) 2025 PaddlePaddle Authors. All Rights Reserved.
2
+ #
3
+ # Licensed under the Apache License, Version 2.0 (the "License");
4
+ # you may not use this file except in compliance with the License.
5
+ # You may obtain a copy of the License at
6
+ #
7
+ # http://www.apache.org/licenses/LICENSE-2.0
8
+ #
9
+ # Unless required by applicable law or agreed to in writing, software
10
+ # distributed under the License is distributed on an "AS IS" BASIS,
11
+ # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12
+ # See the License for the specific language governing permissions and
13
+ # limitations under the License.
14
+
15
+ """Image processor class for PaddleOCR-VL."""
16
+
17
+ import math
18
+ from typing import Dict, List, Optional, Union
19
+
20
+ import numpy as np
21
+ import torch
22
+ from transformers.image_processing_utils import BaseImageProcessor, BatchFeature
23
+ from torchvision.transforms import functional as TF
24
+ from transformers.image_transforms import (
25
+ convert_to_rgb,
26
+ resize,
27
+ to_channel_dimension_format,
28
+ )
29
+ from transformers.image_utils import (
30
+ OPENAI_CLIP_MEAN,
31
+ OPENAI_CLIP_STD,
32
+ ChannelDimension,
33
+ PILImageResampling,
34
+ get_image_size,
35
+ infer_channel_dimension_format,
36
+ is_scaled_image,
37
+ is_valid_image,
38
+ make_list_of_images,
39
+ to_numpy_array,
40
+ valid_images,
41
+ validate_preprocess_arguments,
42
+ )
43
+ from transformers.utils import TensorType, is_vision_available, logging
44
+
45
+
46
+ logger = logging.get_logger(__name__)
47
+
48
+
49
+ if is_vision_available():
50
+ from PIL import Image
51
+
52
+ ImageInput = Union[
53
+ "PIL.Image.Image",
54
+ np.ndarray,
55
+ "torch.Tensor",
56
+ List["PIL.Image.Image"],
57
+ List[np.ndarray],
58
+ List["torch.Tensor"],
59
+ ] # noqa
60
+
61
+
62
+ VideoInput = Union[
63
+ List["PIL.Image.Image"],
64
+ "np.ndarray",
65
+ "torch.Tensor",
66
+ List["np.ndarray"],
67
+ List["torch.Tensor"],
68
+ List[List["PIL.Image.Image"]],
69
+ List[List["np.ndarrray"]],
70
+ List[List["torch.Tensor"]],
71
+ ] # noqa
72
+
73
+
74
+ def make_batched_images(images) -> List[List[ImageInput]]:
75
+ """
76
+ Accepts images in list or nested list format, and makes a list of images for preprocessing.
77
+
78
+ Args:
79
+ images (`Union[List[List[ImageInput]], List[ImageInput], ImageInput]`):
80
+ The input image.
81
+
82
+ Returns:
83
+ list: A list of images.
84
+ """
85
+ if (
86
+ isinstance(images, (list, tuple))
87
+ and isinstance(images[0], (list, tuple))
88
+ and is_valid_image(images[0][0])
89
+ ):
90
+ return [img for img_list in images for img in img_list]
91
+
92
+ elif isinstance(images, (list, tuple)) and is_valid_image(images[0]):
93
+ return images
94
+
95
+ elif is_valid_image(images):
96
+ return [images]
97
+
98
+ raise ValueError(f"Could not make batched images from {images}")
99
+
100
+
101
+ def adjust_size(size, patch_size):
102
+ num_patches = size // patch_size
103
+ if num_patches % 2 != 0: # 如果是奇数,减1
104
+ num_patches -= 1
105
+ return num_patches * patch_size
106
+
107
+
108
+ def make_batched_videos(videos) -> List[VideoInput]:
109
+ if (
110
+ isinstance(videos, (list, tuple))
111
+ and isinstance(videos[0], (list, tuple))
112
+ and is_valid_image(videos[0][0])
113
+ ):
114
+ return videos
115
+
116
+ elif isinstance(videos, (list, tuple)) and is_valid_image(videos[0]):
117
+ if isinstance(videos[0], Image.Image):
118
+ return [videos]
119
+ elif len(videos[0].shape) == 4:
120
+ return [list(video) for video in videos]
121
+
122
+ elif is_valid_image(videos) and len(videos.shape) == 4:
123
+ return [list(videos)]
124
+
125
+ raise ValueError(f"Could not make batched video from {videos}")
126
+
127
+
128
+ def smart_resize(
129
+ height: int,
130
+ width: int,
131
+ factor: int = 28,
132
+ min_pixels: int = 28 * 28 * 130,
133
+ max_pixels: int = 28 * 28 * 1280,
134
+ ):
135
+ """Rescales the image so that the following conditions are met:
136
+
137
+ 1. Both dimensions (height and width) are divisible by 'factor'.
138
+
139
+ 2. The total number of pixels is within the range ['min_pixels', 'max_pixels'].
140
+
141
+ 3. The aspect ratio of the image is maintained as closely as possible.
142
+
143
+ """
144
+ # if height < factor or width < factor:
145
+ # raise ValueError(f"height:{height} or width:{width} must be larger than factor:{factor}")
146
+ # if int(height < factor//4) + int(width < factor//4):
147
+ # raise ValueError(f"height:{height} or width:{width} must be larger than factor:{factor//4}")
148
+
149
+ if height < factor:
150
+ print(f"smart_resize: height={height} < factor={factor}, reset height=factor")
151
+ width = round((width * factor) / height)
152
+ height = factor
153
+
154
+ if width < factor:
155
+ print(f"smart_resize: width={width} < factor={factor}, reset width=factor")
156
+ height = round((height * factor) / width)
157
+ width = factor
158
+
159
+ if max(height, width) / min(height, width) > 200:
160
+ raise ValueError(
161
+ f"absolute aspect ratio must be smaller than 200, got {max(height, width) / min(height, width)}"
162
+ )
163
+ h_bar = round(height / factor) * factor
164
+ w_bar = round(width / factor) * factor
165
+ if h_bar * w_bar > max_pixels:
166
+ beta = math.sqrt((height * width) / max_pixels)
167
+ h_bar = math.floor(height / beta / factor) * factor
168
+ w_bar = math.floor(width / beta / factor) * factor
169
+ elif h_bar * w_bar < min_pixels:
170
+ beta = math.sqrt(min_pixels / (height * width))
171
+ h_bar = math.ceil(height * beta / factor) * factor
172
+ w_bar = math.ceil(width * beta / factor) * factor
173
+ return h_bar, w_bar
174
+
175
+
176
+ class PaddleOCRVLImageProcessor(BaseImageProcessor):
177
+ r"""
178
+ Constructs a Siglip image processor that dynamically resizes images based on the original images.
179
+
180
+ Args:
181
+ do_resize (`bool`, *optional*, defaults to `True`):
182
+ Whether to resize the image's (height, width) dimensions.
183
+ resample (`PILImageResampling`, *optional*, defaults to `Resampling.BICUBIC`):
184
+ Resampling filter to use when resizing the image.
185
+ do_rescale (`bool`, *optional*, defaults to `True`):
186
+ Whether to rescale the image by the specified scale `rescale_factor`.
187
+ rescale_factor (`int` or `float`, *optional*, defaults to `1/255`):
188
+ Scale factor to use if rescaling the image.
189
+ do_normalize (`bool`, *optional*, defaults to `True`):
190
+ Whether to normalize the image.
191
+ image_mean (`float` or `List[float]`, *optional*, defaults to `[0.48145466, 0.4578275, 0.40821073]`):
192
+ Mean to use if normalizing the image. This is a float or list of floats for each channel in the image.
193
+ image_std (`float` or `List[float]`, *optional*, defaults to `[0.26862954, 0.26130258, 0.27577711]`):
194
+ Standard deviation to use if normalizing the image. This is a float or list of floats for each channel in the image.
195
+ do_convert_rgb (`bool`, *optional*, defaults to `True`):
196
+ Whether to convert the image to RGB.
197
+ min_pixels (`int`, *optional*, defaults to `28 * 28 * 130`):
198
+ The min pixels of the image to resize the image.
199
+ max_pixels (`int`, *optional*, defaults to `28 * 28 * 1670`):
200
+ The max pixels of the image to resize the image.
201
+ patch_size (`int`, *optional*, defaults to 14):
202
+ The spacial patch size of the vision encoder.
203
+ temporal_patch_size (`int`, *optional*, defaults to 2):
204
+ The temporal patch size of the vision encoder.
205
+ merge_size (`int`, *optional*, defaults to 2):
206
+ The merge size of the vision encoder to llm encoder.
207
+ """
208
+
209
+ model_input_names = [
210
+ "pixel_values",
211
+ "image_grid_thw",
212
+ "pixel_values_videos",
213
+ "video_grid_thw",
214
+ ]
215
+
216
+ def __init__(
217
+ self,
218
+ do_resize: bool = True,
219
+ resample: PILImageResampling = PILImageResampling.BICUBIC,
220
+ do_rescale: bool = True,
221
+ rescale_factor: Union[int, float] = 1 / 255,
222
+ do_normalize: bool = True,
223
+ image_mean: Optional[Union[float, List[float]]] = None,
224
+ image_std: Optional[Union[float, List[float]]] = None,
225
+ do_convert_rgb: bool = True,
226
+ min_pixels: int = 28 * 28 * 130,
227
+ max_pixels: int = 28 * 28 * 1280,
228
+ patch_size: int = 14,
229
+ temporal_patch_size: int = 1,
230
+ merge_size: int = 2,
231
+ **kwargs,
232
+ ) -> None:
233
+ super().__init__(**kwargs)
234
+ self.do_resize = do_resize
235
+ self.resample = resample
236
+ self.do_rescale = do_rescale
237
+ self.rescale_factor = rescale_factor
238
+ self.do_normalize = do_normalize
239
+ self.image_mean = image_mean if image_mean is not None else OPENAI_CLIP_MEAN
240
+ self.image_std = image_std if image_std is not None else OPENAI_CLIP_STD
241
+ self.min_pixels = min_pixels
242
+ self.max_pixels = max_pixels
243
+ self.patch_size = patch_size
244
+ self.temporal_patch_size = temporal_patch_size
245
+ self.merge_size = merge_size
246
+ self.size = {"min_pixels": min_pixels, "max_pixels": max_pixels} # not used
247
+ self.do_convert_rgb = do_convert_rgb
248
+
249
+ def mvit_rescale(self, image: Image.Image, merge_size: int = 2) -> Image.Image:
250
+ try:
251
+ w, h = image.size
252
+ except:
253
+ raise ValueError(str((type(image), image)))
254
+ patch_size = self.patch_size
255
+
256
+ if (w // patch_size) * (h // patch_size) > self.in_token_limit:
257
+ scale = math.sqrt(
258
+ self.in_token_limit / ((w // patch_size) * (h // patch_size))
259
+ )
260
+ new_w, new_h = int(w * scale), int(h * scale)
261
+
262
+ image = image.resize((new_w, new_h), Image.Resampling.BICUBIC)
263
+ if self.pad_input:
264
+ new_w, new_h = image.size
265
+ pad_size_h = merge_size * patch_size
266
+ pad_size_w = merge_size * patch_size
267
+
268
+ pad_h = (pad_size_h - new_h % pad_size_h) % pad_size_h
269
+ pad_w = (pad_size_w - new_w % pad_size_w) % pad_size_w
270
+
271
+ image = TF.pad(image, (0, 0, pad_w, pad_h))
272
+ else:
273
+ new_w, new_h = image.size
274
+ new_w = new_w - new_w % patch_size
275
+ new_h = new_h - new_h % patch_size
276
+
277
+ new_w = adjust_size(new_w, patch_size)
278
+ new_h = adjust_size(new_h, patch_size)
279
+
280
+ image = TF.center_crop(image, (new_h, new_w))
281
+
282
+ w, h = image.size
283
+ if w // patch_size >= 512 or h // patch_size >= 512:
284
+ new_h = min(patch_size * 510, h)
285
+ new_w = min(patch_size * 510, w)
286
+ image = TF.center_crop(image, (new_h, new_w))
287
+ # raise ValueError("Exceed pos emb")
288
+ return image
289
+
290
+ def _preprocess(
291
+ self,
292
+ images: Union[ImageInput, VideoInput],
293
+ do_resize: bool = None,
294
+ resample: PILImageResampling = None,
295
+ do_rescale: bool = None,
296
+ rescale_factor: float = None,
297
+ do_normalize: bool = None,
298
+ image_mean: Optional[Union[float, List[float]]] = None,
299
+ image_std: Optional[Union[float, List[float]]] = None,
300
+ do_convert_rgb: bool = None,
301
+ data_format: Optional[ChannelDimension] = ChannelDimension.FIRST,
302
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
303
+ ):
304
+ """
305
+ Preprocess an image or batch of images. Copy of the `preprocess` method from `CLIPImageProcessor`.
306
+
307
+ Args:
308
+ images (`ImageInput`):
309
+ Image or batch of images to preprocess. Expects pixel values ranging from 0 to 255. If pixel values range from 0 to 1, set `do_rescale=False`.
310
+ vision_info (`List[Dict]`, *optional*):
311
+ Optional list of dictionaries containing additional information about vision inputs.
312
+ do_resize (`bool`, *optional*, defaults to `self.do_resize`):
313
+ Whether to resize the image.
314
+ resample (`PILImageResampling`, *optional*, defaults to `self.resample`):
315
+ Resampling filter to use if resizing the image. This can be one of the `PILImageResampling` enums.
316
+ do_rescale (`bool`, *optional*, defaults to `self.do_rescale`):
317
+ Whether to rescale the image.
318
+ rescale_factor (`float`, *optional*, defaults to `self.rescale_factor`):
319
+ Scale factor to use if rescaling the image.
320
+ do_normalize (`bool`, *optional*, defaults to `self.do_normalize`):
321
+ Whether to normalize the image.
322
+ image_mean (`float` or `List[float]`, *optional*, defaults to `self.image_mean`):
323
+ Mean to use if normalizing the image. Can be a float or a list of floats corresponding to the number of channels in the image.
324
+ image_std (`float` or `List[float]`, *optional*, defaults to `self.image_std`):
325
+ Standard deviation to use if normalizing the image. Can be a float or a list of floats corresponding to the number of channels in the image.
326
+ do_convert_rgb (`bool`, *optional*, defaults to `self.do_convert_rgb`):
327
+ Whether to convert the image to RGB.
328
+ data_format (`ChannelDimension`, *optional*, defaults to `ChannelDimension.FIRST`):
329
+ The channel dimension format for the output image. Can be one of:
330
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
331
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
332
+ - Unset: Use the channel dimension format of the input image.
333
+ input_data_format (`ChannelDimension` or `str`, *optional*):
334
+ The channel dimension format for the input image. Can be one of:
335
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
336
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
337
+ - `"none"` or `ChannelDimension.NONE`: image in (height, width) format. - `"none"` or `ChannelDimension.NONE`: image in (height, width) format.
338
+ """
339
+ images = make_list_of_images(images)
340
+
341
+ if input_data_format is None:
342
+ # We assume that all images have the same channel dimension format.
343
+ input_data_format = ChannelDimension.LAST if isinstance(images[0], Image.Image) else infer_channel_dimension_format(images[0])
344
+
345
+ if do_convert_rgb:
346
+ images = [convert_to_rgb(image) for image in images]
347
+
348
+ # All transformations expect numpy arrays.
349
+ images = [to_numpy_array(image) for image in images]
350
+
351
+ if is_scaled_image(images[0]) and do_rescale:
352
+ logger.warning_once(
353
+ "It looks like you are trying to rescale already rescaled images. If the input"
354
+ " images have pixel values between 0 and 1, set `do_rescale=False` to avoid rescaling them again."
355
+ )
356
+
357
+ height, width = get_image_size(images[0], channel_dim=input_data_format)
358
+ resized_height, resized_width = height, width
359
+ processed_images = []
360
+
361
+ for image in images:
362
+ if do_resize:
363
+ resized_height, resized_width = smart_resize(
364
+ height,
365
+ width,
366
+ factor=self.patch_size * self.merge_size,
367
+ min_pixels=self.min_pixels,
368
+ max_pixels=self.max_pixels,
369
+ )
370
+ image = resize(
371
+ image,
372
+ size=(resized_height, resized_width),
373
+ resample=resample,
374
+ input_data_format=input_data_format,
375
+ )
376
+
377
+ if do_rescale:
378
+ image = self.rescale(
379
+ image, scale=rescale_factor, input_data_format=input_data_format
380
+ )
381
+
382
+ if do_normalize:
383
+ image = self.normalize(
384
+ image=image,
385
+ mean=image_mean,
386
+ std=image_std,
387
+ input_data_format=input_data_format,
388
+ )
389
+ image = to_channel_dimension_format(
390
+ image, data_format, input_channel_dim=input_data_format
391
+ )
392
+ processed_images.append(image)
393
+
394
+ patches = np.array(processed_images)
395
+ if data_format == ChannelDimension.LAST:
396
+ patches = patches.transpose(0, 3, 1, 2)
397
+ if patches.shape[0] == 1:
398
+ patches = np.tile(patches, (self.temporal_patch_size, 1, 1, 1))
399
+ init_patches = patches
400
+ channel = patches.shape[1]
401
+ grid_t = patches.shape[0] // self.temporal_patch_size
402
+ grid_h, grid_w = (
403
+ resized_height // self.patch_size,
404
+ resized_width // self.patch_size,
405
+ )
406
+ patches = patches.reshape(
407
+ grid_t,
408
+ self.temporal_patch_size,
409
+ channel,
410
+ grid_h,
411
+ self.patch_size,
412
+ grid_w,
413
+ self.patch_size,
414
+ )
415
+ patches = patches.transpose(0, 3, 5, 2, 1, 4, 6)
416
+ assert self.temporal_patch_size == 1
417
+ flatten_patches = patches.reshape(
418
+ grid_t * grid_h * grid_w, channel, self.patch_size, self.patch_size
419
+ )
420
+ return flatten_patches, (grid_t, grid_h, grid_w)
421
+
422
+ def preprocess(
423
+ self,
424
+ images: ImageInput,
425
+ videos: VideoInput = None,
426
+ do_resize: bool = None,
427
+ size: Dict[str, int] = None,
428
+ resample: PILImageResampling = None,
429
+ do_rescale: bool = None,
430
+ rescale_factor: float = None,
431
+ do_normalize: bool = None,
432
+ image_mean: Optional[Union[float, List[float]]] = None,
433
+ image_std: Optional[Union[float, List[float]]] = None,
434
+ do_convert_rgb: bool = None,
435
+ return_tensors: Optional[Union[str, TensorType]] = None,
436
+ data_format: Optional[ChannelDimension] = ChannelDimension.FIRST,
437
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
438
+ ):
439
+ """
440
+ Args:
441
+ images (`ImageInput`):
442
+ Image to preprocess. Expects a single or batch of images with pixel values ranging from 0 to 255. If
443
+ passing in images with pixel values between 0 and 1, set `do_rescale=False`.
444
+ videos (`VideoInput`):
445
+ Video to preprocess. Expects a single or batch of videos with pixel values ranging from 0 to 255. If
446
+ passing in videos with pixel values between 0 and 1, set `do_rescale=False`.
447
+ do_resize (`bool`, *optional*, defaults to `self.do_resize`):
448
+ Whether to resize the image.
449
+ size (`Dict[str, int]`, *optional*, defaults to `self.size`):
450
+ Size of the image after resizing. Shortest edge of the image is resized to size["shortest_edge"], with
451
+ the longest edge resized to keep the input aspect ratio.
452
+ resample (`int`, *optional*, defaults to `self.resample`):
453
+ Resampling filter to use if resizing the image. This can be one of the enum `PILImageResampling`. Only
454
+ has an effect if `do_resize` is set to `True`.
455
+ do_rescale (`bool`, *optional*, defaults to `self.do_rescale`):
456
+ Whether to rescale the image.
457
+ rescale_factor (`float`, *optional*, defaults to `self.rescale_factor`):
458
+ Rescale factor to rescale the image by if `do_rescale` is set to `True`.
459
+ do_normalize (`bool`, *optional*, defaults to `self.do_normalize`):
460
+ Whether to normalize the image.
461
+ image_mean (`float` or `List[float]`, *optional*, defaults to `self.image_mean`):
462
+ Image mean to use for normalization. Only has an effect if `do_normalize` is set to `True`.
463
+ image_std (`float` or `List[float]`, *optional*, defaults to `self.image_std`):
464
+ Image standard deviation to use for normalization. Only has an effect if `do_normalize` is set to
465
+ `True`.
466
+ do_convert_rgb (`bool`, *optional*, defaults to `self.do_convert_rgb`):
467
+ Whether to convert the image to RGB.
468
+ return_tensors (`str` or `TensorType`, *optional*):
469
+ The type of tensors to return. Can be one of:
470
+ - Unset: Return a list of `np.ndarray`.
471
+ - `TensorType.TENSORFLOW` or `'tf'`: Return a batch of type `tf.Tensor`.
472
+ - `TensorType.PYTORCH` or `'pt'`: Return a batch of type `torch.Tensor`.
473
+ - `TensorType.NUMPY` or `'np'`: Return a batch of type `np.ndarray`.
474
+ - `TensorType.JAX` or `'jax'`: Return a batch of type `jax.numpy.ndarray`.
475
+ data_format (`ChannelDimension` or `str`, *optional*, defaults to `ChannelDimension.FIRST`):
476
+ The channel dimension format for the output image. Can be one of:
477
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
478
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
479
+ - Unset: Use the channel dimension format of the input image.
480
+ input_data_format (`ChannelDimension` or `str`, *optional*):
481
+ The channel dimension format for the input image. If unset, the channel dimension format is inferred
482
+ from the input image. Can be one of:
483
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
484
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
485
+ - `"none"` or `ChannelDimension.NONE`: image in (height, width) format.
486
+
487
+ """
488
+ do_resize = do_resize if do_resize is not None else self.do_resize
489
+ size = size if size is not None else self.size
490
+ resample = resample if resample is not None else self.resample
491
+ do_rescale = do_rescale if do_rescale is not None else self.do_rescale
492
+ rescale_factor = (
493
+ rescale_factor if rescale_factor is not None else self.rescale_factor
494
+ )
495
+ do_normalize = do_normalize if do_normalize is not None else self.do_normalize
496
+ image_mean = image_mean if image_mean is not None else self.image_mean
497
+ image_std = image_std if image_std is not None else self.image_std
498
+ do_convert_rgb = (
499
+ do_convert_rgb if do_convert_rgb is not None else self.do_convert_rgb
500
+ )
501
+
502
+ if images is not None:
503
+ images = make_batched_images(images)
504
+ if videos is not None:
505
+ videos = make_batched_videos(videos)
506
+
507
+ if images is not None and not valid_images(images):
508
+ raise ValueError(
509
+ "Invalid image type. Must be of type PIL.Image.Image, numpy.ndarray, "
510
+ "torch.Tensor, tf.Tensor or jax.ndarray."
511
+ )
512
+
513
+ validate_preprocess_arguments(
514
+ rescale_factor=rescale_factor,
515
+ do_normalize=do_normalize,
516
+ image_mean=image_mean,
517
+ image_std=image_std,
518
+ do_resize=do_resize,
519
+ size=size,
520
+ resample=resample,
521
+ )
522
+
523
+ if images is not None:
524
+ pixel_values, vision_grid_thws = [], []
525
+ for image in images:
526
+ patches, image_grid_thw = self._preprocess(
527
+ image,
528
+ do_resize=do_resize,
529
+ resample=resample,
530
+ do_rescale=do_rescale,
531
+ rescale_factor=rescale_factor,
532
+ do_normalize=do_normalize,
533
+ image_mean=image_mean,
534
+ image_std=image_std,
535
+ data_format=data_format,
536
+ do_convert_rgb=do_convert_rgb,
537
+ input_data_format=input_data_format,
538
+ )
539
+ pixel_values.extend(patches)
540
+ vision_grid_thws.append(image_grid_thw)
541
+ pixel_values = np.array(pixel_values)
542
+ vision_grid_thws = np.array(vision_grid_thws)
543
+ data = {"pixel_values": pixel_values, "image_grid_thw": vision_grid_thws}
544
+
545
+ if videos is not None:
546
+ pixel_values, vision_grid_thws = [], []
547
+ for images in videos:
548
+ patches, video_grid_thw = self._preprocess(
549
+ images,
550
+ do_resize=do_resize,
551
+ resample=resample,
552
+ do_rescale=do_rescale,
553
+ rescale_factor=rescale_factor,
554
+ do_normalize=do_normalize,
555
+ image_mean=image_mean,
556
+ image_std=image_std,
557
+ data_format=data_format,
558
+ do_convert_rgb=do_convert_rgb,
559
+ input_data_format=input_data_format,
560
+ )
561
+ pixel_values.extend(patches)
562
+ vision_grid_thws.append(video_grid_thw)
563
+ pixel_values = np.array(pixel_values)
564
+ vision_grid_thws = np.array(vision_grid_thws)
565
+ data = {
566
+ "pixel_values_videos": pixel_values,
567
+ "video_grid_thw": vision_grid_thws,
568
+ }
569
+
570
+ return BatchFeature(data=data, tensor_type=return_tensors)
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d1b2613b26009260e1f98cf691f14da9f9051ad3880a172cc428d1f7f05934e7
3
+ size 1811280336
modeling_paddleocr_vl.py ADDED
The diff for this file is too large to render. See raw diff
 
processing_paddleocr_vl.py ADDED
@@ -0,0 +1,293 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright (c) 2025 PaddlePaddle Authors. All Rights Reserved.
2
+ #
3
+ # Licensed under the Apache License, Version 2.0 (the "License");
4
+ # you may not use this file except in compliance with the License.
5
+ # You may obtain a copy of the License at
6
+ #
7
+ # http://www.apache.org/licenses/LICENSE-2.0
8
+ #
9
+ # Unless required by applicable law or agreed to in writing, software
10
+ # distributed under the License is distributed on an "AS IS" BASIS,
11
+ # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12
+ # See the License for the specific language governing permissions and
13
+ # limitations under the License.
14
+
15
+ from typing import List, Union
16
+ import numpy as np
17
+ import torch
18
+ from transformers.feature_extraction_utils import BatchFeature
19
+ from transformers.processing_utils import (
20
+ ProcessingKwargs,
21
+ ProcessorMixin,
22
+ Unpack,
23
+ VideosKwargs,
24
+ )
25
+ from transformers.tokenization_utils_base import PreTokenizedInput, TextInput
26
+
27
+
28
+ ImageInput = Union[
29
+ "PIL.Image.Image",
30
+ np.ndarray,
31
+ "torch.Tensor",
32
+ List["PIL.Image.Image"],
33
+ List[np.ndarray],
34
+ List["torch.Tensor"],
35
+ ] # noqa
36
+
37
+
38
+ VideoInput = Union[
39
+ List["PIL.Image.Image"],
40
+ "np.ndarray",
41
+ "torch.Tensor",
42
+ List["np.ndarray"],
43
+ List["torch.Tensor"],
44
+ List[List["PIL.Image.Image"]],
45
+ List[List["np.ndarrray"]],
46
+ List[List["torch.Tensor"]],
47
+ ] # noqa
48
+
49
+
50
+ class PaddleOCRVLVideosProcessorKwargs(VideosKwargs, total=False):
51
+ fps: Union[List[float], float]
52
+
53
+
54
+ class PaddleOCRVLProcessorKwargs(ProcessingKwargs, total=False):
55
+ videos_kwargs: PaddleOCRVLVideosProcessorKwargs
56
+ _defaults = {
57
+ "text_kwargs": {
58
+ "padding": False,
59
+ },
60
+ "videos_kwargs": {"fps": 2.0},
61
+ }
62
+
63
+
64
+ class PaddleOCRVLProcessor(ProcessorMixin):
65
+ r"""
66
+ [`PaddleOCRVLProcessor`] offers all the functionalities of [`SiglipImageProcessor`] and [`Qwen2TokenizerFast`]. See the
67
+ [`~PaddleOCRVLProcessor.__call__`] and [`~PaddleOCRVLProcessor.decode`] for more information.
68
+ Args:
69
+ image_processor ([`SiglipImageProcessor`], *optional*):
70
+ The image processor is a required input.
71
+ tokenizer ([`Qwen2TokenizerFast`], *optional*):
72
+ The tokenizer is a required input.
73
+ chat_template (`str`, *optional*): A Jinja template which will be used to convert lists of messages
74
+ in a chat into a tokenizable string.
75
+ """
76
+
77
+ attributes = ["image_processor", "tokenizer"]
78
+ valid_kwargs = [
79
+ "chat_template",
80
+ "image_std",
81
+ "min_pixels",
82
+ "image_mean",
83
+ "merge_size",
84
+ "image_processor_type",
85
+ "temporal_patch_size",
86
+ "patch_size",
87
+ "max_pixels",
88
+ ]
89
+
90
+ image_processor_class = "AutoImageProcessor"
91
+ tokenizer_class = "AutoTokenizer"
92
+
93
+ def __init__(
94
+ self, image_processor=None, tokenizer=None, chat_template=None, **kwargs
95
+ ):
96
+ self.image_token = (
97
+ "<|IMAGE_PLACEHOLDER|>"
98
+ if not hasattr(tokenizer, "image_token")
99
+ else tokenizer.image_token
100
+ )
101
+ self.video_token = (
102
+ "<|video_pad|>"
103
+ if not hasattr(tokenizer, "video_token")
104
+ else tokenizer.video_token
105
+ )
106
+ super().__init__(image_processor, tokenizer, chat_template=chat_template)
107
+
108
+ def __call__(
109
+ self,
110
+ images: ImageInput = None,
111
+ text: Union[
112
+ TextInput, PreTokenizedInput, List[TextInput], List[PreTokenizedInput]
113
+ ] = None,
114
+ videos: VideoInput = None,
115
+ **kwargs: Unpack[PaddleOCRVLProcessorKwargs],
116
+ ) -> BatchFeature:
117
+ """
118
+ Main method to prepare for the model one or several sequences(s) and image(s). This method forwards the `text`
119
+ and `kwargs` arguments to Qwen2TokenizerFast's [`~Qwen2TokenizerFast.__call__`] if `text` is not `None` to encode
120
+ the text. To prepare the vision inputs, this method forwards the `vision_infos` and `kwrags` arguments to
121
+ SiglipImageProcessor's [`~SiglipImageProcessor.__call__`] if `vision_infos` is not `None`.
122
+
123
+ Args:
124
+ images (`PIL.Image.Image`, `np.ndarray`, `torch.Tensor`, `List[PIL.Image.Image]`, `List[np.ndarray]`, `List[torch.Tensor]`):
125
+ The image or batch of images to be prepared. Each image can be a PIL image, NumPy array or PyTorch
126
+ tensor. Both channels-first and channels-last formats are supported.
127
+ text (`str`, `List[str]`, `List[List[str]]`):
128
+ The sequence or batch of sequences to be encoded. Each sequence can be a string or a list of strings
129
+ (pretokenized string). If the sequences are provided as list of strings (pretokenized), you must set
130
+ `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).
131
+ videos (`np.ndarray`, `torch.Tensor`, `List[np.ndarray]`, `List[torch.Tensor]`):
132
+ The image or batch of videos to be prepared. Each video can be a 4D NumPy array or PyTorch
133
+ tensor, or a nested list of 3D frames. Both channels-first and channels-last formats are supported.
134
+ return_tensors (`str` or [`~utils.TensorType`], *optional*):
135
+ If set, will return tensors of a particular framework. Acceptable values are:
136
+ - `'tf'`: Return TensorFlow `tf.constant` objects.
137
+ - `'pt'`: Return PyTorch `torch.Tensor` objects.
138
+ - `'np'`: Return NumPy `np.ndarray` objects.
139
+ - `'jax'`: Return JAX `jnp.ndarray` objects.
140
+
141
+ Returns:
142
+ [`BatchFeature`]: A [`BatchFeature`] with the following fields:
143
+
144
+ - **input_ids** -- List of token ids to be fed to a model. Returned when `text` is not `None`.
145
+ - **attention_mask** -- List of indices specifying which tokens should be attended to by the model (when
146
+ `return_attention_mask=True` or if *"attention_mask"* is in `self.model_input_names` and if `text` is not
147
+ `None`).
148
+ - **pixel_values** -- Pixel values to be fed to a model. Returned when `images` is not `None`.
149
+ - **pixel_values_videos** -- Pixel values of videos to be fed to a model. Returned when `videos` is not `None`.
150
+ - **image_grid_thw** -- List of image 3D grid in LLM. Returned when `images` is not `None`.
151
+ - **video_grid_thw** -- List of video 3D grid in LLM. Returned when `videos` is not `None`.
152
+ - **second_per_grid_ts** -- List of video seconds per time grid. Returned when `videos` is not `None`.
153
+ """
154
+ output_kwargs = self._merge_kwargs(
155
+ PaddleOCRVLProcessorKwargs,
156
+ tokenizer_init_kwargs=self.tokenizer.init_kwargs,
157
+ **kwargs,
158
+ )
159
+
160
+ if images is not None:
161
+ image_inputs = self.image_processor(images=images, return_tensors="pt")
162
+ image_inputs["pixel_values"] = image_inputs["pixel_values"]
163
+ image_grid_thw = image_inputs["image_grid_thw"]
164
+
165
+ else:
166
+ image_inputs = {}
167
+ image_grid_thw = None
168
+
169
+ if videos is not None:
170
+ # TODO: add video processing
171
+ videos_inputs = self.image_processor(
172
+ images=None, videos=videos, **output_kwargs["images_kwargs"]
173
+ )
174
+ video_grid_thw = videos_inputs["video_grid_thw"]
175
+
176
+ fps = output_kwargs["videos_kwargs"].pop("fps", 2.0)
177
+ if isinstance(fps, (int, float)):
178
+ second_per_grid_ts = [
179
+ self.image_processor.temporal_patch_size / fps
180
+ ] * len(video_grid_thw)
181
+ elif hasattr(fps, "__len__") and len(fps) == len(video_grid_thw):
182
+ second_per_grid_ts = [
183
+ self.image_processor.temporal_patch_size / tmp for tmp in fps
184
+ ]
185
+ else:
186
+ raise ValueError(
187
+ f"The length of fps ({len(fps) if hasattr(fps, '__len__') else fps}) must be equal to the length of video_grid_thw ({len(video_grid_thw)}) or fps should be a single number."
188
+ )
189
+ videos_inputs.update(
190
+ {"second_per_grid_ts": torch.tensor(second_per_grid_ts)}
191
+ )
192
+
193
+ else:
194
+ videos_inputs = {}
195
+ video_grid_thw = None
196
+
197
+ if not isinstance(text, list):
198
+ text = [text]
199
+
200
+ if image_grid_thw is not None:
201
+ index = 0
202
+ for i in range(len(text)):
203
+ while self.image_token in text[i]:
204
+ text[i] = text[i].replace(
205
+ self.image_token,
206
+ "<|placeholder|>"
207
+ * (
208
+ image_grid_thw[index].prod()
209
+ // self.image_processor.merge_size
210
+ // self.image_processor.merge_size
211
+ ),
212
+ 1,
213
+ )
214
+ index += 1
215
+ text[i] = text[i].replace("<|placeholder|>", self.image_token)
216
+
217
+ if video_grid_thw is not None:
218
+ index = 0
219
+ for i in range(len(text)):
220
+ while self.video_token in text[i]:
221
+ text[i] = text[i].replace(
222
+ self.video_token,
223
+ "<|placeholder|>"
224
+ * (
225
+ video_grid_thw[index].prod()
226
+ // self.image_processor.merge_size
227
+ // self.image_processor.merge_size
228
+ ),
229
+ 1,
230
+ )
231
+ index += 1
232
+ text[i] = text[i].replace("<|placeholder|>", self.video_token)
233
+
234
+ text_inputs = self.tokenizer(text, **output_kwargs["text_kwargs"])
235
+
236
+ return BatchFeature(data={**text_inputs, **image_inputs, **videos_inputs})
237
+
238
+ def batch_decode(self, *args, **kwargs):
239
+ """
240
+ This method forwards all its arguments to Qwen2TokenizerFast's [`~PreTrainedTokenizer.batch_decode`]. Please
241
+ refer to the docstring of this method for more information.
242
+ """
243
+ return self.tokenizer.batch_decode(*args, **kwargs)
244
+
245
+ def decode(self, *args, **kwargs):
246
+ """
247
+ This method forwards all its arguments to Qwen2TokenizerFast's [`~PreTrainedTokenizer.decode`]. Please refer to
248
+ the docstring of this method for more information.
249
+ """
250
+ return self.tokenizer.decode(*args, **kwargs)
251
+
252
+ def post_process_image_text_to_text(
253
+ self,
254
+ generated_outputs,
255
+ skip_special_tokens=True,
256
+ clean_up_tokenization_spaces=False,
257
+ **kwargs,
258
+ ):
259
+ """
260
+ Post-process the output of the model to decode the text.
261
+
262
+ Args:
263
+ generated_outputs (`torch.Tensor` or `np.ndarray`):
264
+ The output of the model `generate` function. The output is expected to be a tensor of shape `(batch_size, sequence_length)`
265
+ or `(sequence_length,)`.
266
+ skip_special_tokens (`bool`, *optional*, defaults to `True`):
267
+ Whether or not to remove special tokens in the output. Argument passed to the tokenizer's `batch_decode` method.
268
+ Clean_up_tokenization_spaces (`bool`, *optional*, defaults to `False`):
269
+ Whether or not to clean up the tokenization spaces. Argument passed to the tokenizer's `batch_decode` method.
270
+ **kwargs:
271
+ Additional arguments to be passed to the tokenizer's `batch_decode method`.
272
+
273
+ Returns:
274
+ `List[str]`: The decoded text.
275
+ """
276
+ return self.tokenizer.batch_decode(
277
+ generated_outputs,
278
+ skip_special_tokens=skip_special_tokens,
279
+ clean_up_tokenization_spaces=clean_up_tokenization_spaces,
280
+ **kwargs,
281
+ )
282
+
283
+ @property
284
+ def model_input_names(self):
285
+ tokenizer_input_names = self.tokenizer.model_input_names
286
+ image_processor_input_names = self.image_processor.model_input_names
287
+ names_from_processor = list(
288
+ dict.fromkeys(tokenizer_input_names + image_processor_input_names)
289
+ )
290
+ return names_from_processor + ["second_per_grid_ts"]
291
+
292
+
293
+ __all__ = ["PaddleOCRVLProcessor", "PaddleOCRVLProcessor"]
processor_config.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor": {
3
+ "auto_map": {
4
+ "AutoImageProcessor": "image_processing_paddleocr_vl.PaddleOCRVLImageProcessor",
5
+ "AutoProcessor": "processing_paddleocr_vl.PaddleOCRVLProcessor"
6
+ },
7
+ "do_convert_rgb": true,
8
+ "do_normalize": true,
9
+ "do_rescale": true,
10
+ "do_resize": true,
11
+ "image_mean": [
12
+ 0.5,
13
+ 0.5,
14
+ 0.5
15
+ ],
16
+ "image_processor_type": "PaddleOCRVLImageProcessor",
17
+ "image_std": [
18
+ 0.5,
19
+ 0.5,
20
+ 0.5
21
+ ],
22
+ "merge_size": 2,
23
+ "patch_size": 14,
24
+ "resample": 3,
25
+ "rescale_factor": 0.00392156862745098,
26
+ "size": {
27
+ "longest_edge": 1003520,
28
+ "shortest_edge": 112896
29
+ },
30
+ "temporal_patch_size": 1
31
+ },
32
+ "processor_class": "PaddleOCRVLProcessor"
33
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c8a215a59183d0d0781adc33bacd3ce6162716f7fd568fb30234a74d69803a7d
3
+ size 11189060
tokenizer_config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "auto_map": {
4
+ "AutoProcessor": "processing_paddleocr_vl.PaddleOCRVLProcessor"
5
+ },
6
+ "backend": "tokenizers",
7
+ "bos_token": "<s>",
8
+ "clean_up_tokenization_spaces": false,
9
+ "cls_token": "<|begin_of_sentence|>",
10
+ "eos_token": "</s>",
11
+ "image_token": "<|IMAGE_PLACEHOLDER|>",
12
+ "is_local": false,
13
+ "legacy": true,
14
+ "local_files_only": false,
15
+ "mask_token": "<mask:1>",
16
+ "model_max_length": 131072,
17
+ "model_specific_special_tokens": {
18
+ "image_token": "<|IMAGE_PLACEHOLDER|>"
19
+ },
20
+ "pad_token": "<unk>",
21
+ "processor_class": "PaddleOCRVLProcessor",
22
+ "sep_token": "<|end_of_sentence|>",
23
+ "sp_model_kwargs": {},
24
+ "spaces_between_special_tokens": false,
25
+ "tokenizer_class": "TokenizersBackend",
26
+ "unk_token": "<unk>",
27
+ "use_default_system_prompt": false
28
+ }
usage.py ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Read one electrical one-line diagram image into JSON.
2
+
3
+ python usage.py drawing.png [--device cuda|cpu]
4
+
5
+ Apache-2.0, Copyright 2026 Decosa.
6
+ """
7
+ import argparse
8
+ import json
9
+ import re
10
+
11
+ import torch
12
+ from PIL import Image
13
+ from transformers import AutoModelForImageTextToText, AutoProcessor
14
+
15
+ REPO = "decosaai/decosa-oneline-reader-paddleocr-vl"
16
+ PROMPT = "One-line Diagram Recognition:"
17
+ MAX_PIXELS = 1003520 # the processor's cap
18
+
19
+
20
+ def fit(im: Image.Image) -> Image.Image:
21
+ """Scale to just under the processor's pixel cap; small or poor scans are enlarged (as in training)."""
22
+ w, h = im.size
23
+ s = (MAX_PIXELS * 0.98 / (w * h)) ** 0.5
24
+ if s < 1 or w < 1100:
25
+ im = im.resize((max(28, int(w * s)), max(28, int(h * s))), Image.BICUBIC)
26
+ return im
27
+
28
+
29
+ def main() -> None:
30
+ ap = argparse.ArgumentParser()
31
+ ap.add_argument("image")
32
+ ap.add_argument("--model", default=REPO)
33
+ ap.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
34
+ a = ap.parse_args()
35
+ proc = AutoProcessor.from_pretrained(a.model)
36
+ dtype = torch.bfloat16 if a.device == "cuda" else torch.float32
37
+ model = AutoModelForImageTextToText.from_pretrained(a.model, dtype=dtype).to(a.device).eval()
38
+ im = fit(Image.open(a.image).convert("L")).convert("RGB")
39
+ msgs = [{"role": "user", "content": [{"type": "image", "image": im}, {"type": "text", "text": PROMPT}]}]
40
+ x = proc.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True, return_dict=True,
41
+ return_tensors="pt").to(a.device)
42
+ with torch.inference_mode():
43
+ out = model.generate(**x, max_new_tokens=900, do_sample=False, use_cache=True)
44
+ text = proc.decode(out[0][x["input_ids"].shape[-1]:], skip_special_tokens=True)
45
+ text = re.sub(r"<\|[a-z_]+\|>|</s>|<s>|<pad>", "", text).strip()
46
+ try:
47
+ print(json.dumps(json.loads(text), indent=1))
48
+ except json.JSONDecodeError:
49
+ print(text)
50
+
51
+
52
+ if __name__ == "__main__":
53
+ main()