Mayo commited on
Commit
205395b
·
unverified ·
1 Parent(s): 7571f52

Add complete model card metadata

Browse files
Files changed (1) hide show
  1. README.md +44 -0
README.md ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: gpl-3.0
3
+ library_name: candle
4
+ pipeline_tag: image-to-text
5
+ tags:
6
+ - candle
7
+ - ocr
8
+ - image-to-text
9
+ - manga
10
+ - comic
11
+ - multilingual
12
+ - safetensors
13
+ ---
14
+
15
+ # MIT 48px OCR
16
+
17
+ A SafeTensors conversion of the 48-pixel OCR model used by [`zyddnys/manga-image-translator`](https://github.com/zyddnys/manga-image-translator) and BallonsTranslator. The model recognizes cropped comic text lines and also predicts foreground and background colors.
18
+
19
+ ## Model details
20
+
21
+ - Input text height: 48 pixels
22
+ - Maximum configured width: 8100 pixels
23
+ - Visual backbone: ConvNeXt feature extractor
24
+ - Sequence model: four-layer Transformer encoder and five-layer Transformer decoder
25
+ - Embedding dimension: 320
26
+ - Attention heads: 4
27
+ - Default beam size: 5
28
+ - Default maximum sequence length: 255
29
+
30
+ ## Files
31
+
32
+ - `model.safetensors`: converted model weights
33
+ - `config.json`: architecture and decoding configuration
34
+ - `alphabet-all-v7.txt`: tokenizer alphabet and special-token vocabulary
35
+
36
+ Token IDs are `0` for padding, `1` for beginning-of-sequence, and `2` for end-of-sequence. `<SP>` represents a space.
37
+
38
+ ## Intended use and limitations
39
+
40
+ The model expects already detected, cropped, and normalized comic text regions. It does not locate text on a page. Recognition quality depends strongly on crop quality, text scale, language coverage in the supplied alphabet, and image degradation. Training data details and evaluation metrics are not included with this conversion.
41
+
42
+ ## License
43
+
44
+ GPL-3.0, following the upstream manga-image-translator and BallonsTranslator implementations.