Mayo commited on
Add complete model card metadata
Browse files
README.md
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: gpl-3.0
|
| 3 |
+
library_name: candle
|
| 4 |
+
pipeline_tag: image-to-text
|
| 5 |
+
tags:
|
| 6 |
+
- candle
|
| 7 |
+
- ocr
|
| 8 |
+
- image-to-text
|
| 9 |
+
- manga
|
| 10 |
+
- comic
|
| 11 |
+
- multilingual
|
| 12 |
+
- safetensors
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# MIT 48px OCR
|
| 16 |
+
|
| 17 |
+
A SafeTensors conversion of the 48-pixel OCR model used by [`zyddnys/manga-image-translator`](https://github.com/zyddnys/manga-image-translator) and BallonsTranslator. The model recognizes cropped comic text lines and also predicts foreground and background colors.
|
| 18 |
+
|
| 19 |
+
## Model details
|
| 20 |
+
|
| 21 |
+
- Input text height: 48 pixels
|
| 22 |
+
- Maximum configured width: 8100 pixels
|
| 23 |
+
- Visual backbone: ConvNeXt feature extractor
|
| 24 |
+
- Sequence model: four-layer Transformer encoder and five-layer Transformer decoder
|
| 25 |
+
- Embedding dimension: 320
|
| 26 |
+
- Attention heads: 4
|
| 27 |
+
- Default beam size: 5
|
| 28 |
+
- Default maximum sequence length: 255
|
| 29 |
+
|
| 30 |
+
## Files
|
| 31 |
+
|
| 32 |
+
- `model.safetensors`: converted model weights
|
| 33 |
+
- `config.json`: architecture and decoding configuration
|
| 34 |
+
- `alphabet-all-v7.txt`: tokenizer alphabet and special-token vocabulary
|
| 35 |
+
|
| 36 |
+
Token IDs are `0` for padding, `1` for beginning-of-sequence, and `2` for end-of-sequence. `<SP>` represents a space.
|
| 37 |
+
|
| 38 |
+
## Intended use and limitations
|
| 39 |
+
|
| 40 |
+
The model expects already detected, cropped, and normalized comic text regions. It does not locate text on a page. Recognition quality depends strongly on crop quality, text scale, language coverage in the supplied alphabet, and image degradation. Training data details and evaluation metrics are not included with this conversion.
|
| 41 |
+
|
| 42 |
+
## License
|
| 43 |
+
|
| 44 |
+
GPL-3.0, following the upstream manga-image-translator and BallonsTranslator implementations.
|