Reza2kn commited on
Commit
f528e13
·
verified ·
1 Parent(s): e76c92d

Update model card with mini benchmark limitations

Browse files
Files changed (1) hide show
  1. README.md +39 -18
README.md CHANGED
@@ -12,38 +12,59 @@ tags:
12
  - ocr
13
  - document-ai
14
  - surya
 
15
  ---
16
 
17
  # Surya OCR 2 MLX 8-bit G64
18
 
19
- This repository contains a converted/quantized artifact derived from [datalab-to/surya-ocr-2](https://huggingface.co/datalab-to/surya-ocr-2).
20
 
21
- ## What is included
22
 
23
- - Source model: `datalab-to/surya-ocr-2`
24
- - Runtime/format: MLX / mlx-vlm
25
- - Quantization: 8-bit affine weight quantization, group size 64
26
- - Vision weights included: yes, included in the MLX checkpoint
27
- - Created for: local OCR/document-understanding experiments and parity testing
28
 
29
- ## Validation status
 
 
 
 
30
 
31
- Mini olmOCR-bench score: **79.2% ± 6.2%**. This did not meet the current >98% parity target against the source mini baseline. Failures concentrated on long tiny text and old scans.
32
 
33
- ## Known caveats
 
 
 
34
 
35
- Early quantization artifact. Runtime smoke/load is available, but benchmark parity is not yet acceptable.
36
 
37
- ## Files
38
 
39
- MLX `model.safetensors`, tokenizer, processor, config, generation config, and chat template.
40
 
41
- ## Usage
42
 
43
- Use with `mlx-vlm` by passing this repository or a local clone path to `mlx_vlm.load`.
44
 
45
- ## Provenance
46
 
47
- This artifact was generated non-destructively from the original Hugging Face checkpoint. It is not a new fine-tune.
48
 
49
- If you need production parity, compare against the original model on your own document distribution before deployment.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
  - ocr
13
  - document-ai
14
  - surya
15
+ - experimental
16
  ---
17
 
18
  # Surya OCR 2 MLX 8-bit G64
19
 
20
+ This repository contains an **experimental quantized** artifact derived from [datalab-to/surya-ocr-2](https://huggingface.co/datalab-to/surya-ocr-2).
21
 
22
+ This 8-bit MLX quant is the most useful Apple-side artifact from the current batch. It keeps perfect mini-section scores on arxiv math, headers/footers, multi-column, old-scans-math, tables, and baseline checks, but it currently fails the old-scans mini split and is weak on long tiny text.
23
 
24
+ ## What is included
 
 
 
 
25
 
26
+ - Source model: `datalab-to/surya-ocr-2`
27
+ - Runtime/format: MLX / mlx-vlm
28
+ - Quantization: 8-bit affine weight quantization, group size 64
29
+ - Vision weights included: Yes. The MLX checkpoint includes the model vision weights and processor assets.
30
+ - Processor/tokenizer assets: included
31
 
32
+ ## Mini olmOCR-bench results
33
 
34
+ | Candidate | Overall | Arxiv math | Headers/footers | Long tiny text | Multi-column | Old scans | Old scans math | Tables | Baseline |
35
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
36
+ | Source mini baseline | 91.0% ± 6.3% | 100.0% | 100.0% | 100.0% | 100.0% | 33.3% | 100.0% | 100.0% | 94.7% |
37
+ | Surya OCR 2 MLX 8-bit G64 | 79.2% ± 6.2% | 100.0% | 100.0% | 33.3% | 100.0% | 0.0% | 100.0% | 100.0% | 100.0% |
38
 
 
39
 
 
40
 
41
+ ## How to read the benchmark table
42
 
43
+ This is an early quant release with transparent limitations. The table uses our local 40-test mini slice of `allenai/olmOCR-bench`, with 3 samples from each named section plus the benchmark baseline checks. It is **not** the full public score and it is **not** a claim of >98% parity.
44
 
45
+ The useful signal is the split behavior: this artifact is currently strong on clean academic/math, headers/footers, multi-column layouts, tables, old-scan math, and baseline OCR checks, but it should not be used for old degraded scans and is weak on long tiny text.
46
 
47
+ ## Recommended use
48
 
49
+ Use this checkpoint for local experimentation and constrained OCR workloads whose documents resemble the passing sections above. Avoid using it as a production replacement for the original model on degraded historical scans, very small dense body text, or workloads requiring full benchmark parity.
50
 
51
+
52
+ ## Loading
53
+
54
+ ```python
55
+ from mlx_vlm import load, generate
56
+
57
+ model, processor = load("Reza2kn/surya-ocr-2-mlx-8bit-g64")
58
+ # Pass images/documents through the same Surya/MLX-VLM prompting path used by your app.
59
+ ```
60
+
61
+ ## Limitations
62
+
63
+ - This is not a full-parity release yet.
64
+ - Do **not** use this artifact for degraded old scans; the current mini split score is 0.0% there.
65
+ - Do **not** use this artifact for long tiny text unless you independently validate your data; the current mini split score is 33.3%.
66
+ - Math-heavy and table/layout-heavy mini examples looked good in this slice, but full olmOCR-bench is still pending.
67
+
68
+ ## Provenance
69
+
70
+ Generated non-destructively from the original Hugging Face checkpoint. This is not a fine-tune. The goal of publishing this artifact now is transparency: the files are usable for the passing workload slices above, and the known failing slices are documented clearly.