File size: 7,594 Bytes
fffb88c 845ff16 4f7ab5a 845ff16 4f7ab5a fffb88c 4f7ab5a fffb88c 4f7ab5a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 | # Benchmark protocol and claim boundaries
## Preliminary additional server comparison: Phoenix vs Baseer__Nakba
These values were produced in an owner-run server experiment and supplied as aggregate results. The model repository does not yet contain the raw line predictions, sample counts, preprocessing manifest, or runnable evaluation bundle for this experiment. They are therefore labelled **preliminary** and must not be described as independently reproduced or sealed.
| Metric | Phoenix | Baseer__Nakba | Difference |
|---|---:|---:|---:|
| Unweighted macro CER over 4 datasets | **9.71%** | 28.67% | −18.97 pp; **66.1% relative reduction** |
| WER | **34.37%** | 59.52% | −25.15 pp; **42.3% relative reduction** |
| Exact-line accuracy | 13.38% | **26.38%** | Baseer +13.00 pp |
| Parameters | **4,988,946** | ≈3.75B | Phoenix ≈752× smaller |
| File size | **19,936,101 bytes** | ≈7.53 GB | Phoenix ≈378× smaller |
| Dataset | Phoenix CER | Baseer CER | Winner | Phoenix relative CER reduction |
|---|---:|---:|---|---:|
| Agapet | **11.13%** | 35.65% | Phoenix | 68.8% |
| Muharaf | **11.93%** | 25.51% | Phoenix | 53.2% |
| Omar | 6.88% | **0.48%** | Baseer__Nakba | — |
| RASAM | **8.90%** | 53.06% | Phoenix | 83.2% |
The macro is an unweighted average over dataset-level CER, not a character-weighted aggregate. Recomputing from the displayed rows gives Phoenix 9.71% exactly and Baseer 28.675%; the source displays 28.67%. WER and exact-line aggregation denominators were not supplied. The result supports a strong efficiency and average-CER claim, but not universal superiority: Baseer wins Omar and exact-line accuracy.
Machine-readable details and publication constraints are in `baseer_server_benchmark.json`.
## Sealed comparison used for the release decision
- Arms: exp8 and exp9.
- Decoder: greedy only; no language model.
- Total evaluated lines: 22,442.
- Split unit: manuscript or document where applicable.
- exp9 was not tuned after these results were opened.
| Set | Lines | exp8 CER | exp9 CER | exp9 WER | exp9 exact-line rate |
|---|---:|---:|---:|---:|---:|
| Agapet | 10,594 | 22.124% | 17.861% | 58.798% | 0.123% |
| Omar | 11,684 | 17.718% | 11.835% | 42.917% | 3.637% |
| TariMa | 164 | 10.386% | 10.718% | 38.886% | 4.268% |
## Same-protocol model diagnostic
### Protocol
- Inputs: the same pre-cropped single-line PAGE samples for all model arms.
- References: the same raw references; punctuation and digits are counted.
- Decoder: Kraken greedy; no language model and no text normalization.
- Command shape: `ketos test -f page -B 2` on CUDA.
- Metric: CER is `100 − char accuracy`; word accuracy is reported independently.
### Results
Each cell is **CER / word accuracy**.
| Guard | Lines | Guard-list SHA-256 prefix | Original Muharaf | exp6 | exp8 | **exp9** |
|---|---:|---|---:|---:|---:|---:|
| Agapet | 991 | `cc347dc100c7778b` | 33.15 / 16.62 | 23.82 / 24.71 | 20.08 / 33.58 | **9.47 / 67.40** |
| Omar | 1,143 | `3e6b577ebb1cc1bf` | 9.26 / 60.08 | 26.88 / 21.25 | 12.06 / 53.90 | **8.91 / 63.95** |
| RASAM | 1,789 | `61ea9c5bc6172bbb` | 39.01 / 12.06 | 9.04 / 66.78 | 7.95 / 69.90 | **7.67 / 70.46** |
| Muharaf | 920 | `0379f87eef74bbf3` | 13.28 / **64.15** | 36.47 / 15.74 | 12.72 / 58.30 | **12.31** / 59.32 |
| **Unweighted macro over comparable guards** | **4,843** | — | 23.68 / 38.23 | 24.05 / 32.12 | 13.20 / 53.92 | **9.59 / 65.28** |
exp9 achieved the lowest CER on 4/4 comparable guards and the highest word accuracy on 3/4. Relative to the original Muharaf model, the macro-CER reduction is **59.5%**. The Muharaf-domain row is deliberately not reduced to a single “win”: exp9 has lower CER, while the original Muharaf model has higher word accuracy.
### TariMa guard excluded from ranking
| Guard | Lines | Original Muharaf CER | exp6 CER | exp8 CER | exp9 CER |
|---|---:|---:|---:|---:|---:|
| TariMa manuscript guard | 369 | 41.93% | 0.93% | 1.40% | 2.99% |
This row is excluded from the macro and from winner counts. It is manuscript-independent for exp9 but the held manuscripts were used when training exp6/exp8, so it cannot fairly rank all four models.
### Model identities
| Arm | SHA-256 |
|---|---|
| Original Muharaf | `726052869fa5aaf253d906d1c228e87bf80203a0cdcc8c3de50873d2697df178` |
| exp6 | `3d6d3c3df380355e2390db33c9c981487e2d3660deb928da61a352902bc6a79b` |
| exp8 | `d13c0996c14dbf805eeeeb4c718663b9686055863721b326e1cfe5fdeef82707` |
| exp9 | `2896fef9d9665cbb82fba8faa3bf0c628ac6a5cc62eb678f40c707db833aebea` |
Full hashes and machine-readable aggregate scores are stored in `benchmark_summary.json`. The original local evaluation artifact is not included because it contains workstation paths.
This diagnostic re-evaluates the original Muharaf recognizer, exp6, exp8, and exp9 on the same guard lists and raw references. These guards are development data, so the table supports domain diagnosis and protocol-matched comparison, not an independent final-generalization claim.
## Claims that are allowed
- exp9 outperformed exp8 on the sealed Agapet and Omar sets.
- exp9 regressed slightly on sealed TariMa.
- the same-protocol guard table compares the four listed model files fairly on those exact guard inputs.
- Athar provides system-level functionality that a recognizer weight file alone does not provide.
- under the supplied server protocol, Phoenix had lower CER on 3/4 datasets, lower unweighted macro CER, and lower WER than Baseer__Nakba while being roughly 752× smaller by parameter count.
- Baseer__Nakba had lower Omar CER and higher exact-line accuracy in the same supplied comparison.
## Claims that are not allowed
- “exp9 beats the published 18.1% Muharaf result” without reproducing its official protocol.
- “best Arabic HTR model”.
- comparing Arabic-normalized CER with another system's raw CER.
- treating guard-set results as unseen final-test evidence.
- attributing Athar's retrieval, review, or export capabilities to the exp9 weights.
- calling the Baseer comparison independently reproduced, sealed, or universally representative until its raw predictions, counts, preprocessing, and evaluation code are archived.
## Historical development results
These rows document the project's progression, but they use **different evaluation sets**. They must not be read as a single head-to-head leaderboard.
| Model | Evaluation set | LM | CER | WER | Character accuracy | Word accuracy |
|---|---|---:|---:|---:|---:|---:|
| Muharaf baseline | RASAM-test | No | 37.91% | 88.12% | 62.09% | 11.88% |
| exp4A | RASAM-test | No | 7.64% | 27.77% | 92.36% | 72.23% |
| exp6 | RASAM-test | No | 7.48% | 27.17% | 92.52% | 72.83% |
| exp6 + order-8 LM | RASAM-test | Yes | 6.76% | 27.17% | 93.24% | 72.83% |
| exp6 + order-8 LM | TariMa-test | Yes | 8.68% | 37.84% | 91.32% | 62.16% |
| exp7_easy | Easy-Family | No | 5.94% | 22.38% | 94.06% | 77.62% |
| exp7_hard | Hard-Family | No | 13.70% | 45.61% | 86.30% | 54.39% |
### Historical assistance-layer measurements
| Component | Measurement | Result |
|---|---|---:|
| Confidence-gated LM | useful interventions on auto-seg run | 12/18 (67%) |
| Contextual candidates | difficult lines improved / harmed | 6/10 / 0 |
| Catalog search | correct work ranked first | 5/5 |
| Retrieval correction | Qur'an al-ʿAlaq match with reference | 98.3% |
Historical logic-domain LM perplexity was 53.14 for the general model, 40.70 for the self-trained model (−23.4%), and 20.25 for the mixed model (−61.9%). These are language-model diagnostics, not exp9 visual-recognizer CER results.
|