factlogic's picture
Rename public model to Phoenix and clarify evidence levels
845ff16 verified
|
Raw
History Blame
7.59 kB

Benchmark protocol and claim boundaries

Preliminary additional server comparison: Phoenix vs Baseer__Nakba

These values were produced in an owner-run server experiment and supplied as aggregate results. The model repository does not yet contain the raw line predictions, sample counts, preprocessing manifest, or runnable evaluation bundle for this experiment. They are therefore labelled preliminary and must not be described as independently reproduced or sealed.

Metric Phoenix Baseer__Nakba Difference
Unweighted macro CER over 4 datasets 9.71% 28.67% −18.97 pp; 66.1% relative reduction
WER 34.37% 59.52% −25.15 pp; 42.3% relative reduction
Exact-line accuracy 13.38% 26.38% Baseer +13.00 pp
Parameters 4,988,946 ≈3.75B Phoenix ≈752× smaller
File size 19,936,101 bytes ≈7.53 GB Phoenix ≈378× smaller
Dataset Phoenix CER Baseer CER Winner Phoenix relative CER reduction
Agapet 11.13% 35.65% Phoenix 68.8%
Muharaf 11.93% 25.51% Phoenix 53.2%
Omar 6.88% 0.48% Baseer__Nakba
RASAM 8.90% 53.06% Phoenix 83.2%

The macro is an unweighted average over dataset-level CER, not a character-weighted aggregate. Recomputing from the displayed rows gives Phoenix 9.71% exactly and Baseer 28.675%; the source displays 28.67%. WER and exact-line aggregation denominators were not supplied. The result supports a strong efficiency and average-CER claim, but not universal superiority: Baseer wins Omar and exact-line accuracy.

Machine-readable details and publication constraints are in baseer_server_benchmark.json.

Sealed comparison used for the release decision

  • Arms: exp8 and exp9.
  • Decoder: greedy only; no language model.
  • Total evaluated lines: 22,442.
  • Split unit: manuscript or document where applicable.
  • exp9 was not tuned after these results were opened.
Set Lines exp8 CER exp9 CER exp9 WER exp9 exact-line rate
Agapet 10,594 22.124% 17.861% 58.798% 0.123%
Omar 11,684 17.718% 11.835% 42.917% 3.637%
TariMa 164 10.386% 10.718% 38.886% 4.268%

Same-protocol model diagnostic

Protocol

  • Inputs: the same pre-cropped single-line PAGE samples for all model arms.
  • References: the same raw references; punctuation and digits are counted.
  • Decoder: Kraken greedy; no language model and no text normalization.
  • Command shape: ketos test -f page -B 2 on CUDA.
  • Metric: CER is 100 − char accuracy; word accuracy is reported independently.

Results

Each cell is CER / word accuracy.

Guard Lines Guard-list SHA-256 prefix Original Muharaf exp6 exp8 exp9
Agapet 991 cc347dc100c7778b 33.15 / 16.62 23.82 / 24.71 20.08 / 33.58 9.47 / 67.40
Omar 1,143 3e6b577ebb1cc1bf 9.26 / 60.08 26.88 / 21.25 12.06 / 53.90 8.91 / 63.95
RASAM 1,789 61ea9c5bc6172bbb 39.01 / 12.06 9.04 / 66.78 7.95 / 69.90 7.67 / 70.46
Muharaf 920 0379f87eef74bbf3 13.28 / 64.15 36.47 / 15.74 12.72 / 58.30 12.31 / 59.32
Unweighted macro over comparable guards 4,843 23.68 / 38.23 24.05 / 32.12 13.20 / 53.92 9.59 / 65.28

exp9 achieved the lowest CER on 4/4 comparable guards and the highest word accuracy on 3/4. Relative to the original Muharaf model, the macro-CER reduction is 59.5%. The Muharaf-domain row is deliberately not reduced to a single “win”: exp9 has lower CER, while the original Muharaf model has higher word accuracy.

TariMa guard excluded from ranking

Guard Lines Original Muharaf CER exp6 CER exp8 CER exp9 CER
TariMa manuscript guard 369 41.93% 0.93% 1.40% 2.99%

This row is excluded from the macro and from winner counts. It is manuscript-independent for exp9 but the held manuscripts were used when training exp6/exp8, so it cannot fairly rank all four models.

Model identities

Arm SHA-256
Original Muharaf 726052869fa5aaf253d906d1c228e87bf80203a0cdcc8c3de50873d2697df178
exp6 3d6d3c3df380355e2390db33c9c981487e2d3660deb928da61a352902bc6a79b
exp8 d13c0996c14dbf805eeeeb4c718663b9686055863721b326e1cfe5fdeef82707
exp9 2896fef9d9665cbb82fba8faa3bf0c628ac6a5cc62eb678f40c707db833aebea

Full hashes and machine-readable aggregate scores are stored in benchmark_summary.json. The original local evaluation artifact is not included because it contains workstation paths.

This diagnostic re-evaluates the original Muharaf recognizer, exp6, exp8, and exp9 on the same guard lists and raw references. These guards are development data, so the table supports domain diagnosis and protocol-matched comparison, not an independent final-generalization claim.

Claims that are allowed

  • exp9 outperformed exp8 on the sealed Agapet and Omar sets.
  • exp9 regressed slightly on sealed TariMa.
  • the same-protocol guard table compares the four listed model files fairly on those exact guard inputs.
  • Athar provides system-level functionality that a recognizer weight file alone does not provide.
  • under the supplied server protocol, Phoenix had lower CER on 3/4 datasets, lower unweighted macro CER, and lower WER than Baseer__Nakba while being roughly 752× smaller by parameter count.
  • Baseer__Nakba had lower Omar CER and higher exact-line accuracy in the same supplied comparison.

Claims that are not allowed

  • “exp9 beats the published 18.1% Muharaf result” without reproducing its official protocol.
  • “best Arabic HTR model”.
  • comparing Arabic-normalized CER with another system's raw CER.
  • treating guard-set results as unseen final-test evidence.
  • attributing Athar's retrieval, review, or export capabilities to the exp9 weights.
  • calling the Baseer comparison independently reproduced, sealed, or universally representative until its raw predictions, counts, preprocessing, and evaluation code are archived.

Historical development results

These rows document the project's progression, but they use different evaluation sets. They must not be read as a single head-to-head leaderboard.

Model Evaluation set LM CER WER Character accuracy Word accuracy
Muharaf baseline RASAM-test No 37.91% 88.12% 62.09% 11.88%
exp4A RASAM-test No 7.64% 27.77% 92.36% 72.23%
exp6 RASAM-test No 7.48% 27.17% 92.52% 72.83%
exp6 + order-8 LM RASAM-test Yes 6.76% 27.17% 93.24% 72.83%
exp6 + order-8 LM TariMa-test Yes 8.68% 37.84% 91.32% 62.16%
exp7_easy Easy-Family No 5.94% 22.38% 94.06% 77.62%
exp7_hard Hard-Family No 13.70% 45.61% 86.30% 54.39%

Historical assistance-layer measurements

Component Measurement Result
Confidence-gated LM useful interventions on auto-seg run 12/18 (67%)
Contextual candidates difficult lines improved / harmed 6/10 / 0
Catalog search correct work ranked first 5/5
Retrieval correction Qur'an al-ʿAlaq match with reference 98.3%

Historical logic-domain LM perplexity was 53.14 for the general model, 40.70 for the self-trained model (−23.4%), and 20.25 for the mixed model (−61.9%). These are language-model diagnostics, not exp9 visual-recognizer CER results.