File size: 17,216 Bytes
af3342d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
# أثر — An Assistive Hybrid Pipeline for Arabic Manuscript HTR and Scholarly Review

*أثر*

---

## English Abstract

Historical Arabic manuscripts remain largely locked in page images: calligraphic variation, dense
marginalia, page degradation, and domain-specific terminology defeat generic OCR. We present
اثر, an **assistive hybrid pipeline** that couples layout segmentation and Kraken-based
Handwritten Text Recognition (HTR) with six transparent assistance layers: a **confidence-gated

character n-gram language model**, a **layout router** that separates framed logic/philosophy pages
from general pages, a **human-in-the-loop manual/PAGE-XML segmentation** mode, a **contextual LLM

reading assistant**, a **retrieval-augmented corrector** that aligns lines with published digital texts
(Qur'an, ḥadīth), and a **relevance-ranked manuscript catalog**. Building on open resources (Muharaf,
Kraken, eScriptorium), our strongest verified general model with an order-8 language model reaches
**6.76% CER** on RASAM-test and **8.68% CER** on TariMa-test. The gated language model improves
CER on three separated datasets without harming clean pages (e.g., 4.34%→3.29% on the demo set),
and a closed-loop **self-training** experiment reduces logic-domain language-model perplexity by
**62%** (53.14→20.25). The contribution is a reproducible, review-oriented pipeline that keeps the
original transcription intact and surfaces uncertainty rather than silently overwriting it.

**Keywords:** Arabic manuscript HTR; confidence-gated language model; layout routing;
retrieval-augmented correction; human-in-the-loop.

---

## الخلاصة العربية


تبقى المخطوطات العربية التاريخية حبيسة الصور؛ فتنوع الخطوط، وكثافة الحواشي، وتلف الصفحات، والمصطلحات المتخصصة تجعل أنظمة OCR العامة غير كافية. نقدم اثر، وهو خط معالجة هجين مساعد يجمع بين تجزئة التخطيط والتعرف على الخط اليدوي (HTR) بمحرك Kraken، مع ست طبقات مساعدة شفافة: نموذج لغوي حرفي موجه بالثقة، وراوتر تخطيط يميز صفحات المنطق والفلسفة المؤطرة عن الصفحات العامة، ووضع قص يدوي واستيراد PAGE-XML بمشاركة الإنسان، ومساعد قراءة سياقي بنموذج لغوي كبير، ومصحح بالاسترجاع يطابق الأسطر مع نصوص رقمية منشورة، وفهرس مخطوطات مرتب بحسب الصلة. يحقق أقوى نموذج عام متحقق منه معدل خطأ حرفي 6.76% على RASAM-test و 8.68% على TariMa-test عند استخدام نموذج لغوي من الرتبة الثامنة. كما تحسن البوابة اللغوية النتائج على ثلاث مجموعات منفصلة دون الإضرار بالصفحات النظيفة، وتخفض تجربة التدريب الذاتي حيرة النموذج اللغوي في مجال المنطق بنسبة 62%. يهدف النظام إلى تسريع مراجعة المخطوطات مع حفظ النص الأصلي وإظهار مواضع الشك بدلا من استبدالها بصمت.

**الكلمات المفتاحية:** تعرّف الخط اليدوي العربي؛ نموذج لغوي مُبوَّب؛ راوتر تخطيط؛ التصحيح بالاسترجاع؛
مشاركة الإنسان.

---

## 1. Introduction

Digitizing Arabic manuscripts as images does not make them readable, searchable, or analyzable.
Scholarly review needs more than a raw transcript: it needs uncertainty flagging, alternative
readings, term explanation, and cross-referencing with known texts. Generic OCR fails on
handwritten Arabic because a single base shape maps to several letters depending on dots and
scribal convention, and because dense commentary and frames break naïve reading order.

اثر does not start from zero. It builds on open foundations — Muharaf handwritten-Arabic
resources [1], eScriptorium [2], and Kraken [3] — and extends them into an **assistive hybrid

pipeline for recognition *and* review together**. Our contribution is at the engineering and
experimental level: (i) domain fine-tuning that cuts CER from a 37.91% generic baseline to 6.76%;
(ii) a **confidence-gated** language-model decoder that helps hard lines without harming clean ones;
(iii) six transparent assistance layers added around — never replacing — the recognizer; and
(iv) two research probes (retrieval-augmented correction and closed-loop self-training) validated on
real pages.

## 2. Project Problem

Historical Arabic manuscripts exhibit calligraphic variation (Naskh, Maghrebi, Eastern), page
degradation, dense marginalia around framed main text, non-standard orthography, and
domain-specific vocabulary (logic, philosophy, jurisprudence, theology) where one uncertain word can
change an argument. Off-the-shelf recognizers both mis-read these pages and, critically, hide their
uncertainty — unsafe for scholarship.

## 3. Objectives

1. Reduce CER on heterogeneous Arabic manuscripts through data preparation and domain fine-tuning.
2. Add a language layer that helps difficult lines **without degrading easy ones**.
3. Handle complex framed layouts (main text vs. marginalia) via routing and human-in-the-loop tools.
4. Surface uncertainty and provide transparent, non-destructive reading assistance.
5. Keep the whole pipeline reproducible and honest — every claim tied to a measured result.

## 4. Proposed Solution — Hybrid Architecture

The system is a staged pipeline (Figure 1): conservative preprocessing → **layout routing** →
segmentation → HTR recognition → **confidence-gated LM decoding** → a **transparent post-OCR

assistance layer** (candidate readings, retrieval-augmented correction, explanation, catalog search),
plus a **human-in-the-loop** manual/PAGE-XML segmentation path for the hardest pages.

Unlike a plain OCR tool, the assistance layers are *additive and reversible*: the original recognizer
output is preserved, and every suggestion is a labeled annotation the reviewer accepts by a click.

## 5. Innovation and Originality

- **Confidence-gated shallow fusion:** the LM is applied *only* to low-confidence lines, with a
  large-change guard that rejects and flags aggressive rewrites (`uncertain`) — combining the gain of
  language modeling with do-no-harm on clean text.
- **Layout router with frame detection:** a vertical-rule fill test (`cf>0.92`) routes only genuinely
  framed logic pages to the specialized model, fixing a false-positive that had destroyed a normal
  page (58.3%→4.6% CER).
- **Retrieval-augmented correction (RAC):** manuscript lines are aligned to published digital texts;
  Qur'anic/ḥadīth quotations are corrected against the canonical source with a reference.
- **Closed-loop self-training:** the model transcribes more of its own manuscript, high-confidence
  lines become pseudo-labels in the decoder's own encoding, improving the domain LM by 62% perplexity.
- **Human-in-the-loop:** users import eScriptorium PAGE-XML or draw line boxes; column-based
  classification then separates main text from marginalia.

## 6. Methodology

**Preprocessing (safe path).** Recognition models operate on **grayscale (L)** images; binarization
(nlbin), enhancement, resizing, and deskew are **disabled by default**, matching the eScriptorium
setting, because binarizing an L-model input distorts strokes.

**Segmentation & routing.** A fine-tuned baseline segmentation model (`seg_finetuned_v4`, trained on
319 RASAM pages) produces HTR-compatible baselines. A router inspects ruling geometry and density;
only pages with a confirmed rectangular frame use the specialized logic segmentation/recognition
models, otherwise the general models are used.

**Recognition.** Kraken-compatible CTC HTR models (`transcription_finetuned_best` general;
`logic_philosophy_v2_rec` logic). Fine-tuning used a warmup/freeze recipe (`--warmup 2`,
`--freeze-backbone 2`) protecting the visual backbone while sequence layers adapt.

**Gated LM decoding.** An offline character-level 8-gram LM is combined with acoustic confidence
during beam search:

$$\mathrm{Score}(y\mid x) = \log P_{\mathrm{HTR}}(y\mid x) + \alpha \log P_{\mathrm{LM}}(y) + \beta\,|y| \qquad (1)$$

with $\alpha{=}0.5,\ \beta{=}0.3$, beam $=10$. A per-line **confidence gate** (mean character
probability): if confidence $\ge 0.95$ keep greedy; else apply the LM; if the LM changes $>50\%$ of
the line, reject and flag `uncertain`.

**Post-OCR assistance.** Candidate readings (visual beam + contextual LLM), retrieval-augmented
correction, term/page explanation, and catalog search are stored as separate annotations.

**Evaluation.** CER (primary) and WER on frozen splits; the test subset is never used for training
decisions. Three page sets are kept **separate and never mixed**: GT-segmented, manual/GT-demo, and
automatic-segmentation.

## 7. Tools and Technologies

Kraken (HTR engine) · eScriptorium (annotation/PAGE-XML) · PyTorch on **NVIDIA RTX 4050 (CUDA + FP16)** ·
FastAPI backend + PWA frontend · offline character n-gram LM (shallow fusion) · OpenAI GPT-5 family for
the contextual assistant (`gpt-5.4-mini` light tasks, `gpt-5.5` line correction) · OpenITI/Tanzil digital
texts for retrieval. No hardware-speed claims are made in the evidence tables.

## 8. Results / Practical Demonstration

### 8.1 Recognition model comparison (Verified, `metrics_summary.csv`)



The pipeline is a *progression*: a weak generic baseline improved by domain training, mixed-domain

data, and LM rescoring (Table 1, Figure 2).



**Table 1. Model performance (source: `metrics_summary.csv`).**



| Model | Dataset | LM | CER % | WER % | Char Acc % | Word Acc % |

|---|---|---|---:|---:|---:|---:|

| muharaf_rec_best (baseline) | RASAM-test | No | 37.91 | 88.12 | 62.09 | 11.88 |

| exp4A (RASAM-trained) | RASAM-test | No | 7.64 | 27.77 | 92.36 | 72.23 |

| exp6 (RASAM+TariMa general) | RASAM-test | No | 7.48 | 27.17 | 92.52 | 72.83 |

| **exp6 + LM (order-8)** | RASAM-test | Yes | **6.76** | 27.17 | 93.24 | 72.83 |

| exp6 + LM (order-8) | TariMa-test | Yes | 8.68 | 37.84 | 91.32 | 62.16 |

| exp7_easy (specialized) | Easy-Family-test | No | 5.94 | 22.38 | 94.06 | 77.62 |
| exp7_hard (specialized) | Hard-Family-test | No | 13.70 | 45.61 | 86.30 | 54.39 |



Domain fine-tuning cuts CER **37.91%→7.48%**; the order-8 LM adds a further **7.48→6.76%** on

RASAM and **10.27→8.68%** on TariMa (zero-shot domain gap).



### 8.2 Confidence-gated LM — three separated datasets (source: `LM_CONFIDENCE_GATE_REPORT.md`)

**Table 2. Gated LM (threshold <0.95) vs. greedy vs. always-LM (CER %).**

| Dataset | Greedy | Always-LM | **Gated (<0.95)** |
|---|---:|---:|---:|
| D1 clean (GT-seg) | 0.39 | 0.56 | **0.30** |
| D2 demo (GT/manual) | 4.34 | 2.56 | **3.29** |
| D3 auto-seg (v4) | 3.61 | 2.28 | **3.41** |
| Mean (3 sets) | 2.78 | 1.80 | **2.33** |

Always-LM helps hard sets but **harms the clean set** (0.39→0.56). The gate captures most of the gain
while *improving* even the clean set (0.39→0.30) and lowering WER (D2 15.30→10.48). This validates
gating as the safe operating point.

### 8.3 Full-pipeline per-page benchmark (source: `FEATURES_BENCHMARK.md`)



Independent run on 13 held-out pages (automatic segmentation, GPU). Normal pages average **7.63% CER**;

the two 1926-family pages (a known hard automatic-segmentation layout) average 27.30% and drag the

overall mean to 10.65%. During this benchmark we found and fixed a router false-positive (a density

detector had sent normal pages to the logic path, causing a full failure on 1926_0077).

**Table 3. Selected per-page CER (full pipeline, auto-seg).**

| Page | CER % | WER % |
|---|---:|---:|
| 1983_177819 (Eastern) | 1.99 | 7.4 |

| 1977_0136 | 2.92 | 12.7 |
| 417_0011 (after router fix) | 4.61 | 22.5 |

| 1944_0006 | 8.36 | 15.4 |
| 1926_0144 (hard layout) | 29.81 | 54.3 |



Assistance-layer measurements on the same run: gated-LM interventions **12/18 improved** (67% useful,

1 harmful); contextual candidate suggestions **6/10 improved, 0 harmful** on the hardest low-confidence

lines; catalog search returned the correct work first on **5/5** queries after relevance re-ranking.



### 8.4 Self-training the logic LM (source: `RESEARCH_SELF_TRAINING.md`)



Transcribing 60 pages of a logic manuscript (GT pages excluded) yielded **418** high-confidence

(≥0.95) pseudo-labeled lines in the decoder's own visual encoding. A mixed LM (general + pseudo)

reduces perplexity on held-out logic GT by **62%** (Figure 3), confirming that the closed loop models

the domain far better; on the easy GT pages this does not yet move CER (greedy already at 4.21%), but

the mixed LM is consistently the least harmful — the gain is expected on harder/degraded pages.



**Table 4. Logic-domain LM perplexity (lower is better).**



| Language model | Perplexity | vs. general |

|---|---:|---:|

| General (char8_trainval) | 53.14 | — |
| Self-trained (logic only, 418 lines) | 40.70 | −23.4% |
| **Mixed (general + self-trained)** | **20.25** | **−61.9%** |

### 8.5 Retrieval-augmented correction (source: `docs/RETRIEVAL_AUGMENTED_CORRECTION.md`)

Aligning OCR lines to a digital store (Qur'an, Ṣaḥīḥ al-Bukhārī) recovers canonical text with a
reference: e.g., an embedded verse is identified as *Sūrat al-ʿAlaq 1* (98.3% match to the canonical verse),
and prophetic-biography lines match Bukhārī at 95–98%. A calibrated threshold (0.78) cleanly rejects
non-quoting text — no false matches.

## 9. Applications and Impact

The system turns page images into searchable, reviewable text with visible uncertainty: it accelerates
scholarly transcription of logic/philosophy/legal manuscripts, supports catalog-level identification of
works and authors, and — via retrieval — flags and corrects canonical quotations. Because assistance is
additive and reversible, it is safe for critical-edition workflows.

## 10. Challenges and Limitations

- Automatic segmentation still fails on the densest framed layouts (1926 family ≈27% CER); the
  human-in-the-loop PAGE-XML path recovers these (logic pages 32.9%→3.2–3.6% with GT crops).
- Character-level LM rescoring leaves some word-boundary (spacing) errors; measured spacing errors are
  ~16% of all character errors — a target for future word-boundary correction. `Not measured`: isolated
  effect of post-processing normalization; order-10 LM; end-to-end latency benchmark.
- The custom logic subset is 32 pages; inventory reports **1,000** complete lines while the prior paper
  cites **≈1,200** (author audit) — documented rather than silently reconciled.
- Results are validated only for the included datasets/splits; public-data provenance is preserved.

## 11. Conclusion

اثر is a disciplined, reproducible **assistive** HTR pipeline for Arabic manuscripts. The
strongest verified configuration reaches **6.76% CER** (RASAM) / **8.68%** (TariMa). Beyond
recognition, a confidence-gated language model, a layout router, human-in-the-loop segmentation, a
contextual reading assistant, retrieval-augmented correction, and a relevance-ranked catalog form a
transparent review system that preserves the manuscript text and surfaces uncertainty. Two research
probes — retrieval correction and closed-loop self-training (−62% perplexity) — show the architecture
can improve itself without additional hand-labeled data.

## References

[1] M. Saeed, A. Chan, A. Mijar, J. Moukarzel, G. Habchi, C. Younes, A. Elias, C.-W. Wong, A. Khater.
*Muharaf: Manuscripts of Handwritten Arabic Dataset for Cursive Text Recognition.* NeurIPS 2024
Datasets and Benchmarks Track, 2024.
[2] The eScriptorium Project. *eScriptorium: A Digital Text Production Pipeline for Print and

Handwritten Texts using Machine Learning Techniques.* Project documentation.
[3] B. Kiessling. *Kraken: a universal text recognizer for the humanities.* Journal of Open Source
Software, 4(42):1764, 2019. DOI: 10.21105/joss.01764.
[4] A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber. *Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks.* ICML, 2006.
[5] OpenITI and Tanzil Project. *Open Islamicate Texts Initiative and verified Qur'anic text resources.* Project documentation.

---
*Data ethics: the corpus combines public manuscript resources and project-curated material (≈1,200

author-transcribed lines; automated inventory 1,000). Public-source provenance remains visible; the

cleaning, organization, manual transcription, evaluation, and system integration are the project

contributions. No numbers in this paper are estimated — each is drawn from a committed result file or

marked `Not measured`.*