Kraken baseline segmenter for Cairo Genizah fragments (fine-tuned from BLLA 2026)
π Cairo Genizah AI Transcription β formal analysis and write-up of the project The full account of the methods, training data, benchmarks and error analysis behind the Cairo Genizah AI transcription models. Search the transcribed corpus at Cairo Genizah AI.
A kraken 7 baseline-segmentation model fine-tuned on Cairo Genizah manuscript pages β the first Genizah-specific segmenter we know of that is publicly available. Recognition models for these hands exist (the MiDRASH project published its recognizers on Zenodo), but line segmentation is where a generic model loses most of the text on Genizah fragments: tight columns, damaged edges, script far larger than the segmenter was trained on, marginal glosses and rotated scans. This model closes a large part of that gap.
- Base: kraken's BLLA base model (Benjamin Kiessling, ALMAnaCH/Inria, 2026, 10.5281/zenodo.22879549, Apache-2.0), a general print-and-handwriting segmenter trained on Latin, Arabic and Greek script corpora.
- Fine-tuning data: 2,000 pages of Hebrew, Aramaic and Judaeo-Arabic literary manuscripts from the Cairo Genizah, drawn from the KTIV project of the National Library of Israel (at most 3 pages per manuscript; 112 manuscripts held out for validation; manuscripts that overlap our evaluation benchmarks removed). Line targets were derived from the KTIV word coordinates with an ink-estimated baseline; writing KTIV does not transcribe was painted out so that no text is taught as background.
- Files:
ktivseg_e10.safetensors(the released model, epoch 10), plus the epoch-3 and epoch-9 checkpoints for comparison, andviews.py, the test-time multi-view wrapper described below.
Results
Character error rate of the same recognizer (MiDRASH Gen_01) reading the lines each segmenter finds,
on two held-out benchmarks: 140 religious (literary) pages with KTIV ground truth, and 131 documentary
fragments from the Princeton Geniza Project. Lower is better; medians over pages.
| segmenter | religious CER | two-column pages | single-column pages | documentary CER |
|---|---|---|---|---|
| kraken 4 default | 0.680 | 0.504 | 0.654 | 0.358 |
| BLLA 2026, binarised input | 0.587 | 0.399 | 0.587 | 0.297 |
| BLLA 2026 + multi-view | 0.462 | 0.299 | 0.404 | 0.272 |
| this model (RGB input) | 0.409 | 0.314 | 0.461 | 0.264 |
| this model + multi-view | 0.336 | 0.248 | 0.361 | 0.253 |
For scale, MiDRASH's own published reads of the matched religious pages score 0.240 β so with multi-view this open model brings a stock kraken pipeline to within a tenth of a CER point of that, from 0.68.
In a two-reader consensus pipeline (kraken read vs. a vision-language-model read of the same page, 89 documentary pages), the share of lines on which both readers agree rose from 11.1 % (kraken 4) and 12.8 % (BLLA 2026) to 16.6 % with this model plus multi-view.
Training: val_bl_f1 0.935 (base) β 0.952 by epoch 4, flat through epoch 10; we stopped there and released
epoch 10, the best of epochs 3/9/10 on both benchmarks.
Usage
Standard kraken 7 CLI (segment the RGB page, not a binarised one β the model was fine-tuned on colour pages):
pip install "kraken>=7.0"
kraken -i page.jpg lines.json segment -bl -i ktivseg_e10.safetensors
kraken -i page.jpg out.txt segment -bl -i ktivseg_e10.safetensors ocr -m <recognition_model>.mlmodel
Python:
from PIL import Image
from kraken import blla
from kraken.lib import vgsl
net = vgsl.TorchVGSLModel.load_model("ktivseg_e10.safetensors")
seg = blla.segment(Image.open("page.jpg").convert("RGB"), model=net) # baselines + polygons
Multi-view reading (recommended for large script and rotated scans)
BLLA resizes every page to a fixed input height (1,800 px), so the size of the writing at the model input
depends only on how tall the page is relative to its text. A third of the religious benchmark reaches the
model with writing larger than 97 % of the training pages, and recovery collapses there. views.py builds
views of a page β the page scaled to the model height and padded at the bottom so the text shrinks by a
factor 0.5 or 0.33, plus 90Β°/270Β° rotations for vertically running scans β segments each view, maps the
geometry back to full resolution, and lets the best-reading view win (recognition always runs on the
full-resolution page). It is pure PIL geometry with no kraken dependency; the production wiring is in
src/services/kraken_microservice/k7/ of the code repository below.
Limitations
- Trained on literary Genizah hands via KTIV; documentary hands (letters, legal deeds) improve too but were not in the fine-tuning set.
- Like the base model it detects baselines rather than top-lines for Hebrew; recognition models trained on
kraken baselines (MiDRASH
Gen_01included) expect exactly that. - Pages wider than 2,600 px were excluded from training for memory reasons; very large scans are best read through the multi-view wrapper.
- The KTIV images and transcriptions used for fine-tuning are not redistributed here; only the weights are.
Provenance and credits
- Base model: Kiessling, B. (2026). BLLA base model. Zenodo. https://doi.org/10.5281/zenodo.22879549 (Apache-2.0). kraken: https://kraken.re
- Fine-tuning pages: National Library of Israel, KTIV β the International Collection of Digitized Hebrew Manuscripts (images and transcriptions). Documentary evaluation: the Princeton Geniza Project.
- Recognizer used in all evaluations: MiDRASH
Gen_01(Zenodo). - Training recipe, PageXML exporter, evaluation harness and the two-reader pipeline:
https://github.com/AIStream-Peelout/historical-document-analysis
(
src/finetuning/kraken/segtrain_ktiv.sh,export_ktiv_pagexml.py; evaluation write-up indocs/kraken7_segmenter_ab.md).
If you use this model, please cite the base model above and credit NLI KTIV for the manuscripts.