Kraken baseline segmenter for Cairo Genizah fragments (fine-tuned from BLLA 2026)

πŸ“„ Cairo Genizah AI Transcription β€” formal analysis and write-up of the project The full account of the methods, training data, benchmarks and error analysis behind the Cairo Genizah AI transcription models. Search the transcribed corpus at Cairo Genizah AI.

A kraken 7 baseline-segmentation model fine-tuned on Cairo Genizah manuscript pages β€” the first Genizah-specific segmenter we know of that is publicly available. Recognition models for these hands exist (the MiDRASH project published its recognizers on Zenodo), but line segmentation is where a generic model loses most of the text on Genizah fragments: tight columns, damaged edges, script far larger than the segmenter was trained on, marginal glosses and rotated scans. This model closes a large part of that gap.

  • Base: kraken's BLLA base model (Benjamin Kiessling, ALMAnaCH/Inria, 2026, 10.5281/zenodo.22879549, Apache-2.0), a general print-and-handwriting segmenter trained on Latin, Arabic and Greek script corpora.
  • Fine-tuning data: 2,000 pages of Hebrew, Aramaic and Judaeo-Arabic literary manuscripts from the Cairo Genizah, drawn from the KTIV project of the National Library of Israel (at most 3 pages per manuscript; 112 manuscripts held out for validation; manuscripts that overlap our evaluation benchmarks removed). Line targets were derived from the KTIV word coordinates with an ink-estimated baseline; writing KTIV does not transcribe was painted out so that no text is taught as background.
  • Files: ktivseg_e10.safetensors (the released model, epoch 10), plus the epoch-3 and epoch-9 checkpoints for comparison, and views.py, the test-time multi-view wrapper described below.

Results

Character error rate of the same recognizer (MiDRASH Gen_01) reading the lines each segmenter finds, on two held-out benchmarks: 140 religious (literary) pages with KTIV ground truth, and 131 documentary fragments from the Princeton Geniza Project. Lower is better; medians over pages.

segmenter religious CER two-column pages single-column pages documentary CER
kraken 4 default 0.680 0.504 0.654 0.358
BLLA 2026, binarised input 0.587 0.399 0.587 0.297
BLLA 2026 + multi-view 0.462 0.299 0.404 0.272
this model (RGB input) 0.409 0.314 0.461 0.264
this model + multi-view 0.336 0.248 0.361 0.253

For scale, MiDRASH's own published reads of the matched religious pages score 0.240 β€” so with multi-view this open model brings a stock kraken pipeline to within a tenth of a CER point of that, from 0.68.

In a two-reader consensus pipeline (kraken read vs. a vision-language-model read of the same page, 89 documentary pages), the share of lines on which both readers agree rose from 11.1 % (kraken 4) and 12.8 % (BLLA 2026) to 16.6 % with this model plus multi-view.

Training: val_bl_f1 0.935 (base) β†’ 0.952 by epoch 4, flat through epoch 10; we stopped there and released epoch 10, the best of epochs 3/9/10 on both benchmarks.

Usage

Standard kraken 7 CLI (segment the RGB page, not a binarised one β€” the model was fine-tuned on colour pages):

pip install "kraken>=7.0"
kraken -i page.jpg lines.json segment -bl -i ktivseg_e10.safetensors
kraken -i page.jpg out.txt segment -bl -i ktivseg_e10.safetensors ocr -m <recognition_model>.mlmodel

Python:

from PIL import Image
from kraken import blla
from kraken.lib import vgsl

net = vgsl.TorchVGSLModel.load_model("ktivseg_e10.safetensors")
seg = blla.segment(Image.open("page.jpg").convert("RGB"), model=net)   # baselines + polygons

Multi-view reading (recommended for large script and rotated scans)

BLLA resizes every page to a fixed input height (1,800 px), so the size of the writing at the model input depends only on how tall the page is relative to its text. A third of the religious benchmark reaches the model with writing larger than 97 % of the training pages, and recovery collapses there. views.py builds views of a page β€” the page scaled to the model height and padded at the bottom so the text shrinks by a factor 0.5 or 0.33, plus 90Β°/270Β° rotations for vertically running scans β€” segments each view, maps the geometry back to full resolution, and lets the best-reading view win (recognition always runs on the full-resolution page). It is pure PIL geometry with no kraken dependency; the production wiring is in src/services/kraken_microservice/k7/ of the code repository below.

Limitations

  • Trained on literary Genizah hands via KTIV; documentary hands (letters, legal deeds) improve too but were not in the fine-tuning set.
  • Like the base model it detects baselines rather than top-lines for Hebrew; recognition models trained on kraken baselines (MiDRASH Gen_01 included) expect exactly that.
  • Pages wider than 2,600 px were excluded from training for memory reasons; very large scans are best read through the multi-view wrapper.
  • The KTIV images and transcriptions used for fine-tuning are not redistributed here; only the weights are.

Provenance and credits

  • Base model: Kiessling, B. (2026). BLLA base model. Zenodo. https://doi.org/10.5281/zenodo.22879549 (Apache-2.0). kraken: https://kraken.re
  • Fine-tuning pages: National Library of Israel, KTIV β€” the International Collection of Digitized Hebrew Manuscripts (images and transcriptions). Documentary evaluation: the Princeton Geniza Project.
  • Recognizer used in all evaluations: MiDRASH Gen_01 (Zenodo).
  • Training recipe, PageXML exporter, evaluation harness and the two-reader pipeline: https://github.com/AIStream-Peelout/historical-document-analysis (src/finetuning/kraken/segtrain_ktiv.sh, export_ktiv_pagexml.py; evaluation write-up in docs/kraken7_segmenter_ab.md).

If you use this model, please cite the base model above and credit NLI KTIV for the manuscripts.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support