dfine-det-large-baseline-hebrew-samaritan-stage1
Stage-1 product fine-tune of dfine-det — D-FINE Large (HGNetv2-B4) predicting text-line baselines as polylines (B-spline control points) for medieval Hebrew and Samaritan manuscripts (PAGE-XML).
Hub repo:
johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1
Primary file:best_cbad_f1.safetensors(~116 MB)
SHA256:ef7c32254905d72f1768ba2043cf092b5b040e5643fba83fba2bbbc13b70643d
Init chain: Peterande/D-FINE dfine_l_obj2coco → johnlockejrr/dfine-det-large-baseline-stage0 (cbad_f1_max ≈ 0.893) → this Stage-1.
Model summary
| Architecture | PolylineDFINE — D-FINE detection core + polyline head |
| Size class | Large (D-FINE-L / backbone B4) |
| Parameters | ~30.1 M trainable |
| Queries | 300 |
| Geometry | (K=8) cubic B-spline control points per line + height |
| Canvas | (1280\times1280) letterbox |
| Init | Stage-0 best_cbad_f1.safetensors (strict=False load path) |
| Task | Document baseline / text-line detection → PAGE or ALTO export |
| Scripts | Hebrew (square / medieval) + Samaritan |
| Recommended conf | 0.50 (locked on unseen Hebrew holdout) |
| Reading order (export) | Prefer --reading-order rtl for PAGE/ALTO serialize |
Detection is set prediction of scored polylines. Polygonal line environments (BLLA) are applied only at PAGE/ALTO serialize time, not in the training objective.
Intended use
Use for
- Baseline detection on Hebrew and Samaritan manuscript pages (PAGE-XML workflows)
- Downstream HTR / layout pipelines that need reliable text-line geometry
- Research on Stage-0 → Stage-1 transfer for script-specific corpora
Not for
- Official ICDAR cBAD 2019 bake-offs (wrong track / domain)
- Production OCR transcription (detects lines; does not recognize text)
- Replacing layout region detectors (paragraphs, tables, etc.)
Training pipeline
Peterande D-FINE-L (Objects365→COCO)
│
▼
Stage-0 multiscript polyline pretrain
(~46.5k train / 1.9k val; cbad_f1_max ≈ 0.893 @ ~ep39)
Hub: johnlockejrr/dfine-det-large-baseline-stage0
│
▼
Stage-1 Hebrew + Samaritan fine-tune ← this release
(1758 train / 261 val; monitor peak cbad_f1_max ≈ 0.9599 @ ep109)
Config: configs/hebrew_samaritan_stage1.yaml.
Fine-tuning data (Stage-1)
PAGE-XML baselines from:
| Corpus | Role | Pages (train / val) |
|---|---|---|
Hebrew_Medieval-seg |
Square / medieval Hebrew | 333 / 37 |
sam_44_mss_pango_additional |
Samaritan (44+ MSS pack) | 1425 / 224 |
| Combined | 1758 / 261 |
Split policy (seed 42; see configs/hebrew_samaritan_combined_split.json):
- Samaritan: manuscript-level holdout (~12% pages target) — 43 train MSS / 7 val MSS (no page from a val MS appears in train).
- Hebrew: page-level split within
Hebrew_Medieval-seg.
Compiled to Arrow (simplify_eps=0.01 → B-spline (K=8)) under outputs/hebrew_samaritan_combined/arrow/.
Training recipe (Stage-1)
| Hyperparameter | Value |
|---|---|
| Init weights | Stage-0 PRETRAIN/stage0/best_cbad_f1.safetensors |
| Optimizer | AdamW; base LR (5\times10^{-5}) (backbone (0.1\times)) |
| Schedule | Linear warmup + cosine (Lightning defaults as in dfine-det train) |
| Precision | bf16-mixed |
| Effective batch | 8 (micro-batch 8 × accum 1 on this run) |
| Cap / early-stop | max 150 epochs; quit: early on cbad_f1_max; patience 20; min_epochs 15; min_delta 0.0005 |
| Best snapshot | epoch 109 (checkpoints/best-epoch=109.ckpt → best_cbad_f1.safetensors) |
| Monitor peak | cbad_f1_max ≈ 0.9599 |
| Conf sweep (train) | {0.1, 0.2, 0.3, 0.4, 0.5, 0.6}; fast_conf_sweep: true |
| Losses / matcher | loss_class=8.0, loss_height=3.0, cost_y=3.0 |
| Augment | Enabled (--augment / YAML augment: true) |
| Match distance | 20 px on the 1280 canvas |
| Seed | 42 |
Evaluation
Metric: cBAD-style line F1 — Hungarian matching of densified polylines with mean bidirectional Chamfer cost; TP if cost ≤ dist_thresh (20 px).
Operating point (recommended)
Lock conf = 0.50 using the unseen Hebrew holdout (not the Stage-1 val peak). On Stage-1 val, 0.50 is within ~0.0015 F1 of the val-best conf (0.40).
1) Unseen Hebrew holdout (primary report)
Source: /opt/DATA/DATASETS/multi-export-hebrew (674 PAGE pages).
Excluded any page already in Stage-1 train/val (exact basename) or Stage-0 lists (exact / post-__ stem) → 377 pages.
List: configs/multi_export_hebrew_unseen_test.lst.
dfine-det --config configs/hebrew_samaritan_stage1.yaml -d cuda:0 test \
--weights best_cbad_f1.safetensors \
--val-arrow outputs/hebrew_samaritan_stage1/arrow_unseen_test/val.arrow \
--conf 0.50
| conf | cbad_f1 | precision | recall | mean_chamfer |
|---|---|---|---|---|
| 0.10 | 0.7369 | 0.6010 | 0.9523 | 5.62 |
| 0.20 | 0.8750 | 0.8288 | 0.9268 | 6.51 |
| 0.30 | 0.8901 | 0.8851 | 0.8951 | 7.86 |
| 0.40 | 0.8901 | 0.9070 | 0.8738 | 7.14 |
| 0.50 | 0.8977 | 0.9321 | 0.8657 | 5.91 |
| 0.60 | 0.8935 | 0.9464 | 0.8462 | 5.18 |
Headline (unseen @ 0.50): F1 0.898 · P 0.932 · R 0.866
2) Stage-1 in-domain val (early-stop split)
261 pages (hebrew_samaritan_combined val Arrow). Optimistic vs true holdout.
| conf | cbad_f1 | precision | recall | mean_chamfer |
|---|---|---|---|---|
| 0.10 | 0.7937 | 0.6623 | 0.9902 | 3.07 |
| 0.20 | 0.9317 | 0.8928 | 0.9741 | 4.30 |
| 0.30 | 0.9494 | 0.9394 | 0.9597 | 5.10 |
| 0.40 | 0.9553 | 0.9602 | 0.9504 | 4.76 |
| 0.50 | 0.9538 | 0.9718 | 0.9364 | 4.14 |
| 0.60 | 0.9503 | 0.9814 | 0.9211 | 3.55 |
Train monitor cbad_f1_max ≈ 0.9599 @ ep109 (best across the same conf grid during training).
Caveats
- Unseen set is Hebrew multi-export pages not in Stage-0/1 lists; it is not a Samaritan-only test and not manuscript-disjoint vs every possible Hebrew source.
- Stage-1 val F1 is not interchangeable with unseen F1 (~0.95 vs ~0.90).
- Fixed conf 0.1 (train log
cbad_f1) is far below operating F1 — always sweep or use 0.50.
How to use
Install
pip install -e ".[dev]" # from the dfine-det repo; see QUICKSTART for torch/CUDA notes
Inference (PAGE XML)
dfine-det -d cuda:0 infer \
--weights best_cbad_f1.safetensors \
--image page.jpg \
-o page.xml \
--format page \
--conf 0.5 \
--reading-order rtl
Evaluate
dfine-det --config configs/hebrew_samaritan_stage1.yaml -d cuda:0 test \
--weights best_cbad_f1.safetensors \
--val-arrow path/to/val.arrow \
--conf 0.50
Files in this release
| File | Description |
|---|---|
best_cbad_f1.safetensors |
Recommended Stage-1 weights (monitor peak ≈ 0.9599) |
best_0.9599.safetensors |
Same snapshot, score-tagged filename |
checkpoints/best-epoch=109.ckpt |
Optional Lightning resume state |
README.md |
This model card |
Upload the safetensors (+ this README) for Hub users; Lightning .ckpt is optional.
Limitations
- Domain gap remains between Stage-1 val (
0.95 F1) and unseen Hebrew (0.90 F1). - (K=8) control points underfit strongly curved / damaged lines.
- Vertical “ghost” doubles can still appear on some layouts;
cost_y/loss_heightmitigate but do not eliminate them. - Private / institutional manuscript images used in training are not redistributed with the weights; respect each corpus license.
- Do not cite these scores as cBAD 2019 / Orli bake-off results.
Citation & credits
D-FINE (architecture & COCO/Objects365 pretrain)
@inproceedings{peng2025dfine,
title = {D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement},
author = {Peng, Yansong and Li, Hebei and Wu, Peixi and Zhang, Yueyi and Sun, Xiaoyan and Wu, Feng},
booktitle = {The Thirteenth International Conference on Learning Representations},
year = {2025},
url = {https://arxiv.org/abs/2410.13842}
}
- Code: github.com/Peterande/D-FINE
- Weights: huggingface.co/Peterande/D-FINE
Stage-0 polyline pretrain
@software{dfine_det_stage0_large,
title = {dfine-det Large Stage-0: Multiscript Polyline Baseline Pretrain},
author = {John Locke Jrr, johnlockejrr},
year = {2026},
url = {https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0}
}
This Stage-1 Hebrew/Samaritan fine-tune
@software{dfine_det_hebsam_stage1,
title = {dfine-det Large Stage-1: Hebrew and Samaritan Polyline Baseline Detection},
author = {John Locke Jrr, johnlockejrr},
year = {2026},
note = {Fine-tuned from johnlockejrr/dfine-det-large-baseline-stage0; operating conf 0.50},
url = {https://huggingface.co/johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1}
}
Additional notices
- Vendored D-FINE modules: see
THIRD_PARTY_NOTICES.md(Apache-2.0, © 2024 The D-FINE Authors). - PAGE polygonalization helpers derive from Kraken/BLLA (Apache-2.0, © Benjamin Kiessling) — used at export only.
- Training corpora: follow each dataset’s original license / institutional terms.
License
Apache License 2.0 for the dfine-det code and these weights, consistent with D-FINE’s Apache-2.0 release. Downstream users must also comply with licenses of any datasets used in further fine-tuning.
Model tree for johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1
Base model
Peterande/D-FINEPaper for johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1
Evaluation results
- cbad_f1 @ conf=0.50 on multi-export-hebrew unseen holdout (excl. Stage-0/1 pages)test set self-reported0.898
- Precision @ conf=0.50 on multi-export-hebrew unseen holdout (excl. Stage-0/1 pages)test set self-reported0.932
- Recall @ conf=0.50 on multi-export-hebrew unseen holdout (excl. Stage-0/1 pages)test set self-reported0.866
- cbad_f1_max (train monitor / conf sweep) on Hebrew Medieval + Samaritan Stage-1 valvalidation set self-reported0.960
- cbad_f1 @ conf=0.40 (post-hoc sweep) on Hebrew Medieval + Samaritan Stage-1 valvalidation set self-reported0.955
- cbad_f1 @ conf=0.50 (locked operating point) on Hebrew Medieval + Samaritan Stage-1 valvalidation set self-reported0.954