dfine-det-large-baseline-hebrew-samaritan-stage1

Stage-1 product fine-tune of dfine-det — D-FINE Large (HGNetv2-B4) predicting text-line baselines as polylines (B-spline control points) for medieval Hebrew and Samaritan manuscripts (PAGE-XML).

Hub repo: johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1
Primary file: best_cbad_f1.safetensors (~116 MB)
SHA256: ef7c32254905d72f1768ba2043cf092b5b040e5643fba83fba2bbbc13b70643d

Init chain: Peterande/D-FINE dfine_l_obj2coco → johnlockejrr/dfine-det-large-baseline-stage0 (cbad_f1_max ≈ 0.893) → this Stage-1.


Model summary

Architecture PolylineDFINE — D-FINE detection core + polyline head
Size class Large (D-FINE-L / backbone B4)
Parameters ~30.1 M trainable
Queries 300
Geometry (K=8) cubic B-spline control points per line + height
Canvas (1280\times1280) letterbox
Init Stage-0 best_cbad_f1.safetensors (strict=False load path)
Task Document baseline / text-line detection → PAGE or ALTO export
Scripts Hebrew (square / medieval) + Samaritan
Recommended conf 0.50 (locked on unseen Hebrew holdout)
Reading order (export) Prefer --reading-order rtl for PAGE/ALTO serialize

Detection is set prediction of scored polylines. Polygonal line environments (BLLA) are applied only at PAGE/ALTO serialize time, not in the training objective.


Intended use

Use for

  • Baseline detection on Hebrew and Samaritan manuscript pages (PAGE-XML workflows)
  • Downstream HTR / layout pipelines that need reliable text-line geometry
  • Research on Stage-0 → Stage-1 transfer for script-specific corpora

Not for

  • Official ICDAR cBAD 2019 bake-offs (wrong track / domain)
  • Production OCR transcription (detects lines; does not recognize text)
  • Replacing layout region detectors (paragraphs, tables, etc.)

Training pipeline

Peterande D-FINE-L (Objects365→COCO)
        │
        ▼
Stage-0 multiscript polyline pretrain
  (~46.5k train / 1.9k val; cbad_f1_max ≈ 0.893 @ ~ep39)
  Hub: johnlockejrr/dfine-det-large-baseline-stage0
        │
        ▼
Stage-1 Hebrew + Samaritan fine-tune  ← this release
  (1758 train / 261 val; monitor peak cbad_f1_max ≈ 0.9599 @ ep109)

Config: configs/hebrew_samaritan_stage1.yaml.


Fine-tuning data (Stage-1)

PAGE-XML baselines from:

Corpus Role Pages (train / val)
Hebrew_Medieval-seg Square / medieval Hebrew 333 / 37
sam_44_mss_pango_additional Samaritan (44+ MSS pack) 1425 / 224
Combined 1758 / 261

Split policy (seed 42; see configs/hebrew_samaritan_combined_split.json):

  • Samaritan: manuscript-level holdout (~12% pages target) — 43 train MSS / 7 val MSS (no page from a val MS appears in train).
  • Hebrew: page-level split within Hebrew_Medieval-seg.

Compiled to Arrow (simplify_eps=0.01 → B-spline (K=8)) under outputs/hebrew_samaritan_combined/arrow/.


Training recipe (Stage-1)

Hyperparameter Value
Init weights Stage-0 PRETRAIN/stage0/best_cbad_f1.safetensors
Optimizer AdamW; base LR (5\times10^{-5}) (backbone (0.1\times))
Schedule Linear warmup + cosine (Lightning defaults as in dfine-det train)
Precision bf16-mixed
Effective batch 8 (micro-batch 8 × accum 1 on this run)
Cap / early-stop max 150 epochs; quit: early on cbad_f1_max; patience 20; min_epochs 15; min_delta 0.0005
Best snapshot epoch 109 (checkpoints/best-epoch=109.ckpt → best_cbad_f1.safetensors)
Monitor peak cbad_f1_max ≈ 0.9599
Conf sweep (train) {0.1, 0.2, 0.3, 0.4, 0.5, 0.6}; fast_conf_sweep: true
Losses / matcher loss_class=8.0, loss_height=3.0, cost_y=3.0
Augment Enabled (--augment / YAML augment: true)
Match distance 20 px on the 1280 canvas
Seed 42

Evaluation

Metric: cBAD-style line F1 — Hungarian matching of densified polylines with mean bidirectional Chamfer cost; TP if cost ≤ dist_thresh (20 px).

Operating point (recommended)

Lock conf = 0.50 using the unseen Hebrew holdout (not the Stage-1 val peak). On Stage-1 val, 0.50 is within ~0.0015 F1 of the val-best conf (0.40).

1) Unseen Hebrew holdout (primary report)

Source: /opt/DATA/DATASETS/multi-export-hebrew (674 PAGE pages).
Excluded any page already in Stage-1 train/val (exact basename) or Stage-0 lists (exact / post-__ stem) → 377 pages.
List: configs/multi_export_hebrew_unseen_test.lst.

dfine-det --config configs/hebrew_samaritan_stage1.yaml -d cuda:0 test \
  --weights best_cbad_f1.safetensors \
  --val-arrow outputs/hebrew_samaritan_stage1/arrow_unseen_test/val.arrow \
  --conf 0.50
conf cbad_f1 precision recall mean_chamfer
0.10 0.7369 0.6010 0.9523 5.62
0.20 0.8750 0.8288 0.9268 6.51
0.30 0.8901 0.8851 0.8951 7.86
0.40 0.8901 0.9070 0.8738 7.14
0.50 0.8977 0.9321 0.8657 5.91
0.60 0.8935 0.9464 0.8462 5.18

Headline (unseen @ 0.50): F1 0.898 · P 0.932 · R 0.866

2) Stage-1 in-domain val (early-stop split)

261 pages (hebrew_samaritan_combined val Arrow). Optimistic vs true holdout.

conf cbad_f1 precision recall mean_chamfer
0.10 0.7937 0.6623 0.9902 3.07
0.20 0.9317 0.8928 0.9741 4.30
0.30 0.9494 0.9394 0.9597 5.10
0.40 0.9553 0.9602 0.9504 4.76
0.50 0.9538 0.9718 0.9364 4.14
0.60 0.9503 0.9814 0.9211 3.55

Train monitor cbad_f1_max ≈ 0.9599 @ ep109 (best across the same conf grid during training).

Caveats

  • Unseen set is Hebrew multi-export pages not in Stage-0/1 lists; it is not a Samaritan-only test and not manuscript-disjoint vs every possible Hebrew source.
  • Stage-1 val F1 is not interchangeable with unseen F1 (~0.95 vs ~0.90).
  • Fixed conf 0.1 (train log cbad_f1) is far below operating F1 — always sweep or use 0.50.

How to use

Install

pip install -e ".[dev]"   # from the dfine-det repo; see QUICKSTART for torch/CUDA notes

Inference (PAGE XML)

dfine-det -d cuda:0 infer \
  --weights best_cbad_f1.safetensors \
  --image page.jpg \
  -o page.xml \
  --format page \
  --conf 0.5 \
  --reading-order rtl

Evaluate

dfine-det --config configs/hebrew_samaritan_stage1.yaml -d cuda:0 test \
  --weights best_cbad_f1.safetensors \
  --val-arrow path/to/val.arrow \
  --conf 0.50

Files in this release

File Description
best_cbad_f1.safetensors Recommended Stage-1 weights (monitor peak ≈ 0.9599)
best_0.9599.safetensors Same snapshot, score-tagged filename
checkpoints/best-epoch=109.ckpt Optional Lightning resume state
README.md This model card

Upload the safetensors (+ this README) for Hub users; Lightning .ckpt is optional.


Limitations

  • Domain gap remains between Stage-1 val (0.95 F1) and unseen Hebrew (0.90 F1).
  • (K=8) control points underfit strongly curved / damaged lines.
  • Vertical “ghost” doubles can still appear on some layouts; cost_y / loss_height mitigate but do not eliminate them.
  • Private / institutional manuscript images used in training are not redistributed with the weights; respect each corpus license.
  • Do not cite these scores as cBAD 2019 / Orli bake-off results.

Citation & credits

D-FINE (architecture & COCO/Objects365 pretrain)

@inproceedings{peng2025dfine,
  title     = {D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement},
  author    = {Peng, Yansong and Li, Hebei and Wu, Peixi and Zhang, Yueyi and Sun, Xiaoyan and Wu, Feng},
  booktitle = {The Thirteenth International Conference on Learning Representations},
  year      = {2025},
  url       = {https://arxiv.org/abs/2410.13842}
}

Stage-0 polyline pretrain

@software{dfine_det_stage0_large,
  title  = {dfine-det Large Stage-0: Multiscript Polyline Baseline Pretrain},
  author = {John Locke Jrr, johnlockejrr},
  year   = {2026},
  url    = {https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0}
}

This Stage-1 Hebrew/Samaritan fine-tune

@software{dfine_det_hebsam_stage1,
  title  = {dfine-det Large Stage-1: Hebrew and Samaritan Polyline Baseline Detection},
  author = {John Locke Jrr, johnlockejrr},
  year   = {2026},
  note   = {Fine-tuned from johnlockejrr/dfine-det-large-baseline-stage0; operating conf 0.50},
  url    = {https://huggingface.co/johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1}
}

Additional notices

  • Vendored D-FINE modules: see THIRD_PARTY_NOTICES.md (Apache-2.0, © 2024 The D-FINE Authors).
  • PAGE polygonalization helpers derive from Kraken/BLLA (Apache-2.0, © Benjamin Kiessling) — used at export only.
  • Training corpora: follow each dataset’s original license / institutional terms.

License

Apache License 2.0 for the dfine-det code and these weights, consistent with D-FINE’s Apache-2.0 release. Downstream users must also comply with licenses of any datasets used in further fine-tuning.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1

Base model

Peterande/D-FINE
Finetuned
(2)
this model

Paper for johnlockejrr/dfine-det-large-baseline-hebrew-samaritan-stage1

Evaluation results