--- language: - fr license: mit library_name: transformers pipeline_tag: text-classification base_model: microsoft/layoutlmv3-base tags: - transformers - safetensors - layoutlmv3 - text-classification - sequence-classification - document-layout-analysis - document-understanding - bulletin-officiel - french-government - sommaire - ocr - document-ai - private datasets: - custom metrics: - accuracy - f1 widget: - text: "SOMMAIRE" example_title: "Sommaire banner" - text: "Dahir n° 1-23-45 portant promulgation" example_title: "Titre segment" model-index: - name: sommaire-layoutlmv3-classifier results: - task: type: text-classification name: Sommaire segment layout classification (9 classes) dataset: type: custom name: SommaireTOC BO segments (held-out test, n=1800) metrics: - type: f1 name: F1 macro (test) value: 0.9738 - type: accuracy name: Accuracy (test) value: 0.9867 - type: f1 name: F1 micro (test) value: 0.9867 - type: f1 name: F1 weighted (test) value: 0.9867 - task: type: text-classification name: Validation (checkpoint selection) dataset: type: custom name: SommaireTOC BO segments (val) metrics: - type: f1 name: F1 macro (val, best) value: 0.9680 - type: accuracy name: Accuracy (val, best) value: 0.9861 - type: epoch name: Best epoch value: 15 --- # SommaireTOC Segment Classifier — LayoutLMv3 (gated) Fine-tuned [`microsoft/layoutlmv3-base`](https://huggingface.co/microsoft/layoutlmv3-base) for **segment-level layout classification** on French **Bulletin Officiel (BO)** SOMMAIRE / TOC pages. Each annotated region (page crop image + OCR words + word boxes) is encoded as one sequence; the model predicts **one of 9 layout classes** (`LayoutLMv3ForSequenceClassification`). > **Gated model** — trained on proprietary/internal BO annotations. Best weights from run `run_20260809_031930` (checkpoint-21990 / epoch 15). ## Model description | Property | Value | |----------|-------| | Architecture | `LayoutLMv3ForSequenceClassification` | | Base model | `microsoft/layoutlmv3-base` | | Task | Segment-level sequence classification | | Classes | **9** SommaireTOC layout roles | | Input | Page/region image + word tokens + word boxes (0–1000) | | Parameters | ~126M | | Best checkpoint | `checkpoint-21990` (epoch **15**) | | Selection metric | `eval_f1_macro` | ### Layout classes (9) | id | label | Role | |---:|-------|------| | 0 | `Titre` | TOC title line | | 1 | `Page_title` | Page number associated with a title | | 2 | `Som_Section` | Section banner inside sommaire | | 3 | `Sommaire` | SOMMAIRE header / banner | | 4 | `Page` | Standalone page number | | 5 | `Meta_Data` | Header/footer metadata | | 6 | `Tex_PAR` | Text block — PARTICULIER | | 7 | `Tex_GEN` | Text block — GENERAL | | 8 | `Tex_AVIS` | Text block — AVIS | ## Training data | Item | Value | |------|------:| | Segments (pipeline scale) | 15,398 | | Corrected GT titles (ref) | 3,697 | | Task | `sommaire_segment_classification` | | Run id | `run_20260809_031930` | | Domain | French Bulletin Officiel SOMMAIRE pages | See [`metrics/run_config.json`](metrics/run_config.json). ## Training hyperparameters | Parameter | Value | |-----------|------:| | Epochs requested | 30 | | Epochs to best | **15** (early stopping patience 5; run stopped ~20) | | Batch size (per device) | 4 | | Gradient accumulation | 2 | | Effective batch size | 8 | | Learning rate | 2e-5 | | Warmup ratio | 0.1 | | Weight decay | 0.01 | | FP16 | true | | Optimizer metric | `f1_macro` | | Best val F1 macro | **0.9680** | | Best global step | 21990 | Summaries: [`metrics/trainer_summary.json`](metrics/trainer_summary.json), [`metrics/val_best_metrics.json`](metrics/val_best_metrics.json) ## Evaluation results ### Held-out test set (n=1800) From [`metrics/test_metrics.json`](metrics/test_metrics.json): | Metric | Value | |--------|------:| | **Accuracy** | **0.9867** | | **F1 macro** | **0.9738** | | **F1 micro** | **0.9867** | | **F1 weighted** | **0.9867** | | Precision macro | 0.9700 | | Recall macro | 0.9790 | | Mean confidence | 0.9971 | ### Per-class F1 (test) | Class | Precision | Recall | F1 | Support | |-------|----------:|-------:|---:|--------:| | Meta_Data | 1.00 | 0.99 | 0.994 | 90 | | Page | 0.94 | 0.95 | 0.945 | 100 | | Page_title | 0.99 | 0.99 | 0.989 | 593 | | Som_Section | 0.98 | 0.97 | 0.978 | 231 | | Sommaire | 1.00 | 0.99 | 0.995 | 96 | | Tex_AVIS | 0.88 | 1.00 | 0.933 | 7 | | Tex_GEN | 1.00 | 0.92 | 0.960 | 26 | | Tex_PAR | 0.95 | 1.00 | 0.974 | 38 | | Titre | 1.00 | 0.99 | 0.994 | 619 | ### Validation (best checkpoint) | Metric | Value | |--------|------:| | F1 macro | 0.9680 | | Accuracy | 0.9861 | | Epoch | 15 | ## Plots ![Confusion matrix (new)](plots/confusion_matrix_new.png) ![Per-class F1 compare](plots/per_class_f1_compare.png) ![Overall metrics compare](plots/overall_metrics_compare.png) ![Val F1 macro](plots/val_f1_macro.png) ![Train loss](plots/train_loss.png) ![Val loss](plots/val_loss.png) ## Example pages ![BO_6954 page 1](examples/BO_6954_Fr__page_001.png) ![BO_6954 page 2](examples/BO_6954_Fr__page_002.png) ![BO_6962 page 1](examples/BO_6962_fr__page_001.png) ## Usage ```python from transformers import AutoProcessor, LayoutLMv3ForSequenceClassification import torch from PIL import Image repo = "AvoCahDoe/sommaire-layoutlmv3-classifier" processor = AutoProcessor.from_pretrained(repo, apply_ocr=False) model = LayoutLMv3ForSequenceClassification.from_pretrained(repo) model.eval() image = Image.open("region_crop.png").convert("RGB") # words / boxes: list[str], list[list[int]] normalized 0–1000 encoding = processor( image, words=words, boxes=boxes, return_tensors="pt", truncation=True, padding="max_length", max_length=512, ) with torch.no_grad(): logits = model(**encoding).logits pred_id = int(logits.argmax(-1).item()) print(model.config.id2label[pred_id]) ``` ## Repository layout ``` model.safetensors, config.json tokenizer / preprocessor files metrics/ # test, val, run_config plots/ # curves + confusion / comparisons examples/ # BO page images ``` ## Intended use - Classification stage of SommaireTOC after region proposal - Labels feed title / page extraction and `pdf_page → BO_page` mapping ## Limitations - Tuned for French BO SOMMAIRE segments; other layouts may degrade - Rare classes (`Tex_AVIS`) have low support in test - Private weights; do not redistribute without authorization - Requires word-level OCR + boxes (processor `apply_ocr=False` in the snippet above) ## Related models - Region proposal: [`AvoCahDoe/sommaire-pp-doclayout-l-proposal`](https://huggingface.co/AvoCahDoe/sommaire-pp-doclayout-l-proposal) - Earlier public segment classifier: [`AvoCahDoe/layoutlmv3-bo-segments`](https://huggingface.co/AvoCahDoe/layoutlmv3-bo-segments) ## Citation If you use this model in work derived from the SommaireTOC pipeline, please cite the Hub repo and base model `microsoft/layoutlmv3-base`.