Instructions to use kaan84/sinhala-print-trocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kaan84/sinhala-print-trocr with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="kaan84/sinhala-print-trocr")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("kaan84/sinhala-print-trocr", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Pre-registered Research Protocol
1. Administrative record
- Project: TrOCR Sinhala Printed Lines
- Protocol version: 1.0
- Protocol date: 2026-07-28
- Researcher: Kalana Hewapathirana
- Planned repository:
kaan84/trocr-sinhala-printed-lines - Phase: Phase 0, before training
Commit this file before training. Any later change must be appended to the deviation log with its date, rationale, and expected impact.
2. Research question
Does synthetic-data fine-tuning of Ransaka/TrOCR-Sinhala improve printed, line-level Sinhala OCR on real documents?
3. Hypotheses
Primary hypothesis H1
Fine-tuning the starting checkpoint on Ransaka/sinhala_synthetic_ocr-large will reduce CER on a real printed Sinhala line benchmark relative to the unchanged checkpoint and Tesseract 5.5.
Secondary hypothesis H2
The fine-tuned checkpoint will reduce WER relative to the unchanged checkpoint and Tesseract 5.5.
Secondary hypothesis H3
Performance will be lower on older or visibly degraded documents than on newer or cleaner documents.
4. Experimental unit and scope
The evaluation unit is one cropped text-line image paired with one reference transcription. Page-level and line-level metrics must not be combined in the same comparison table without an explicit unit label.
Included:
- Printed Sinhala
- Real and synthetic text lines
- CER and WER
- Starting checkpoint and Tesseract comparisons
- Optional Surya comparison
Excluded:
- Handwriting
- End-to-end page layout analysis
- Production-readiness claims
- Hyperparameter optimization outside the declared contingency rules
5. Assets and fixed revisions
Before training, replace TO_RECORD with immutable Hub commit hashes:
| Asset | Role | Revision |
|---|---|---|
Ransaka/TrOCR-Sinhala |
Starting model | TO_RECORD |
Ransaka/sinhala_synthetic_ocr-large |
Train/validation | TO_RECORD |
avishadilhara/sinhala-ocr-lk-acts-1010 |
Real evaluation source | TO_RECORD |
The experiment manifest must also record the Transformers, Datasets, PyTorch, CUDA, jiwer, Pillow, NumPy, and Tesseract versions.
6. Data splits
Synthetic dataset
Use the published dataset's single train split.
- Split seed: 42
- Train share: 90%
- Validation share: 10%
- Split method: deterministic shuffled index split
- No sample may appear in both partitions
- Save the exact indices to version-controlled JSON files
- Record SHA-256 checksums of the index files
Real-document benchmark
Use only the published test partition for the final real-document analysis. Do not use its labels for training or hyperparameter selection.
Because the source is page-level, create line crops and transcriptions using one versioned procedure. Record:
- Source page identifier
- Line identifier
- Crop coordinates
- Reference transcription
- Segmentation method and version
- Any exclusion reason
- Document year or period, if available
Freeze the line benchmark before final evaluation. Manual corrections made after viewing model predictions invalidate the final test and must trigger a new held-out test set.
7. Preprocessing
Images
- Convert to RGB
- Use the processor associated with the starting checkpoint
- Do not apply test-time enhancement unless declared as a separate experimental condition
- Preserve raw evaluation crops
Text normalization
Report both raw and normalized metrics.
The fixed normalized condition:
- Unicode normalize with NFC.
- Convert line breaks and tabs to spaces.
- Collapse consecutive Unicode whitespace to one ASCII space.
- Strip leading and trailing whitespace.
- Preserve Sinhala characters, punctuation, numerals, and letter case.
Do not remove punctuation or diacritics in the primary analysis. Any alternative normalization is exploratory and must be labeled accordingly.
8. Training configuration
| Parameter | Fixed value |
|---|---|
| Model | Ransaka/TrOCR-Sinhala |
| Optimizer | AdamW |
| Learning rate | 4e-5 |
| Epochs | 10 |
| Effective batch size | 8 |
| Precision | FP16 if supported |
| Seeds | 42, 123, 2026 |
| Generation beams | 4 |
| Save strategy | Every epoch |
| Selection rule | Lowest validation CER |
Gradient accumulation may change only to preserve effective batch size 8. Physical batch size may be lowered for memory limits.
9. Contingency rules
Permitted without redefining hypotheses:
- Reduce physical batch size and increase gradient accumulation.
- Resume from the latest complete checkpoint after infrastructure interruption.
- Use gradient checkpointing if the GPU cannot hold the declared effective batch.
- Disable FP16 only if the hardware or software fails to support it reliably.
Not permitted in the confirmatory experiment:
- Selecting a new learning rate after inspecting final-test results
- Increasing epochs because final-test performance is weak
- Training on real-test labels
- Removing difficult samples based on prediction errors
- Reporting only the best seed
Any other adjustment becomes exploratory and must be separated from confirmatory results.
10. Baselines
Required:
- Unchanged
Ransaka/TrOCR-Sinhala - Tesseract 5.5
- Fine-tuned TrOCR for all three seeds
Optional:
- Surya OCR, only after recording the exact version and applicable model terms
Every system must receive the same frozen line images. Use system-appropriate decoding without model-specific alteration of references.
11. Metrics
Primary: Character Error Rate.
Secondary: Word Error Rate.
Use jiwer with transformations that exactly implement Section 7. Compute:
- Per-sample edit counts
- Corpus-level raw CER and WER
- Corpus-level normalized CER and WER
- Mean and standard deviation across training seeds
- Paired bootstrap 95% confidence interval for differences, resampling line IDs with 10,000 replicates and bootstrap seed 42
12. Decision rules
H1 is supported only when:
- The fine-tuned model's mean normalized CER is lower than both the starting checkpoint and Tesseract 5.5.
- Improvement over each comparator has a paired-bootstrap 95% confidence interval that excludes zero in the favorable direction.
- All three fine-tuning seeds improve over the starting checkpoint.
H2 uses the same rule with normalized WER.
H3 is descriptive unless each period contains at least 30 evaluation lines. Report counts and confidence intervals and avoid strong generalization when slices are small.
Publish results regardless of outcome. A result above published page-level baselines is not automatically a failure because the evaluation units differ.
13. Error analysis
Before viewing predictions, define these error groups:
- Character substitution
- Character deletion
- Character insertion
- Whitespace or word-boundary error
- Punctuation or numeral error
- Degraded print
- Skew or rotation
- Segmentation failure
- Mixed-language or unsupported symbol
Sample up to 50 errors using a deterministic random seed. Do not select only visually compelling cases.
14. Reporting requirements
The final report must include:
- Every seed
- Mean and standard deviation
- Confidence intervals
- Dataset and code revisions
- Hardware and runtime
- Raw and normalized metrics
- Exact evaluation-unit counts
- Exclusion counts and reasons
- Period-stratified results when valid
- Known limitations
- Deviations from this protocol
15. Licensing gate
The training dataset license is confirmed MIT. The intermediate checkpoint currently has no declared license. Training experiments may be prepared, but public redistribution of derivative weights is blocked until the checkpoint licensing position is clarified or independently reviewed.
16. Deviation log
| Date | Change | Rationale | Expected impact | Confirmatory or exploratory |
|---|---|---|---|---|
| — | None | — | — | — |