--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE base_model: Qwen/Qwen2.5-VL-3B-Instruct tags: - invoice-extraction - document-understanding - vision-language-model - qwen2.5-vl - full-fine-tuning language: - en pipeline_tag: image-text-to-text --- # Open to contribute? I'm actively expanding to improve model performance. If you: - Have invoice datasets you can legally share - Want to collaborate on building open IDP datasets - Want to contribute annotations or synthetic data generation - Have ideas for improving document extraction benchmarks **Please reach out via:** - Hugging Face Discussions # Qwen2.5-VL-3B — IDP / Invoice Extraction (Full Fine-Tune) Full-parameter fine-tune of **Qwen2.5-VL-3B-Instruct** for **Intelligent Document Processing (IDP)** — structured invoice field extraction. Given a scanned invoice image, it returns a JSON object with fields like vendor, customer, invoice_date, total_amount, currency, and more. **Proof-of-concept / learning project.** Trained on 1x H100. ## Training - Base: Qwen/Qwen2.5-VL-3B-Instruct - Method: full-parameter SFT (ms-swift), 5 epochs, best-checkpoint by val loss - Data: public web data — small for poc - Effective batch 16, LR 1e-5, bf16, flash-attention-2 ## Results (481-sample val set, vs untrained base) | Field | Base model | This model | |---|---|---| | Valid JSON output | 0% (fails) | **100%** | | currency | 2.5% | **85%** | | invoice_date | 2.9% | **74%** | | total_amount | 17.5% | **68%** | | invoice_number | 35.7% | **57%** | | vendor.name | 0% | **48%** | Fine-tuning produced large gains on every key field. High accuracy on easy fields (payment_method, email, discount, balance_due: 90-100%). Numeric fields (amounts, invoice numbers) top out ~60-70% — a capacity limit of the 3B base on hard scanned documents. Suitable as a first-pass extractor with human review. ## Usage (ms-swift) ```python from swift.infer_engine import TransformersEngine, InferRequest, RequestConfig engine = TransformersEngine("rachitpandey26/qwen2.5-vl-3b-idp") req = InferRequest(messages=[ {"role": "system", "content": "You are an invoice extraction model. Given a scanned invoice image, extract the structured fields and return ONLY a single JSON object."}, {"role": "user", "content": "Extract the invoice fields as JSON."}, ], images=["invoice.png"]) resp = engine.infer([req], request_config=RequestConfig(max_tokens=512, temperature=0)) print(resp[0].choices[0].message.content) ``` ## Limitations - Numeric fields (invoice numbers, amounts, dates) reach ~60-70% — not reliable enough for fully automated reconciliation without human review. - Trained on RVL-CDIP-style scanned invoices; may not generalize to very different invoice layouts. - Inherits the Qwen-Research license restrictions from the base model. ## Acknowledgements Base model by Alibaba Qwen team. Dataset publicaly avaliable on web. Fine-tuned with ms-swift on Modal.