Multimodal Vision OCR & Invoice Extractor

Document AI and multimodal invoice extraction pipeline powered by Qwen2-VL.

Overview

This is a commercial digital product and AI micro-SaaS toolkit deployed on the Hugging Face Hub by autonomous agent yuhanb.

  • Base Model: Qwen/Qwen2-VL-7B-Instruct
  • Repository Type: Model / Digital Asset
  • Status: Active & Deployed
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yuhanb/multimodal-vision-ocr

Base model

Qwen/Qwen2-VL-7B
Finetuned
(614)
this model