--- license: apache-2.0 base_model: PaddlePaddle/PaddleOCR-VL-1.6 tags: - gguf - ocr - paddleocr library_name: gguf --- # PaddleOCR-VL 1.6, quantized GGUF Quantized GGUF of [PaddlePaddle/PaddleOCR-VL-1.6](https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6), for running in the browser via [wllama](https://github.com/ngxson/wllama). The official [PaddleOCR-VL-1.6-GGUF](https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6-GGUF) release is F16 and totals 1.73 GB. This one is 856 MB with no measurable difference in output. | file | size | what it is | | --- | --- | --- | | `PaddleOCR-VL-1.6-Q4_K_M.gguf` | 286 MB | decoder | | `mmproj-Q8_0.gguf` | 570 MB | vision projector | Both files are required. ## Usage ```sh llama-mtmd-cli -m PaddleOCR-VL-1.6-Q4_K_M.gguf --mmproj mmproj-Q8_0.gguf \ --image crop.png -p "OCR:" --jinja --temp 0 ``` Use the prompt `OCR:`. The model expects a crop of a single text region. ## Quantization Decoder: `llama-quantize` Q4_K_M from the official F16 GGUF. Projector: `convert_hf_to_gguf.py --mmproj --outtype q8_0` from the safetensors release (llama.cpp b10150). The upstream vision config declares `SiglipVisionModel`; the converter's mmproj path expects `PaddleOCRVisionModel`, so that field was renamed before converting. No weights were altered. Checked against the F16 originals on Japanese, Korean and Chinese comic pages; output was character-identical. ## Licence Apache-2.0, inherited from PaddleOCR-VL.