Pathumma-llm-vision-3.0.0-preview 2B

Pathumma-llm-vision-3.0.0-preview 2B is a vision-language model developed by NECTEC, based on Qwen3.5-2B and further trained for Thai and multilingual OCR and image-text understanding.

The model was trained on 377K OCR samples with a focus on improving OCR performance, particularly for Thai and challenging real-world document and scene-text images.

Model Highlights

  • 🧠 Based on Qwen3.5-2B
  • 🇹🇭 Optimized for Thai OCR
  • 📚 Trained on 377K OCR samples
  • 🖼️ Vision-language image-to-text understanding
  • 📄 Designed for OCR and document understanding
  • 🚀 Intended for efficient and compact deployment

Benchmark

We evaluate our models on ThaiOCRBench, a benchmark designed to assess OCR and document understanding capabilities across Thai and challenging real-world visual content.

ThaiOCRBench Results

Model Document Parsing Fine-grained Text Recognition Full-page OCR Handwritten Content Extraction Text Recognition Document Classification Diagram VQA Cognition VQA Infographics
Typhoon-OCR1.5-2B 0.2355 0.1274 0.7724 0.2808 0.5922 0.4093 0.4063 0.5690 0.5839
Pathumma-LLM-Vision-3.0.0-re 0.4117 0.1478 0.7361 0.3675 0.6742 0.4326 0.4363 0.6623 0.5951
Pathumma-LLM-Vision-3.0.0-preview 0.4872 0.1458 0.7831 0.3286 0.6443 0.2558 0.5784 0.6763 0.6451

Training

The model was trained using 377K OCR samples.

Training Configuration

Parameter Value
Base model Qwen3.5-2B
Training data 377K OCR samples
Training method SFT
Learning rate 1.0e-5
Epochs 3
lr_scheduler cosine
Gradient accumulation 2
Hardware 4 × NVIDIA A100

The training setup was designed for large-scale OCR fine-tuning using 4 NVIDIA A100 GPUs.

Intended Use

Pathumma Vision 3.5-2B is intended for:

  • Thai OCR
  • Scene text recognition
  • Document text extraction
  • Thai document understanding
  • Efficient OCR deployment

Quickstart

Installation

pip install -U transformers
pip install torch torchvision

Using 🤗 Transformers

import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "nectec/Pathumma-llm-vision-3.0.0-preview"

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "path/to/your/image.jpg",
            },
            {
                "type": "text",
                "text": "อ่านข้อความในภาพนี้",
            },
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
)

inputs = inputs.to(model.device)

generated_ids = model.generate(
    **inputs,
    max_new_tokens=512,
)

generated_ids_trimmed = [
    out_ids[len(in_ids):]
    for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)

print(output_text[0])

Contributors

This model was developed by:

  • Theerawat Phromchai
  • Kun Kerdthaisong
  • Thanaporn Pintobtang
  • Khemjira Prachumkhong
  • Teepakorn Lilek
  • Theerasit Issaranon
  • Sarawoot Kongyoung

Acknowledgements

We thank the NECTEC team and contributors involved in the development of Pathumma and Thai-language vision-language resources.

This model is built upon the Qwen3.5 architecture and benefits from the work of the Qwen team.

Citation

If you find Pathumma-llm-vision-3.0.0-preview useful in your research, please cite:

@misc{PathummaVision3,
  author = {
    Phromchai, Theerawat and
    Kerdthaisong, Kun and
    Pintobtang, Thanaporn and
    Prachumkhong, Khemjira and
    Lilek, Teepakorn and
    Issaranon, Theerasit and
    Kongyoung, Sarawoot
  },
  title = {Pathumma Vision 3.5-2B},
  year = {2026},
  url = {https://huggingface.co/nectec/Pathumma-llm-vision-3.5-2b}
}
Downloads last month
62
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nectec/Pathumma-llm-vision-3.0.0-preview

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(331)
this model