Instructions to use ML-Intern-lab/citrus-disease-vlm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ML-Intern-lab/citrus-disease-vlm with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Citrus Disease VLM (Qwen3.5-2B, LoRA)
A LoRA fine-tune of Qwen/Qwen3.5-2B that looks at a photo of a citrus leaf, fruit or shoot, names the disease, pest or nutrient deficiency, explains the cause and symptoms, and recommends both biological/organic and chemical management. Trained with TRL SFTTrainer on the private instruction dataset ML-Intern-lab/citrus-disease-vlm-instruct (3,017 train / 335 test examples, 21 classes).
This repo contains:
- merged model — repo root, bf16 safetensors with the LoRA weights folded in (ready to load with
transformers) adapter/— the standalone LoRA adapter (r=16, α=32) to load on top of the base modeleval/— all eval artifacts: predictions, strict + relaxed metrics, confusion matricesscripts/— the exact training, evaluation and re-scoring scripts used here
Results
Evaluated on the 335-example test split: for each example the model generates an answer to the first user turn, and we check whether the class name from the dataset's knowledge base (KB[label]["name"]) appears in the answer.
| Metric (test, n=335) | Zero-shot Qwen3.5-2B | Fine-tuned (this repo) |
|---|---|---|
| Strict accuracy — KB class name in answer (primary metric) | 14.9% | 52.8% |
| Strict match rate — any KB class name found | 46.6% | 82.1% |
| Relaxed accuracy — + vetted aliases (see note) | 20.3% | 61.2% |
| Relaxed match rate | 57.0% | 91.0% |
| Relaxed accuracy, treat-only questions excluded (n=293) | 12.0% | 65.9% |
| Relaxed match rate, treat-only questions excluded | 52.9% | 97.3% |
Metric notes (read before comparing):
- Strict is the metric as specified: the KB class name (e.g. "Citrus canker") must appear in the generated answer. It structurally under-reports: 42/335 test questions are treat-only conversations where the user names the problem and the model legitimately answers with treatment only, never restating the diagnosis (34 of the 60 "unrecognized" fine-tuned answers are of this type); and 18/38 healthy-image answers use the template "I do not see any disease, pest or deficiency symptoms here", which does not contain the KB name "Healthy citrus tissue".
- Relaxed re-scores the same generated answers (
eval/predictions_*.jsonl) adding a per-class alias table (pathogen/scientific names and distinctive KB phrases, vetted to zero false positives across all 670 answers — e.g. "lasiodiplodia theobromae" → die_back, "aleurocanthus spiniferus" → spiny_whitefly, "no abnormalities" → healthy). The alias table was tuned on this eval set, so treat relaxed as a secondary, optimistic number; the strict column is the primary as-specified metric. Classes with no safe alias (melanose, scab, red_scale, mn_deficiency) gain nothing. - Full JSON with both matchers, per-label breakdowns and confusion matrices:
eval/results_zeroshot.json,eval/results_finetuned.json(strict) andeval/results_*_relaxed.json(relaxed). Re-scoring script:scripts/rescore.py.
Per-category accuracy (strict, primary metric)
| Category | n | Zero-shot | Fine-tuned |
|---|---|---|---|
| disease_bacterial | 94 | 22.3% | 69.1% |
| disease_fungal | 92 | 17.4% | 67.4% |
| pest | 73 | 8.2% | 56.2% |
| nutrient_deficiency | 38 | 18.4% | 23.7% |
| healthy | 38 | 0.0% | 0.0%* |
* Under the strict matcher the healthy class scores 0 because the KB name is literally "Healthy citrus tissue", which the fine-tuned model (which answers with a "no symptoms seen" template) never echoes verbatim. Under the relaxed matcher: healthy 42.1% fine-tuned vs 18.4% zero-shot.
Per-label accuracy (relaxed matcher, fine-tuned model)
Diagnosis is strong for most fungal/bacterial diseases and pests (relaxed): die_back 90%, spiny_whitefly 85%, powdery_mildew 80%, mealybugs 80%, shot_hole 80%, greening 76%, greasy_spot 70%, canker 61%, black_spot 53%, red_scale_sequelae 60%, texas_mite 60%, healthy 42%. Weak / not measurable: nutrient deficiencies (fe 40%, mg 20%, zn 30%, n 0%) — deficiencies are visually confusable with greening blotchy mottle and with each other — and the tiny classes melanose, scab, red_scale, mn_deficiency (1–3 test images each, no safe match alias).
Confusion matrices (relaxed matcher, 21 labels + "unrecognized")
Largest fine-tuning gains (error mass removed, relaxed): greening→fe_deficiency 16→0, canker→unrecognized 16→1, healthy→unrecognized 23→4, spiny_whitefly→unrecognized 15→2, die_back→unrecognized 11→1.
Training configuration
| Base model | Qwen/Qwen3.5-2B (Qwen3_5ForConditionalGeneration) |
| Method | TRL SFTTrainer (native VLM path, images + messages columns), LoRA via PEFT |
| LoRA | r=16, α=32, dropout=0.05, all 186 language-model linear layers (self-attention q/k/v/o, linear-attention in_proj_a/b/z/qkv + out_proj, MLP gate/up/down) via a verified regex; vision tower untouched |
| Precision | bf16, gradient checkpointing |
| Optimizer | AdamW, lr 1e-4, cosine schedule, 3% warmup |
| Batch | per-device 2 × grad-accum 8 = effective 16 |
| Epochs / steps | 2 epochs = 378 optimizer steps, seed 42 |
| Sequence length | uncapped (max_length=None — image tokens never truncated); images ≤768 px longest side (processor pixel budget capped at 589,824) |
| Runtime | 1 h 24 m on 1× A10G (24 GB), ~13.3 s/step (reference linear-attention kernels; causal_conv1d/flash-linear-attention not installed) |
| Metrics | Trackio dashboard — final train loss ≈ 1.10, token accuracy ≈ 0.996 |
Usage
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"ML-Intern-lab/citrus-disease-vlm", dtype=torch.bfloat16, device_map="cuda")
processor = AutoProcessor.from_pretrained("ML-Intern-lab/citrus-disease-vlm")
messages = [{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": "What is wrong with this citrus plant and how do I treat it?"},
]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = processor(text=prompt, images=[pil_image], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(processor.tokenizer.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
To use the adapter alone (base + LoRA):
from peft import PeftModel
base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.5-2B", dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "ML-Intern-lab/citrus-disease-vlm", subfolder="adapter")
Dataset and attribution
Training data: ML-Intern-lab/citrus-disease-vlm-instruct — 21 unified classes (fungal and bacterial diseases, insect and mite pests, nutrient deficiencies, healthy), one image per example, assistant answers generated from a curated knowledge base. Images come from three CC BY 4.0 datasets published in Data in Brief:
- Rauf, H. T. et al. (2019). A citrus fruits and leaves dataset for detection and classification of citrus diseases through machine learning. Data in Brief 26, 104340. Mendeley Data doi:10.17632/3f83gxmv57.2. Hub mirror:
Project-AgML/citrus_fruit_leaf_disease_classification. - Gomez-Flores, W., Garza-Saldana, J. J., Varela-Fuentes, S. E. (2024). CitrusUAT: A dataset of orange Citrus sinensis leaves for abnormality detection using image analysis techniques. Data in Brief 52, 109908. Zenodo doi:10.5281/zenodo.8294078. Hub mirror:
Project-AgML/citrusuat_disease_classification. - Emon, Y. R., Ahad, M. T., Rabbany, G. (2024). Multi-format open-source sweet orange leaf dataset for disease detection, classification, and analysis. Data in Brief 55, 110713. Mendeley Data doi:10.17632/f7cr74mwpj.2. Hub mirror:
Project-AgML/orange_leaf_disease_classification.
This model inherits CC BY 4.0 from the source datasets.
Limitations
- Treatment text is templated from a knowledge base, not written per image; review with an agronomist before real decisions. Chemical recommendations name active ingredients only; registration, dose and pre-harvest intervals differ by country.
- Nutrient deficiencies remain weak (confusable with each other and with greening blotchy mottle); sources were photographed in Pakistan, Mexico and Bangladesh, so expect domain shift on other varieties, backgrounds and cameras.
- The relaxed metric's alias table was vetted on this test set (post-hoc); the strict metric is the primary as-specified number.
- Downloads last month
- -


Task type is invalid.