portal-vlm: qwen25vl-lora-lm

Grounding LoRA (r16/a32), 252 LM-decoder sites, native absolute-pixel coordinate dialect. wave-ui-25k, 1 epoch, best checkpoint at step 372/1489 by held-out click probe.

Part of the portal-vlm release - an independent replication of Ramp Labs' PorTAL (portable task adapters via hypernet-generated LoRA) extended to vision-language models on GUI grounding.

ScreenSpot-v2 overall web split
this artifact 77.1% 72.8%

Reproduce this row without training

git clone https://github.com/robbym-dev/portal-vlm && cd portal-vlm && uv sync
uv run python scripts/eval.py --config configs/qwen25vl_lora_lm.yaml --adapter hf:manihani4/portal-vlm-qwen25vl-lora-lm

Standard PEFT LoRA adapter - also loadable directly with peft.PeftModel.from_pretrained on the pinned base model.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for manihani4/portal-vlm-qwen25vl-lora-lm

Adapter
(257)
this model