LFM2.5-VL-3B-DragOn — GGUF

GGUF quants of EryriLabs/LFM2.5-VL-3B-DragOn: LiquidAI's LFM2.5-VL-3B fine-tuned for drag-and-drop grounding on GUI screenshots. Screenshot + instruction in, {"start":[x,y],"end":[x,y]} out (0-1000 normalised coordinates).

The short version of the story: the base model scores 0.7% on the DragOn public eval, this fine-tune scores 70.8% acc@5 (78.5% acc@10), and it cost about $36 to train. Details, per-domain numbers and caveats are in the main repo.

Files

You need TWO files: a main model quant plus the vision projector (mmproj).

file size note
LFM2.5-VL-3B-DragOn-Q4_K_M.gguf ~1.5 GB good default, runs on almost anything
LFM2.5-VL-3B-DragOn-Q5_K_M.gguf ~1.8 GB
LFM2.5-VL-3B-DragOn-Q6_K.gguf ~2.0 GB recommended if you have the room
LFM2.5-VL-3B-DragOn-Q8_0.gguf ~2.7 GB
LFM2.5-VL-3B-DragOn-F16.gguf ~5.1 GB reference
mmproj-LFM2.5-VL-3B-DragOn-F16.gguf ~0.8 GB vision projector, always required

Note that coordinates are a precision task, so if you see degraded accuracy at Q4, step up a quant before blaming the model.

Usage

llama-server -m LFM2.5-VL-3B-DragOn-Q6_K.gguf \
  --mmproj mmproj-LFM2.5-VL-3B-DragOn-F16.gguf \
  -ngl 99 -c 4096 --temp 0

Then send a chat completion with the image and this exact prompt shape (it's what the model was trained on):

This is a screenshot of a user interface. You must perform a DRAG action.
Task: <your instruction>
Give the drag as JSON with the START point (where the mouse button goes down) and the END point (where it is released), in coordinates normalised to 0-1000 for both x (left->right) and y (top->bottom):
{"start":[x,y],"end":[x,y]}
Output only the JSON.

Multiply by your actual screen size /1000 and you have your drag.

Thanks

To LiquidAI for the base model and to Nathan Bout, Maxime Langevin and Ronan Riochet at Hcompany for the DragOn dataset — see the main repo card for the full credits.

Quantised by Dwain Barnes (EryriLabs), August 2026.

Downloads last month
279
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EryriLabs/LFM2.5-VL-3B-DragOn-GGUF

Quantized
(1)
this model

Dataset used to train EryriLabs/LFM2.5-VL-3B-DragOn-GGUF