--- license: apache-2.0 base_model: Qwen/Qwen3.5-4B library_name: transformers pipeline_tag: image-text-to-text tags: - vision-language - tool-calling - agent - qwen3_5 - local-agents - edge --- ![Karti-Small-VL-4B — it knows when it can't read something](./assets/karti-small-vl-4b.svg) # Karti-Small-VL-4B · v1 **A 4B vision model that knows when it can't read something.** Built for local agents that look at a thing and then call a tool about it — read a panel, check a camera frame, act on a screen. The failure that matters there isn't a wrong caption; it's a **confidently invented identifier**. v1 is trained to hit that boundary and decline instead. | | | |---|---| | Parameters | 4.66 B (4,659,865,088) | | Base | [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) @ `851bf6e8` | | Precision | BF16 · 9.3 GB · merged weights, no adapter | | Context | 32,768 | | Tools | native, `qwen3_xml` parser | | Faster build | [NVFP4](https://huggingface.co/KartiOS/Karti-Small-VL-4B-NVFP4) — 2.5×, Blackwell only | ## Results ![v1 against the untrained base](./assets/karti-vl-4b-results.svg) Same endpoint, same rows, same decode path. | | base | **v1** | |---|---|---| | **Invented-identifier rate** | 0.378 | **0.023** — 16× fewer | | Panel reading | 0.595 | **0.967** | | GUI grounding, ScreenSpot-v2 · n=1271 | 0.718 | **0.907** | | ⤷ small targets · n=426 | 0.580 | **0.883** | | Unparseable click points | 48/1272 | **9/1272** | | General vision, curated · n=220 | 0.727 | **0.732** | | False-refusal, real photos · n=600 | 0.220 | 0.227 | ScreenSpot-v2 and the general-vision slice were **held out**. On the two hardest abstention families the base scored zero; v1 scores 101/112 and 117/119. **Scope, honestly:** part of the grounding gain is learning the normalised 0–1000 coordinate convention shared with the training set — real, but not "38% better at seeing UI". And v1 reads *photographs* no better than the base; what changed is that it stops over-refusing them. The legibility gain is on rendered panels and screens. ## Speed — NVIDIA DGX Spark (GB10) vLLM 0.27.1, BF16, 32k context, FP8 KV cache, temperature 0, single stream. | tokens | tok/s | |---|---| | 128 | 21.0 | | 512 | 21.1 | Fits in ~20 GB. For 2.5× throughput at 3.7 GB, use the [NVFP4 build](https://huggingface.co/KartiOS/Karti-Small-VL-4B-NVFP4) (needs Blackwell, sm_120+). ## Use ```python from transformers import AutoProcessor, AutoModelForImageTextToText m = AutoModelForImageTextToText.from_pretrained( "KartiOS/Karti-Small-VL-4B", dtype="auto", device_map="auto") p = AutoProcessor.from_pretrained("KartiOS/Karti-Small-VL-4B") ``` ```bash vllm serve KartiOS/Karti-Small-VL-4B \ --max-model-len 32768 --kv-cache-dtype fp8 \ --enable-auto-tool-choice --tool-call-parser qwen3_xml \ --default-chat-template-kwargs '{"enable_thinking": false}' ``` Pin a revision — `main` moves. `revision="pre-v1"` is the untrained baseline. ## How it was trained BF16 LoRA (r16, α32, lr 5e-5, 75 steps) merged into these weights. 32.5 M trainable parameters across 248 language modules; the **vision tower is frozen and proven frozen** — all 297 vision tensors are byte-identical to the base. Corpus: 3,569 rows — an authored legibility ladder that places text exactly at the readable/unreadable boundary, plus VizWiz and GUI-Odyssey. The decisive parameter was **the share of the corpus whose answer is "I can't read this."** At 29.1% the model learns the word, not the judgement: 0.955 on the lane, then refusing 14 of 22 answerable photographs. At 14.4% it learns the judgement. That threshold — not the learning rate, not the step count — is what made v1 work. > **Serve the merged weights, not a LoRA adapter.** vLLM 0.27.1 applies this > adapter incompletely — about 0.19 absolute ScreenSpot accuracy, silently. > Qwen3.5 packs its linear-attention projections and this adapter targets those > sub-modules. The adapter is not published for that reason. --- From [Lumbridge](https://lumbridgecorp.com/models). Sibling text model: [`Karti-Small-RSI-3B`](https://huggingface.co/KartiOS/Karti-Small-RSI-3B).