Qwen-Image-2.1 Doodle-in LoRA

A LoRA adapter for Qwen/Qwen-Image-2.1 that takes one photo with a magenta scribble drawn on it plus a short text naming an object, and returns the same photo with the scribble replaced by that object β€” following the scribble's shape and pose, lit consistently with the scene, with everything outside the scribble's neighbourhood unchanged.


Marker convention (exact β€” a Space can reproduce the training inputs byte-for-byte)

Implemented standalone in scribble.py in the dataset repo.

Property Value
Colour pure magenta RGB (255, 0, 255), opaque
Stroke width uniform(0.004, 0.010) Γ— short_side, clamped to [3, 7] px (i.e. 0.4–1.0% of the short side)
Caps / joints round: ImageDraw.line(..., joint="curve") plus filled end circles of radius width/2
Drawing directly on the photo, single colour, no blending

Prompt template: <doodle> Turn the magenta scribble into {caption}. β€” e.g. <doodle> Turn the magenta scribble into a brown leather armchair.

Magenta was chosen deliberately to differ from the red (239, 68, 68) used by object-remover LoRAs, so a remover LoRA and this insert LoRA can share one Space.

Three scribble styles were mixed 50/30/20 in training:

  • outline (50%) β€” mask contour simplified with Douglas–Peucker (eps ∈ [0.2%, 0.6%] of the contour perimeter), wobbled with smooth sinusoidal noise (amplitude 0.5–2% of the short side), with occasional small gaps (0–2 segments removed) and overshoots (up to 1.5% of the short side, 40% per endpoint)
  • detailed (30%) β€” outline plus up to 4 interior strokes taken from a Canny map of the object inside the eroded mask, sparsified to ≀40 points and wobbled
  • blob (20%) β€” a jittered convex hull resampled to 27 angular samples (radius jitter Γ—0.92–1.06, Gaussian noise 0.8% of the short side)

Training pairs additionally had the object mask dilated by 3% of the bounding-box diagonal (extended downwards for contact shadows) and filled with LaMa before the scribble was drawn.

Files in this repo

File What it is
doodle_in_lora_qwen21_gate_up_split.safetensors Recommended. Diffusers-keyed adapter. ai-toolkit saves a fused img_mlp.gate_up LoRA that diffusers silently drops (it stores gate/up as two linears); this file has the fused tensor split, verified delta-identical to the original (max deviation exactly 0.0 on all 448 modules).
doodle_in_lora_qwen21.safetensors Original ai-toolkit (ComfyUI-style) keys. Load this in ComfyUI, or use convert_lora.py to make your own split file.
doodle_in_lora_qwen21_000000500.safetensors (+ 1000, 1500) Intermediate checkpoints. Checkpoint 500 is the selected one (see evaluation).
convert_lora.py The gate/up split converter + numeric verification.
config.yaml The exact ai-toolkit training config.
optimizer.pt Optimizer state (for resuming).
checkpoints/ Checkpoint copies with sibling configs.

Usage (diffusers)

Verified with QwenImage21Pipeline.load_lora_weights at diffusers commit 0121a91f9d419ff7234c8a5923f82c244e6f1914, transformers==5.17.0, torch==2.14.0 β€” the exact code below (modulo repo ids) is what our evaluation jobs executed.

40-step edit with the LoRA

import torch
from PIL import Image
from diffusers import QwenImage21Pipeline

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", dtype=torch.bfloat16).to("cuda")

pipe.load_lora_weights(
    "ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA",
    weight_name="doodle_in_lora_qwen21_gate_up_split.safetensors",
    adapter_name="doodle")
# sanity check used in our eval: a correct load has 448 lora modules;
# a load that dropped the fused gate_up keys has fewer.
n = sum(1 for name, _ in pipe.transformer.named_modules() if "lora_A" in name)
assert n == 448, f"adapter loaded with dropped keys: got {n}, expect 448"

img = pipe(
    prompt="<doodle> Turn the magenta scribble into a brown leather armchair.",
    image=Image.open("input.jpg").convert("RGB"),
    num_inference_steps=40,
    true_cfg_scale=1.0,          # no CFG, as on the model card
    width=1024, height=768,      # multiples of 32; match your input aspect
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]

6-step turbo (both adapters stacked)

The Viggle turbo LoRA cuts the same edit to 6 steps. Stack both adapters and use the turbo sigmas and scheduler:

from diffusers import FlowMatchEulerDiscreteScheduler
from huggingface_hub import hf_hub_download

base_cfg = dict(pipe.scheduler.config)          # keep BEFORE switching
pipe.load_lora_weights(
    hf_hub_download("Viggle/Qwen-Image-2.1-viggle-turbo",
                    "Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors"),
    adapter_name="turbo")
pipe.set_adapters(["doodle", "turbo"], [1.0, 1.0])
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(
    base_cfg, shift_terminal=None)

img = pipe(
    prompt="<doodle> Turn the magenta scribble into a brown leather armchair.",
    image=Image.open("input.jpg").convert("RGB"),
    num_inference_steps=6,
    true_cfg_scale=1.0,
    sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25],
    width=1024, height=768,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]

To go back to 40 steps, unload or reset adapters and restore the base scheduler: pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(base_cfg) (the turbo branch sets shift_terminal=None, which must not leak into non-turbo runs).

Optional speed fix (same outputs, ~30 s β†’ ~3 ms per ref image on Qwen3-VL's patch-embed conv):

pe = pipe.text_encoder.model.visual.patch_embed
w = pe.proj.weight.reshape(pe.proj.weight.shape[0], -1)
pe.forward = lambda hs: hs.to(w.dtype).flatten(1) @ w.T + pe.proj.bias

Training

Trained with ostris/ai-toolkit sd_trainer, arch: qwen_image_2, from the Hub id (loading from a local diffusers dir can crash with "Cannot copy out of meta tensor" β€” ai-toolkit issue #1067):

  • LoRA linear: 32, linear_alpha: 32; optimizer: adamw8bit, lr: 1e-4, batch_size: 1
  • timestep_type: weighted, noise_scheduler: flowmatch, gradient_checkpointing: true, dtype: bf16, quantize: false, quantize_te: false
  • dataset: folder_path (targets) + control_path (magenta-marked inputs) + caption_ext: txt, resolution: [768], cache_latents_to_disk: true, cache_text_embeddings: true, sampling disabled (issue #1059)
  • 2,000 steps at 768 px; measured 2.15 s/step on one A100 80 GB; training loss 1.197 β†’ a stable ~0.07–0.3 band (final-step loss 0.077); text-embedding caching (Qwen3-VL over the control images) covered 4,773 items in 1,055 s
  • final adapter check: 192/192 lora_B tensors non-zero, max|B| = 0.052
  • full run wall-clock 1 h 38 m on A100-large

Evaluation

Setup. 160 test pairs from Open Images validation (73 outline / 50 detailed / 37 blob; 40 pairs from 23 held-out classes never seen in training). Fixed seeds (sha1(pair_id, 42) per pair). Metrics per edit:

  • Object present β€” OWLv2 (google/owlv2-base-patch16-ensemble) queried with the class name; a hit = score above threshold and box IoU β‰₯ 0.3 with the scribble's bounding box.
  • Shape fidelity β€” SAM 2.1 (facebook/sam2.1-hiera-small) prompted with the scribble bbox on the output; mask IoU against the ground-truth object mask (and against the dilated scribble region).
  • Ink removed β€” fraction of the scribble's stroke pixels still near-magenta in the output (lower is better; near zero everywhere).
  • Background preserved β€” PSNR and LPIPS against the input, outside the union of the dilated scribble region and the LaMa hole.
  • Realism β€” blind, randomised-order pairwise VLM judgement (ours vs base on the same input), 100 pairs, judge Qwen/Qwen2.5-VL-72B-Instruct via Inference Providers.

All metrics computed 160/160 with zero scoring errors; raw per-item files are in eval/full/ckpt500/ of the dataset repo, alongside the generated images.

Main comparison β€” 160 test pairs

Config Steps Prompt OWLv2 hit % bg PSNR ↑ LPIPS ↓ SAM IoU (vs GT) ↑ ink leftover ↓ s/edit ↓
ours (ckpt500) 40 <doodle> 64.4 17.38 0.3644 0.5701 0.0025 20.0
base (a) 40 plain instruction 66.9 18.55 0.3578 0.5790 0.0008 18.3
base (b) 40 <doodle> 66.2 18.34 0.3687 0.5703 0.0082 18.2
ours + Viggle turbo 6 <doodle> 67.5 17.63 0.3509 0.5809 0.0026 4.67
base + Viggle turbo (c) 6 plain instruction 65.6 17.98 0.3628 0.5683 0.0031 4.36
base + Viggle turbo 6 <doodle> 65.6 17.34 0.3823 0.5507 0.0091 4.35

Honest headline: our LoRA does not beat the base model on hit rate at 40 steps (64.4% vs 66.9%) β€” the base model can already follow painted annotations reasonably well, as its card suggests. Where the LoRA does win: at 6 turbo steps it scores the highest hit rate of all six configs (67.5%) and the best LPIPS, and the base-with-<doodle>-prompt combo shows measurably more leftover ink (0.8–0.9%) than ours at the same step counts. The practical takeaway is that the LoRA + turbo stack gives you the best quality at 4.67 s/edit; if you only ever run 40 steps, the plain base model is a competitive baseline.

Single-edit timings on the same hardware (A100 80 GB, from the smoke test): 13.7 s/edit at 40 steps (peak 37.3 GiB), 3.2 s/edit at 6 turbo steps (peak 38.6 GiB). Batched-eval means above are per-edit averages over batch-2 runs (peak 37.7–51.6 GiB).

Checkpoint selection (48-pair fixed subset, 20 steps, <doodle> prompt)

Checkpoint OWLv2 hit % bg PSNR ink leftover
ckpt500 64.6 17.97 0.0009
ckpt1000 56.2 18.04 0.0
ckpt1500 54.2 18.82 0.0
ckpt2000 (final) 62.5 17.40 0.0005

ckpt500 selected β€” best hit rate with PSNR within ~0.9 dB of the best; the mid-training checkpoints trade hit rate away before it partially recovers at 2000.

Breakdowns (hit rate %, 160 pairs)

Slice ours 40 ours t6 base 40
style: outline 65.8 71.2 65.8
style: detailed 62.0 66.0 64.0
style: blob 64.9 62.2 73.0
size: small (<8% of image) 65.5 70.9 65.5
size: medium (8–20%) 64.9 64.9 67.6
size: large (β‰₯20%) 61.3 67.7 67.7
held-out classes (n=40) 65.0 67.5 67.5
seen classes (n=120) 64.2 67.5 66.7

No held-out degradation: the LoRA generalizes to unseen categories (65.0% held-out vs 64.2% seen at 40 steps; 67.5% / 67.5% at turbo). Weakest slice for the LoRA is blob scribbles at 40 steps (64.9% vs the base's 73.0%) β€” the blob style carries the least shape signal, so it benefits least from the marker-trained prior. Full per-slice metrics (PSNR, LPIPS, SAM, ink) are in eval/breakdowns.json.

Realism vote (blind, 100 pairs, judge Qwen2.5-VL-72B)

LoRA 44 β€” base 56 (0 errored pairs). A mild VLM preference for the base model's edits on "which looks more naturally part of the photo" β€” reported as measured.

Out-of-domain check

12 clean scene photos generated with the base model's text-to-image mode (empty kitchen counter Γ—2, living room, park bench Γ—2, office desk Γ—2, beach Γ—2, bedroom nightstand, bathroom counter). Scribbles drawn on empty areas with scribble.py, using outlines borrowed from test-pair ground-truth masks; captions: 6 normal (a ceramic coffee mug, a potted plant, a sleeping cat, a pair of headphones, a beach ball, a table lamp) and 6 absurd (a small dinosaur, a vintage robot, a glass elephant, a teapot with legs, a tiny spaceship, a banana wearing a hat). Each scene Γ— 4 configs (ours_40, base_40, ours_t6, base_t6) = 48 edits.

Contact sheets (input + 4 outputs per scene): eval/ood/sheet_ood_00.jpg … sheet_ood_11.jpg in the dataset repo, with per-scene scribbles and captions in ood_metadata.jsonl.

Known OOD caveats, stated honestly: we did not run automated metrics on these 12 scenes (the 160-pair in-domain numbers above are the quantitative proof) β€” the contact sheets are the check. Watch specifically for: (1) missing or wrong contact shadows β€” training inputs always had a LaMa-filled patch under the scribble, so shadow behaviour on clean empty surfaces is the least-verified property; (2) wrong scale for the scene for absurd requests, where the LoRA may fall back to scale priors of the nearest training class; (3) background redrawing outside the scribble neighbourhood, which shows up as LPIPS/PSNR failures in-domain and should be judged visually here.

Limitations

  • Single-object insertion only; one scribble per edit (the convention and training data are one-object-per-pair).
  • Qwen-Image-2.1's output resolution snaps to multiples of 32, so edits at non-multiple-of-32 input sizes are internally resized; metrics were computed at the snapped output resolution against the correspondingly resized reference.
  • Trained at 768 px; 1024 px was tested in planning but exceeded the training budget including text-embedding caching.
  • The realism vote mildly favours the base model; hit-rate parity at 40 steps means the LoRA's value is concentrated in the 6-step turbo regime and in the shared-Space magenta-marker convention, not in beating the base model everywhere.
  • Non-commercial license (see below) β€” inherited from the Qwen Research License.

Attribution

  • Open Images V7 (train + validation splits): segmentation annotations CC BY 4.0; images listed as CC BY 2.0. Per-image credits (author, original URL, license) are stored in the dataset repo's metadata for every pair: ML-Intern-lab/doodle-in-pairs.
  • LaMa inpainting (via simple-lama-inpainting, Apache 2.0) β€” used only to build training inputs (hole filling under the scribble).
  • Marker idea credited to the red-highlight object-remover LoRAs (e.g. prithivMLmods/QIE-2511-Object-Remover-v2, "Remove the red highlighted object from the scene"); our marker colour is changed to magenta (255, 0, 255) vs their red (239, 68, 68) so both can share one Space.
  • Viggle turbo (Viggle/Qwen-Image-2.1-viggle-turbo) β€” used for the 6-step evaluation, not in training.
  • Metrics: OWLv2, SAM 2.1, LPIPS (AlexNet), Qwen/Qwen2.5-VL-72B-Instruct (captions and realism vote) via Inference Providers.

License

This adapter is a derivative of Qwen-Image-2.1 and is distributed under the Qwen RESEARCH LICENSE AGREEMENT (Release Date: September 20, 2026) β€” non-commercial, research/evaluation use only. The full text is in LICENSE in this repo (copied from the base model). Commercial use requires a separate license from Qwen (model-business@notice.qwencloud.com).

Required attribution notice, retained verbatim:

"Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."

Change notice: the modified files in this distribution are the LoRA adapter safetensors (doodle_in_lora_qwen21*.safetensors), convert_lora.py, config.yaml and optimizer.pt β€” a low-rank adapter and its tooling trained on top of Qwen-Image-2.1; no base-model weights or files were modified.

Built with Qwen.

Downloads last month
576
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA

Adapter
(98)
this model

Dataset used to train ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA

Spaces using ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA 3

Collection including ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA