Instructions to use ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
Qwen-Image-2.1 Doodle-in LoRA
A LoRA adapter for Qwen/Qwen-Image-2.1 that takes one photo with a magenta scribble drawn on it plus a short text naming an object, and returns the same photo with the scribble replaced by that object β following the scribble's shape and pose, lit consistently with the scene, with everything outside the scribble's neighbourhood unchanged.
- Base model: Qwen/Qwen-Image-2.1 (one model for generation and editing; editing mode is triggered by passing
image=) - Training data: 6,042 pairs built from Open Images V7 β
ML-Intern-lab/doodle-in-pairs - Trainer: ostris/ai-toolkit
sd_trainer,arch: qwen_image_2, @ commit6468a2f - Demo Space:
ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRAβ draw a magenta outline, name the object, compare checkpoints and the base model - Evaluation: 160-pair test set (40 pairs from 23 held-out classes) against the base model at 40 steps and with the Viggle 6-step turbo LoRA
Marker convention (exact β a Space can reproduce the training inputs byte-for-byte)
Implemented standalone in scribble.py in the dataset repo.
| Property | Value |
|---|---|
| Colour | pure magenta RGB (255, 0, 255), opaque |
| Stroke width | uniform(0.004, 0.010) Γ short_side, clamped to [3, 7] px (i.e. 0.4β1.0% of the short side) |
| Caps / joints | round: ImageDraw.line(..., joint="curve") plus filled end circles of radius width/2 |
| Drawing | directly on the photo, single colour, no blending |
Prompt template: <doodle> Turn the magenta scribble into {caption}. β e.g. <doodle> Turn the magenta scribble into a brown leather armchair.
Magenta was chosen deliberately to differ from the red (239, 68, 68) used by object-remover LoRAs, so a remover LoRA and this insert LoRA can share one Space.
Three scribble styles were mixed 50/30/20 in training:
- outline (50%) β mask contour simplified with DouglasβPeucker (
eps β [0.2%, 0.6%]of the contour perimeter), wobbled with smooth sinusoidal noise (amplitude 0.5β2% of the short side), with occasional small gaps (0β2 segments removed) and overshoots (up to 1.5% of the short side, 40% per endpoint) - detailed (30%) β outline plus up to 4 interior strokes taken from a Canny map of the object inside the eroded mask, sparsified to β€40 points and wobbled
- blob (20%) β a jittered convex hull resampled to 27 angular samples (radius jitter Γ0.92β1.06, Gaussian noise 0.8% of the short side)
Training pairs additionally had the object mask dilated by 3% of the bounding-box diagonal (extended downwards for contact shadows) and filled with LaMa before the scribble was drawn.
Files in this repo
| File | What it is |
|---|---|
doodle_in_lora_qwen21_gate_up_split.safetensors |
Recommended. Diffusers-keyed adapter. ai-toolkit saves a fused img_mlp.gate_up LoRA that diffusers silently drops (it stores gate/up as two linears); this file has the fused tensor split, verified delta-identical to the original (max deviation exactly 0.0 on all 448 modules). |
doodle_in_lora_qwen21.safetensors |
Original ai-toolkit (ComfyUI-style) keys. Load this in ComfyUI, or use convert_lora.py to make your own split file. |
doodle_in_lora_qwen21_000000500.safetensors (+ 1000, 1500) |
Intermediate checkpoints. Checkpoint 500 is the selected one (see evaluation). |
convert_lora.py |
The gate/up split converter + numeric verification. |
config.yaml |
The exact ai-toolkit training config. |
optimizer.pt |
Optimizer state (for resuming). |
checkpoints/ |
Checkpoint copies with sibling configs. |
Usage (diffusers)
Verified with QwenImage21Pipeline.load_lora_weights at diffusers commit 0121a91f9d419ff7234c8a5923f82c244e6f1914, transformers==5.17.0, torch==2.14.0 β the exact code below (modulo repo ids) is what our evaluation jobs executed.
40-step edit with the LoRA
import torch
from PIL import Image
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", dtype=torch.bfloat16).to("cuda")
pipe.load_lora_weights(
"ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA",
weight_name="doodle_in_lora_qwen21_gate_up_split.safetensors",
adapter_name="doodle")
# sanity check used in our eval: a correct load has 448 lora modules;
# a load that dropped the fused gate_up keys has fewer.
n = sum(1 for name, _ in pipe.transformer.named_modules() if "lora_A" in name)
assert n == 448, f"adapter loaded with dropped keys: got {n}, expect 448"
img = pipe(
prompt="<doodle> Turn the magenta scribble into a brown leather armchair.",
image=Image.open("input.jpg").convert("RGB"),
num_inference_steps=40,
true_cfg_scale=1.0, # no CFG, as on the model card
width=1024, height=768, # multiples of 32; match your input aspect
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
6-step turbo (both adapters stacked)
The Viggle turbo LoRA cuts the same edit to 6 steps. Stack both adapters and use the turbo sigmas and scheduler:
from diffusers import FlowMatchEulerDiscreteScheduler
from huggingface_hub import hf_hub_download
base_cfg = dict(pipe.scheduler.config) # keep BEFORE switching
pipe.load_lora_weights(
hf_hub_download("Viggle/Qwen-Image-2.1-viggle-turbo",
"Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors"),
adapter_name="turbo")
pipe.set_adapters(["doodle", "turbo"], [1.0, 1.0])
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(
base_cfg, shift_terminal=None)
img = pipe(
prompt="<doodle> Turn the magenta scribble into a brown leather armchair.",
image=Image.open("input.jpg").convert("RGB"),
num_inference_steps=6,
true_cfg_scale=1.0,
sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25],
width=1024, height=768,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
To go back to 40 steps, unload or reset adapters and restore the base scheduler: pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(base_cfg) (the turbo branch sets shift_terminal=None, which must not leak into non-turbo runs).
Optional speed fix (same outputs, ~30 s β ~3 ms per ref image on Qwen3-VL's patch-embed conv):
pe = pipe.text_encoder.model.visual.patch_embed
w = pe.proj.weight.reshape(pe.proj.weight.shape[0], -1)
pe.forward = lambda hs: hs.to(w.dtype).flatten(1) @ w.T + pe.proj.bias
Training
Trained with ostris/ai-toolkit sd_trainer, arch: qwen_image_2, from the Hub id (loading from a local diffusers dir can crash with "Cannot copy out of meta tensor" β ai-toolkit issue #1067):
- LoRA
linear: 32,linear_alpha: 32;optimizer: adamw8bit,lr: 1e-4,batch_size: 1 timestep_type: weighted,noise_scheduler: flowmatch,gradient_checkpointing: true,dtype: bf16,quantize: false,quantize_te: false- dataset:
folder_path(targets) +control_path(magenta-marked inputs) +caption_ext: txt,resolution: [768],cache_latents_to_disk: true,cache_text_embeddings: true, sampling disabled (issue #1059) - 2,000 steps at 768 px; measured 2.15 s/step on one A100 80 GB; training loss 1.197 β a stable ~0.07β0.3 band (final-step loss 0.077); text-embedding caching (Qwen3-VL over the control images) covered 4,773 items in 1,055 s
- final adapter check: 192/192
lora_Btensors non-zero, max|B| = 0.052 - full run wall-clock 1 h 38 m on A100-large
Evaluation
Setup. 160 test pairs from Open Images validation (73 outline / 50 detailed / 37 blob; 40 pairs from 23 held-out classes never seen in training). Fixed seeds (sha1(pair_id, 42) per pair). Metrics per edit:
- Object present β OWLv2 (
google/owlv2-base-patch16-ensemble) queried with the class name; a hit = score above threshold and box IoU β₯ 0.3 with the scribble's bounding box. - Shape fidelity β SAM 2.1 (
facebook/sam2.1-hiera-small) prompted with the scribble bbox on the output; mask IoU against the ground-truth object mask (and against the dilated scribble region). - Ink removed β fraction of the scribble's stroke pixels still near-magenta in the output (lower is better; near zero everywhere).
- Background preserved β PSNR and LPIPS against the input, outside the union of the dilated scribble region and the LaMa hole.
- Realism β blind, randomised-order pairwise VLM judgement (ours vs base on the same input), 100 pairs, judge
Qwen/Qwen2.5-VL-72B-Instructvia Inference Providers.
All metrics computed 160/160 with zero scoring errors; raw per-item files are in eval/full/ckpt500/ of the dataset repo, alongside the generated images.
Main comparison β 160 test pairs
| Config | Steps | Prompt | OWLv2 hit % | bg PSNR β | LPIPS β | SAM IoU (vs GT) β | ink leftover β | s/edit β |
|---|---|---|---|---|---|---|---|---|
| ours (ckpt500) | 40 | <doodle> |
64.4 | 17.38 | 0.3644 | 0.5701 | 0.0025 | 20.0 |
| base (a) | 40 | plain instruction | 66.9 | 18.55 | 0.3578 | 0.5790 | 0.0008 | 18.3 |
| base (b) | 40 | <doodle> |
66.2 | 18.34 | 0.3687 | 0.5703 | 0.0082 | 18.2 |
| ours + Viggle turbo | 6 | <doodle> |
67.5 | 17.63 | 0.3509 | 0.5809 | 0.0026 | 4.67 |
| base + Viggle turbo (c) | 6 | plain instruction | 65.6 | 17.98 | 0.3628 | 0.5683 | 0.0031 | 4.36 |
| base + Viggle turbo | 6 | <doodle> |
65.6 | 17.34 | 0.3823 | 0.5507 | 0.0091 | 4.35 |
Honest headline: our LoRA does not beat the base model on hit rate at 40 steps (64.4% vs 66.9%) β the base model can already follow painted annotations reasonably well, as its card suggests. Where the LoRA does win: at 6 turbo steps it scores the highest hit rate of all six configs (67.5%) and the best LPIPS, and the base-with-<doodle>-prompt combo shows measurably more leftover ink (0.8β0.9%) than ours at the same step counts. The practical takeaway is that the LoRA + turbo stack gives you the best quality at 4.67 s/edit; if you only ever run 40 steps, the plain base model is a competitive baseline.
Single-edit timings on the same hardware (A100 80 GB, from the smoke test): 13.7 s/edit at 40 steps (peak 37.3 GiB), 3.2 s/edit at 6 turbo steps (peak 38.6 GiB). Batched-eval means above are per-edit averages over batch-2 runs (peak 37.7β51.6 GiB).
Checkpoint selection (48-pair fixed subset, 20 steps, <doodle> prompt)
| Checkpoint | OWLv2 hit % | bg PSNR | ink leftover |
|---|---|---|---|
| ckpt500 | 64.6 | 17.97 | 0.0009 |
| ckpt1000 | 56.2 | 18.04 | 0.0 |
| ckpt1500 | 54.2 | 18.82 | 0.0 |
| ckpt2000 (final) | 62.5 | 17.40 | 0.0005 |
ckpt500 selected β best hit rate with PSNR within ~0.9 dB of the best; the mid-training checkpoints trade hit rate away before it partially recovers at 2000.
Breakdowns (hit rate %, 160 pairs)
| Slice | ours 40 | ours t6 | base 40 |
|---|---|---|---|
| style: outline | 65.8 | 71.2 | 65.8 |
| style: detailed | 62.0 | 66.0 | 64.0 |
| style: blob | 64.9 | 62.2 | 73.0 |
| size: small (<8% of image) | 65.5 | 70.9 | 65.5 |
| size: medium (8β20%) | 64.9 | 64.9 | 67.6 |
| size: large (β₯20%) | 61.3 | 67.7 | 67.7 |
| held-out classes (n=40) | 65.0 | 67.5 | 67.5 |
| seen classes (n=120) | 64.2 | 67.5 | 66.7 |
No held-out degradation: the LoRA generalizes to unseen categories (65.0% held-out vs 64.2% seen at 40 steps; 67.5% / 67.5% at turbo). Weakest slice for the LoRA is blob scribbles at 40 steps (64.9% vs the base's 73.0%) β the blob style carries the least shape signal, so it benefits least from the marker-trained prior. Full per-slice metrics (PSNR, LPIPS, SAM, ink) are in eval/breakdowns.json.
Realism vote (blind, 100 pairs, judge Qwen2.5-VL-72B)
LoRA 44 β base 56 (0 errored pairs). A mild VLM preference for the base model's edits on "which looks more naturally part of the photo" β reported as measured.
Out-of-domain check
12 clean scene photos generated with the base model's text-to-image mode (empty kitchen counter Γ2, living room, park bench Γ2, office desk Γ2, beach Γ2, bedroom nightstand, bathroom counter). Scribbles drawn on empty areas with scribble.py, using outlines borrowed from test-pair ground-truth masks; captions: 6 normal (a ceramic coffee mug, a potted plant, a sleeping cat, a pair of headphones, a beach ball, a table lamp) and 6 absurd (a small dinosaur, a vintage robot, a glass elephant, a teapot with legs, a tiny spaceship, a banana wearing a hat). Each scene Γ 4 configs (ours_40, base_40, ours_t6, base_t6) = 48 edits.
Contact sheets (input + 4 outputs per scene): eval/ood/sheet_ood_00.jpg β¦ sheet_ood_11.jpg in the dataset repo, with per-scene scribbles and captions in ood_metadata.jsonl.
Known OOD caveats, stated honestly: we did not run automated metrics on these 12 scenes (the 160-pair in-domain numbers above are the quantitative proof) β the contact sheets are the check. Watch specifically for: (1) missing or wrong contact shadows β training inputs always had a LaMa-filled patch under the scribble, so shadow behaviour on clean empty surfaces is the least-verified property; (2) wrong scale for the scene for absurd requests, where the LoRA may fall back to scale priors of the nearest training class; (3) background redrawing outside the scribble neighbourhood, which shows up as LPIPS/PSNR failures in-domain and should be judged visually here.
Limitations
- Single-object insertion only; one scribble per edit (the convention and training data are one-object-per-pair).
- Qwen-Image-2.1's output resolution snaps to multiples of 32, so edits at non-multiple-of-32 input sizes are internally resized; metrics were computed at the snapped output resolution against the correspondingly resized reference.
- Trained at 768 px; 1024 px was tested in planning but exceeded the training budget including text-embedding caching.
- The realism vote mildly favours the base model; hit-rate parity at 40 steps means the LoRA's value is concentrated in the 6-step turbo regime and in the shared-Space magenta-marker convention, not in beating the base model everywhere.
- Non-commercial license (see below) β inherited from the Qwen Research License.
Attribution
- Open Images V7 (train + validation splits): segmentation annotations CC BY 4.0; images listed as CC BY 2.0. Per-image credits (author, original URL, license) are stored in the dataset repo's metadata for every pair:
ML-Intern-lab/doodle-in-pairs. - LaMa inpainting (via
simple-lama-inpainting, Apache 2.0) β used only to build training inputs (hole filling under the scribble). - Marker idea credited to the red-highlight object-remover LoRAs (e.g.
prithivMLmods/QIE-2511-Object-Remover-v2, "Remove the red highlighted object from the scene"); our marker colour is changed to magenta (255, 0, 255) vs their red (239, 68, 68) so both can share one Space. - Viggle turbo (
Viggle/Qwen-Image-2.1-viggle-turbo) β used for the 6-step evaluation, not in training. - Metrics: OWLv2, SAM 2.1, LPIPS (AlexNet),
Qwen/Qwen2.5-VL-72B-Instruct(captions and realism vote) via Inference Providers.
License
This adapter is a derivative of Qwen-Image-2.1 and is distributed under the Qwen RESEARCH LICENSE AGREEMENT (Release Date: September 20, 2026) β non-commercial, research/evaluation use only. The full text is in LICENSE in this repo (copied from the base model). Commercial use requires a separate license from Qwen (model-business@notice.qwencloud.com).
Required attribution notice, retained verbatim:
"Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."
Change notice: the modified files in this distribution are the LoRA adapter safetensors (doodle_in_lora_qwen21*.safetensors), convert_lora.py, config.yaml and optimizer.pt β a low-rank adapter and its tooling trained on top of Qwen-Image-2.1; no base-model weights or files were modified.
Built with Qwen.
- Downloads last month
- 576
Model tree for ML-Intern-lab/Qwen-Image-2.1-doodle-in-LoRA
Base model
Qwen/Qwen-Image-2.1