# Qwen-Image-2.1-viewpoint-orbit-LoRA **An RGBA viewpoint-orbit LoRA for [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1).** Give it ONE transparent (RGBA) image of an object and a relative camera instruction; it returns the same object from the requested viewpoint — also as a transparent RGBA image. - **Base model:** [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), bf16, unquantized - **Adapter:** rank 32 / alpha 32, trained 2,000 steps at 768 px (~2.2 s/step on one A100-80GB) with [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit), arch `qwen_image_2`, `rgba: true` - **Training data:** synthetic orbit renders of Google Scanned Objects → [ML-Intern-lab/gso-orbit-rgba](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba) (renders, pose metadata, pair lists and split all included) Try it in the Space: [ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA](https://huggingface.co/spaces/ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA). ## Usage (verified) This exact path was run end-to-end during evaluation (`verify_diffusers_lora.py` on the pinned diffusers commit; loads with **zero dropped-key warnings**, output is RGBA): ```python import torch from huggingface_hub import hf_hub_download from diffusers import QwenImage21Pipeline pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda") lora = hf_hub_download("ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA", "checkpoints/steps2000res768/orbit_alpha_lora_gate_up_split.safetensors") pipe.load_lora_weights(lora) # diffusers-keyed file: zero dropped keys src = ... # your source image, 768x768 RGBA PIL.Image out = pipe(image=src, width=768, height=768, num_inference_steps=40, true_cfg_scale=1.0, prompt=" rotate the camera 90 degrees to the right, eye level. " "The image has alpha channel and the background is transparent.").images[0] out.save("orbit.png") # out.mode == "RGBA" — transparency is preserved ``` Notes: - `true_cfg_scale=1.0` (no CFG) is how the model was evaluated and how it is meant to be run here. - The output is RGBA because the 2.1 VAE is natively RGBA — no matting post-processing involved. - The file `orbit_alpha_lora_gate_up_split.safetensors` is the **diffusers-keyed** adapter. The original trainer output (`checkpoints/steps2000res768/orbit_alpha_lora/orbit_alpha_lora.safetensors`) uses ComfyUI-style keys and mostly loads in diffusers, but silently drops the `img_mlp.gate_up` branch (a fused SwiGLU projection diffusers stores as two separate linears). Use the `_gate_up_split` file; both are verified to give matching metrics. - Verified against `diffusers` @ commit `0121a91f9d419ff7234c8a5923f82c244e6f1914` with `transformers>=5.17`, `peft`, `torch 2.13+`. ## Instruction grammar Azimuth is **relative to the source view** (scanned objects — and user uploads — have no canonical front); elevation is absolute. Every instruction starts with ``. The full training set is exactly these 23 instructions (80–81 pairs each, 1,844 pairs total): | | low angle | eye level | elevated | |---|---|---|---| | rotate 45° left | ✓ | ✓ | ✓ | | rotate 90° left | ✓ | ✓ | ✓ | | rotate 135° left | ✓ | ✓ | ✓ | | rotate 45° right | ✓ | ✓ | ✓ | | rotate 90° right | ✓ | ✓ | ✓ | | rotate 135° right | ✓ | ✓ | ✓ | | rotate 180 degrees (no left/right) | ✓ | ✓ | ✓ | | keep the camera angle (elevation only) | ✓ (low angle) | — | ✓ (elevated) | Verbatim forms: - ` rotate the camera {45|90|135|180} degrees to the {left|right}, {low angle|eye level|elevated}` (180 omits left/right: ` rotate the camera 180 degrees, eye level`) - ` keep the camera angle, {low angle|elevated}` (20% of training pairs) At inference we append `The image has alpha channel and the background is transparent.` — chosen by an A/B on the base model (alpha IoU 0.830 with the suffix vs 0.798 without, on 10 pairs). ## Results vs the base model (ground truth) 160 held-out edits (40 real-scanned objects × 4 targets), 40 inference steps, 768 px, fixed per-pair seeds, metrics against the true renders (LPIPS/PSNR composited on mid-gray): | config | alpha IoU ↑ | LPIPS ↓ | PSNR ↑ | DINOv2 sim ↑ | |---|---|---|---|---| | base model, plain grammar ("rotate the camera 90 degrees to the right") | 0.727 | 0.097 | 21.74 | 0.714 | | base model, `` grammar | 0.731 | 0.099 | 21.74 | 0.713 | | **this LoRA (step-2000 checkpoint)** | **0.794** | **0.090** | **22.34** | **0.734** | By rotation size (alpha IoU): 45° 0.835, 90° **0.773** (base: ~0.67 — the hard bucket, where hidden sides must be invented), 180° 0.835. The step-2000 checkpoint was chosen by best LPIPS with no alpha-IoU loss over {500, 1000, 1500, 2000} on a fixed 48-edit subset; it won on every metric, so no trade-off arose. Per-edit records: `eval/final/`, `eval/ckpt_select/`, `eval/baseline/` in [ML-Intern-lab/gso-orbit-rgba](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba). ## Out-of-domain check 12 text-to-image-generated transparent subjects (cartoon mascot, sneaker, armchair, robot toy, potted plant, game character, perfume bottle, backpack, headphones, coffee mug, desk lamp, rubber duck), orbited 7 views each at eye level by chaining 45° moves — every edit sourced from the **original** image, never from a generated one. Contact sheets and ping-pong turntable GIFs: `eval/ood/turntable_*.gif` in [ML-Intern-lab/gso-orbit-rgba](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba) — e.g. [turntable_sneaker.gif](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba/blob/main/eval/ood/turntable_sneaker.gif), [turntable_robot_toy.gif](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba/blob/main/eval/ood/turntable_robot_toy.gif), [turntable_cartoon_mascot.gif](https://huggingface.co/datasets/ML-Intern-lab/gso-orbit-rgba/blob/main/eval/ood/turntable_cartoon_mascot.gif). Where it breaks, honestly: - **Transparency holds** — per-view alpha coverage stayed at 0.45–0.49 across all 84 views (no mask collapse, no background bleed-through), and no view dropped below 5% coverage. - **Identity drifts with azimuth.** Small 45° steps are faithful; by cumulative 180–315° (chained as legal 135°/90°/45° instructions) silhouettes stay plausible but fine details (logos, straps, handles) sometimes get re-invented rather than preserved. There is no ground truth for these subjects, so this is a qualitative observation from the contact sheets, not a measured number. - **Style transfer of lighting:** the model was trained on clean studio-lit scans; on T2I subjects with baked-in lighting it occasionally carries the source's shading into the new view instead of re-lighting. ## Limitations - Trained exclusively on Google Scanned Objects (household-scale, studio-lit, no ground plane). Performance on other scales/materials is unmeasured; see the OOD section. - Relative azimuth only — there are no absolute poses ("front view") by design; the camera always moves relative to the input view. - Elevation is limited to the trained range (low angle ≈ −20°, eye level, elevated ≈ +40°). - 90° moves are the hardest bucket (hidden sides must be invented); good, but measurably below small rotations. - **Non-commercial** — see License. ## Training details - Data: 24 rendered views (8 azimuths × 3 elevations) per object, 768×768 RGBA, transparent background, three-point lighting, Z-up handled, objects framed to 55–65% of frame; objects whose any-view coverage fell under 8% of the frame were rejected (529 of 1,030 rejected → 501 usable; 461 train / 40 held-out, stratified by category). - Pairs: 4 per training object (1,844), source always eye-level at a random azimuth; target = source azimuth + one of the 7 relative moves (or same azimuth for the 20% elevation-only pairs) at one of 3 elevations; instruction balance 80/81 across all 23. - Trainer: ai-toolkit `sd_trainer`, `control_path` editing, `rgba: true`, `match_target_res` default, bf16 unquantized (`Qwen/Qwen-Image-2.1`), adamw8bit, gradient checkpointing, cached text embeddings, `timestep_type: weighted`, lr 1e-4, batch 1, 2,000 steps, checkpoints every 500. - Loss curve: `checkpoints/steps2000res768/loss_log.jsonl` in this repo (2,000 steps, ~0.01–0.2, no divergence). ## License This adapter is a derivative of Qwen-Image-2.1 and is distributed under the **Qwen RESEARCH LICENSE AGREEMENT — non-commercial use only**. The verbatim license text is included in this repo as [LICENSE](https://huggingface.co/ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA/blob/main/LICENSE). **Change notice** (per license §3b): this repo adds only new LoRA adapter weight files (`safetensors`) trained on top of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1); no base-model files are modified or redistributed. **Built with Qwen.** "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved." The repo name uses "Qwen-Image-2.1" descriptively, as a fine-tune of Qwen-Image-2.1, which license §4c permits; the adapter itself is called Orbit Alpha (file names keep the `orbit_alpha_lora` prefix). ## Attribution - **Base model:** [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) (Qwen Research License). - **Training data:** Google Scanned Objects ([original dataset](https://github.com/GoogleCloudPlatform/tensorflow-graphics/blob/master/docs/scanned_objects.md)), obtained via [suvadityamuk/google-scanned-objects](https://huggingface.co/datasets/suvadityamuk/google-scanned-objects), **CC-BY 4.0**. Renders produced with trimesh + pyrender. - **Camera-LoRA idea:** credit to fal ([fal/Qwen-Image-Edit-2511-Multiple-Angles-LoRA](https://huggingface.co/fal/Qwen-Image-Edit-2511-Multiple-Angles-LoRA)) and dx8152 for the multiple-angles LoRA concept. The relative-azimuth grammar here is deliberately different from fal's absolute poses. - **Tooling:** [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit) for training, [diffusers](https://github.com/huggingface/diffusers) `QwenImage21Pipeline` for inference/eval, LPIPS + DINOv2 for metrics.