| --- |
| license: cc-by-nc-4.0 |
| base_model: stabilityai/stable-diffusion-3-medium-diffusers |
| library_name: peft |
| pipeline_tag: image-to-image |
| language: en |
| tags: |
| - super-resolution |
| - image-super-resolution |
| - extreme-zoom |
| - chain-of-zoom |
| - diffusion |
| - lora |
| - peft |
| - privileged-distillation |
| - faithfulness |
| --- |
| |
| # OracleZoom |
|
|
| **Privileged-Latent Distillation for faithful extreme super-resolution.** |
| A tiny (7.1M-parameter) LoRA adapter that makes Chain-of-Zoom's recursive super-resolution add *faithful* detail instead of hallucinating, all the way to 256x. |
|
|
| [](https://github.com/dipta007/OPD-Zoom) |
| [](https://huggingface.co/stabilityai/stable-diffusion-3-medium-diffusers) |
| [](https://github.com/dipta007/OPD-Zoom) |
| [-red)](https://github.com/dipta007/OPD-Zoom) |
| [](https://creativecommons.org/licenses/by-nc/4.0/) |
|
|
| ## Highlights |
| - **Faithful, not just sharp.** At extreme zoom the backbone must *invent* detail; the question is whether it is faithful or hallucinated. This adapter teaches faithfulness. |
| - **Holds where baselines collapse.** CLIPIQA **0.71 at 256x** while Chain-of-Zoom (CoZ) and five SOTA SR backbones fall to <=0.58; most faithful of all methods at 4x (LPIPS **0.20** vs CoZ 0.22). |
| - **Judged more faithful.** Two cross-family vision-language judges (InternVL + Gemini) prefer this zoom **68-78%** of the time at 64-256x and flag the strongest baseline hallucinating **2-5x more**. |
| - **Tiny and drop-in.** A rank-16 LoRA (**7.1M** trainable params) trained on only **1,000** curated 4K images; it drops into CoZ's recursion with no other changes. |
|
|
| ## Model Overview |
| | Property | Value | |
| |---|---| |
| | Model type | LoRA adapter (PEFT) for a one-step SR backbone | |
| | Backbone (frozen) | OSEDiff on Stable Diffusion 3-medium | |
| | Prompt extractor (frozen) | Qwen2.5-VL-3B-Instruct | |
| | Trainable params | 7.1M | |
| | LoRA | r = 16, alpha = 32, dropout = 0.0 | |
| | LoRA targets | `to_q, to_k, to_v, add_q_proj, add_k_proj, add_v_proj` (SD3 transformer) | |
| | Training data | 1,000 curated 4K photographs (supervised at 4x only) | |
| | Objective | decode-space LPIPS + anchored cycle-consistency - beta_reward * TOPIQ-NR + beta_kl * KL-to-base + EMA | |
| | Key weights | beta_reward 0.4, beta_kl 8.0, w_cyc 1.0, lambda_ema 0.1 (EMA decay 0.95) | |
| | Recursion at test | 4 steps (4x / 16x / 64x / 256x), 512x512 center crop | |
|
|
| ## Method |
| Recursive SR (Chain-of-Zoom) reuses a 4x backbone step after step to reach 16x-256x. Each step is **blind**: it sees only a blurred crop of its own previous output and must invent the missing detail, so errors compound and the invention may be hallucinated. |
|
|
| **Privileged-latent distillation (the idea).** A *privileged teacher* is shown the ground-truth high-resolution patch **at training time only** and distills its real detail into the blind student, in **decode space** (a perceptual loss between the student's decoded image and the real patch). Only a small LoRA adapter is trained; the backbone, VAE, and prompter stay frozen. |
|
|
| **A KL leash keeps the deep reward faithful (the safeguard).** Ground truth exists only at 4x. To carry the distilled prior into the deeper recursion, the student chases a differentiable detail reward (TOPIQ-NR) through the real zoom. Left free, such a reward games the metric with a repetitive crosshatch; **leashed** to the deployed backbone by a KL trust region (a latent distance for a one-step map), it sharpens detail without drifting into hallucination. |
|
|
| Full derivation and ablations are in the [OPD-Zoom repo](https://github.com/dipta007/OPD-Zoom). |
|
|
| ## Quickstart |
| This is the trained artifact of OracleZoom: a PEFT LoRA on the SD3 transformer that OSEDiff uses as the SR backbone inside Chain-of-Zoom. It plugs into the OPD-Zoom pipeline. |
|
|
| ```bash |
| # 1) get the pipeline |
| git clone https://github.com/dipta007/OPD-Zoom && cd OPD-Zoom |
| |
| # 2) get this adapter |
| huggingface-cli download dipta007/OracleZoom --local-dir ckpt/OracleZoom |
| |
| # 3) run the 4-recursion zoom with the adapter as the student SR LoRA |
| python -m opd_zoom.teacher.oracle_infer \ |
| --mode student --pld_lora ckpt/OracleZoom \ |
| --gt_dir <your_images> --out <out_dir> --rec_num 4 |
| ``` |
|
|
| Loading just the adapter with PEFT: |
| ```python |
| from peft import PeftModel |
| # `sd3_transformer` is the SD3Transformer2DModel used by the OSEDiff backbone |
| model = PeftModel.from_pretrained(sd3_transformer, "dipta007/OracleZoom") |
| ``` |
| The adapter is `inference_mode` and merges into the backbone at no added latency; the VLM prompter is unchanged, so per-image inference cost equals Chain-of-Zoom's. |
|
|
| ## Results |
| Under Chain-of-Zoom's exact protocol on a curated 4K benchmark and six test sets (in-domain 4K, DIV8K, DRealSR, RealSR, FFHQ, Flickr2K): |
|
|
| | Axis | Metric | Ours | CoZ / best baseline | |
| |---|---|---|---| |
| | Sharpness (no-reference) | CLIPIQA @256x | **0.71** | <= 0.58 | |
| | Fidelity @4x (ground truth exists) | LPIPS | **0.20** | 0.22 (CoZ) | |
| | Deep faithfulness (MLLM judge, 64-256x) | preferred vs CoZ | **68-78%** | - | |
| | Deep faithfulness | hallucination rate vs CoZ | **2-5x lower** | - | |
|
|
| Sharpness is the axis prior methods are built for; the decisive gap is **faithfulness**, verified by full-reference metrics at 4x and by two cross-family MLLM judges plus a blinded human study past 4x. |
|
|
| ## Intended Use |
| - **In-scope:** research on faithful extreme (recursive) super-resolution; as the SR-backbone adapter inside the Chain-of-Zoom recursion on natural photographs. |
| - **Out-of-scope:** a standalone single-shot SR model (it is a drop-in LoRA for the CoZ loop, not a full model); forensic or evidentiary use (detail past 4x is generated, not recovered); real-camera-zoom claims (the benchmark uses synthetic center-crop zoom). |
|
|
| ## Training |
| Early-stopped on held-out validation at ~epoch 37 (step 9300); best val 0.216. Trained on one 8xH200 node (single GPU trains the adapter). Full config in `train_meta.json` and the [repo](https://github.com/dipta007/OPD-Zoom). |
|
|
| ## Citation |
| ```bibtex |
| @inproceedings{dipta2027oraclezoom, |
| title={OracleZoom: Privileged-Latent Distillation for Faithful Extreme Super-Resolution}, |
| author={Shubhashis Roy Dipta}, |
| year={2027}, |
| note={In submission, WACV 2027}, |
| url={https://github.com/dipta007/OPD-Zoom} |
| } |
| ``` |
| Please also cite Chain-of-Zoom and OSEDiff, whose components this builds on. |
|
|
| ## License |
| Released for **research, non-commercial** use (CC-BY-NC-4.0). This adapter is trained on top of OSEDiff / Stable Diffusion 3 and used with a Qwen2.5-VL prompter inside Chain-of-Zoom; the respective upstream licenses apply to those components. |
|
|