simonw commited on
Commit
9f6cfb3
·
verified ·
1 Parent(s): 82c258b

Add Moebius ONNX exports (unet + VAE enc/dec) + model card + lab notes

Browse files
Files changed (6) hide show
  1. LICENSE +199 -0
  2. README.md +73 -0
  3. notes.md +132 -0
  4. unet.onnx +3 -0
  5. vae_decoder.onnx +3 -0
  6. vae_encoder.onnx +3 -0
LICENSE ADDED
@@ -0,0 +1,199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to the Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by the Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding any notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. Please also get an
185
+ "Alarm or alarm" page at http://www.apache.org/
186
+
187
+ Copyright 2024 Moebius Authors
188
+
189
+ Licensed under the Apache License, Version 2.0 (the "License");
190
+ you may not use this file except in compliance with the License.
191
+ You may obtain a copy of the License at
192
+
193
+ http://www.apache.org/licenses/LICENSE-2.0
194
+
195
+ Unless required by applicable law or agreed to in writing, software
196
+ distributed under the License is distributed on an "AS IS" BASIS,
197
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
198
+ See the License for the specific language governing permissions and
199
+ limitations under the License.
README.md CHANGED
@@ -1,3 +1,76 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ library_name: onnx
4
+ pipeline_tag: image-to-image
5
+ base_model: hustvl/Moebius
6
+ tags:
7
+ - image-inpainting
8
+ - inpainting
9
+ - diffusion
10
+ - onnx
11
+ - onnxruntime-web
12
+ - webgpu
13
+ - in-browser
14
+ language:
15
+ - en
16
  ---
17
+
18
+ # Moebius — ONNX (browser / WebGPU)
19
+
20
+ ONNX exports of the [**Moebius**](https://huggingface.co/hustvl/Moebius) 0.22B lightweight
21
+ image-inpainting model ([hustvl/Moebius](https://github.com/hustvl/Moebius), ECCV'26), packaged
22
+ to run **fully client-side in a web browser** via
23
+ [ONNX Runtime Web](https://onnxruntime.ai/docs/tutorials/web/) on the **WebGPU** backend.
24
+
25
+ No text encoder, no tokenizer — Moebius conditions on a small learned embedding table, so the
26
+ whole pipeline is three graphs plus a short DDIM loop you drive in JS.
27
+
28
+ ## Files
29
+
30
+ | File | Graph | Input → Output | Size (fp32) |
31
+ |------|-------|----------------|-------------|
32
+ | `unet.onnx` | student denoiser (`RemovalModel`: embedding + lambda-DWConv UNet) | `latent (B,9,64,64)`, `timesteps (B,)`, `input_ids (B,10)` → `noise (B,4,64,64)` | ~907 MB |
33
+ | `vae_encoder.onnx` | SD VAE encoder | `image (B,3,512,512)` → `moments (B,8,64,64)` | ~137 MB |
34
+ | `vae_decoder.onnx` | SD VAE decoder | `latent (B,4,64,64)` → `image (B,3,512,512)` | ~198 MB |
35
+
36
+ - Exported at a **static 512×512** resolution (64×64 latent). The model's cross-attention uses a
37
+ relative-position embedding tied to the trained resolution, so spatial size is fixed.
38
+ - The learned-embedding "prompt" conditioning stays inside `unet.onnx` as an `nn.Embedding(20, 3072)`
39
+ gather. For classifier-free guidance: `input_ids` rows `[0..9]` = conditional, `[10..19]` = unconditional.
40
+
41
+ ## Pipeline notes (must match for correct output)
42
+
43
+ - **VAE `scaling_factor = 0.13025`** (this is a custom VAE — *not* the usual SD `0.18215`).
44
+ Encode: `latent = mean(moments[:, :4]) * 0.13025`. Decode: feed `latent / 0.13025`.
45
+ - 9-channel UNet input = `concat([noisy_latent(4), mask(1), masked_image_latent(4)], dim=1)`.
46
+ - Scheduler: DDIM, `beta_start=0.00085`, `beta_end=0.012`, `scaled_linear`, 1000 train steps,
47
+ `clip_sample=false`. 20 steps with `strength≈0.99` ⇒ 19 actual steps.
48
+ - VAE encoder source: [`hustvl/PixelHacker`](https://huggingface.co/hustvl/PixelHacker) `vae/`.
49
+
50
+ A reference TypeScript implementation (DDIM loop, CFG, 9-channel assembly, pre/post-processing)
51
+ that loads these files lives in the accompanying web demo.
52
+
53
+ ## Precision
54
+
55
+ These are **fp32** exports (chosen for guaranteed numeric parity with the reference pipeline).
56
+ Parity vs PyTorch on CPU EP: decoder `max|Δ|≈5.7e-5`, unet `≈3.6e-6`. A full-pipeline check vs the
57
+ Torch reference (identical noise) gives a decoded-image `mean|Δ|≈0.0022`. fp16 halves the download
58
+ but is a quality gamble in the lambda layers and unstable for the VAE — validate before using.
59
+
60
+ ## License & attribution
61
+
62
+ Licensed under **Apache 2.0**, inherited from the upstream
63
+ [hustvl/Moebius](https://huggingface.co/hustvl/Moebius). These artifacts are a format conversion
64
+ (PyTorch → ONNX) of the original weights; all model credit belongs to the original authors.
65
+
66
+ ```bibtex
67
+ @misc{DuanAndXu2026Moebius,
68
+ title = {Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance},
69
+ author = {Kangsheng Duan and Ziyang Xu and Wenyu Liu and Xiaohu Ruan and Xiaoxin Chen and Xinggang Wang},
70
+ year = {2026},
71
+ eprint = {2606.19195},
72
+ archivePrefix = {arXiv},
73
+ primaryClass = {cs.CV},
74
+ url = {https://arxiv.org/abs/2606.19195}
75
+ }
76
+ ```
notes.md ADDED
@@ -0,0 +1,132 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Notes / lab log
2
+
3
+ Running log of what Claude Opus 4.8 in Claude Code figured out. Newest at the bottom of each section.
4
+
5
+ ## Environment
6
+ - macOS (darwin arm64), Apple Silicon. No CUDA → torch CPU/MPS build.
7
+ - System python 3.9.6; using `uv` to manage an isolated env.
8
+ - git-lfs 3.7.1 present. Weights cloned to `/tmp/Moebius/Moebius-weights`
9
+ (pretrained, ft_celebahq, ft_ffhq, ft_places2 — each ~450MB fp32 `.bin`).
10
+ - Code repo at `/tmp/Moebius/Moebius`.
11
+
12
+ ## Code map (what matters for the port)
13
+ - Entry: `infer/infer_moebius.py` → `infer/utils.py:build_pipeline`.
14
+ - Pipeline: `removal/v1_2/pipeline.py` (`RemovalSDXLPipeline_BatchMode`).
15
+ - Model wrapper: `removal/v1_2/removal_model.py` (`RemovalModel` = embedding + diff UNet).
16
+ - Core helpers: `utils_infer.py` (`encode_clean_latents`, `predict_noise`).
17
+ - UNet impl: `model_lib/nets/unet_lambda_prune_lite.py` (+ lambda layers under
18
+ `model_lib/nets/layers/λ/vanillaλ.py`).
19
+ - Config: `config/model_cfg/moebius.yaml`.
20
+
21
+ ## Key findings
22
+ - The CUDA/Triton `fla` dependency is ONLY imported in `model_lib/nets/layers/gla/gla.py`
23
+ (the GLA teacher variant). Moebius's student UNet (lambda-DWConv) does not need it —
24
+ must avoid importing `unet_gla` to keep the graph clean for export.
25
+ - "Prompt" conditioning is a plain `nn.Embedding(20, 3072)`. CFG uses fixed ids:
26
+ cond=[0..9], uncond=[10..19]. So encoder_hidden_states is a constant per branch →
27
+ can be precomputed and baked into the ONNX UNet as a constant, OR passed as input.
28
+ - 9-channel UNet input = cat([noisy_latents(4), resized_mask(1), masked_latents(4)], dim=1).
29
+ - CFG batches uncond+cond into one forward (batch dim ×2), then splits.
30
+ - einsum is used in the λ layers (linear attention). Supported in ONNX; need to check
31
+ ORT-Web WebGPU coverage.
32
+
33
+ ## Phase 1 results (reference inference — DONE)
34
+ - Got the real pipeline running end-to-end on CPU (macOS, torch 2.7.1).
35
+ - Patches needed to load student on CPU/mac:
36
+ - `model_lib/__init__.py`: wrapped teacher `unet_gla` import in try/except (needs `fla`).
37
+ - Don't import `utils_train` (drags in orjson/library); `build_vae` is just
38
+ `AutoencoderKL.from_pretrained(vae_dir)`.
39
+ - **VAE scaling_factor = 0.13025** (NOT the usual SD 0.18215!). Custom VAE.
40
+ block_out_channels = [128,256,512,512] → vae_scale_factor 8. This MUST be hardcoded
41
+ correctly in the JS port or colors/contrast will be wrong.
42
+ - removal_model params = 226.04M confirmed. load_state_dict: all keys matched.
43
+ - Perf: ~8.9 s/step on CPU (×19 steps + CFG ×2 = 38 UNet passes ≈ 2:48 total). WebGPU
44
+ expected far faster. Confirms CPU/WASM is unusable; WebGPU is the whole game.
45
+ - Output saved to reference_out/reference_result.png — plausible inpaint. Mask convention:
46
+ white(255) → 1 → region to inpaint (zeroed in masked_image); black → keep.
47
+ - num_inference_steps=20 with strength=0.99 → DDIM uses 19 steps (drops first).
48
+
49
+ ## Parity strategy
50
+ - Won't try to reproduce torch RNG in JS. For PyTorch↔ONNX parity: dump identical input
51
+ tensors and compare outputs. For the web app: generate noise with a seedable JS RNG;
52
+ diffusion is robust to the particular noise draw, so visual results will be valid even
53
+ if not bit-identical to the torch reference.
54
+ - DDIM `scale_model_input` is identity → skip in TS. Need to reproduce DDIM alphas/betas
55
+ (scaled_linear, beta 0.00085→0.012, 1000 steps) and the DDIM step update in JS.
56
+
57
+ ## Architecture: spatial size is FIXED (important!)
58
+ - Self-attn (attn1): MQSλ with `r=15` → local-context path (Conv3d pos_conv). Spatially
59
+ dynamic, fine at any size.
60
+ - Cross-attn (attn2): MQCλ with NO `r` → global path → `rel_pos_emb` is an
61
+ `nn.Parameter(n*n, m, dim_k, dim_u)` where n = per-block sample_size, m = 10. This is
62
+ TIED to the trained spatial resolution. Different spatial size → wrong/oob indexing.
63
+ - ⇒ Export at STATIC 512×512 image (64×64 latent). Web app resizes user input to 512×512,
64
+ inpaints, resizes result back + pastes. Square only. This is the benchmark resolution.
65
+
66
+ ## ONNX export plan
67
+ - Three graphs, spatial static, batch dynamic where cheap:
68
+ - vae_encoder: (B,3,512,512) → moments (B,8,64,64); JS uses mean=moments[:,:4]*sf.
69
+ - unet (RemovalModel): (B,9,64,64), timesteps(B,), input_ids(B,10) → noise(B,4,64,64).
70
+ Embedding (nn.Embedding 20×3072) stays IN the graph (cheap gather). CFG batches B=2.
71
+ - vae_decoder: (B,4,64,64) → (B,3,512,512).
72
+ - scaling_factor = 0.13025 applied in JS (encode: latent*sf; decode: latent/sf).
73
+
74
+ ## Phase 2 results (ONNX export — DONE)
75
+ - torch.onnx.export (legacy tracer, opset 18) traced all 3 graphs cleanly. No op-coverage
76
+ failures. The einsum/lambda/Conv3d ops all exported.
77
+ - Parity vs PyTorch (CPU EP): decoder 5.7e-5, unet 3.6e-6, encoder mean ch ~2e-2.
78
+ - FULL pipeline parity test (python/onnx_pipeline.py): reimplemented DDIM+CFG+9ch+scaling
79
+ in numpy on the ONNX sessions, vs torch models with identical noise:
80
+ final latents max|Δ| 0.149, decoded image mean|Δ| 0.0022, max 0.090 → visually identical.
81
+ This validates the ENTIRE orchestration I'll port to TS.
82
+ - numpy DDIM vs diffusers DDIMScheduler: step max|Δ| 5e-7, timesteps identical. ✓
83
+
84
+ ## DDIM constants for JS (validated)
85
+ - betas = linspace(sqrt(0.00085), sqrt(0.012), 1000)^2 ; alphas_cumprod = cumprod(1-betas)
86
+ - timesteps(20 steps) = [950,900,...,50,0]; strength 0.99 ⇒ drop first ⇒ [900,...,0] (19).
87
+ - ddim_step (eta=0, clip_sample=False):
88
+ pred_x0 = (sample - sqrt(1-ac_t)*eps) / sqrt(ac_t)
89
+ prev = sqrt(ac_prev)*pred_x0 + sqrt(1-ac_prev)*eps
90
+ ac_prev = alphas_cumprod[prev_t], or final_alpha_cumprod=1.0 when prev_t<0 (last step).
91
+ - noise_offset 0.0357: noise += 0.0357 * randn(B,4,1,1). (optional; small)
92
+
93
+ ## Web pipeline recipe (numpy → TS)
94
+ 1. resize image+mask to 512×512 (mask NEAREST, binarize ≥128).
95
+ 2. img→[-1,1] CHW; masked = img*(1-mask).
96
+ 3. encode img & masked → moments; take mean[:4] * 0.13025.
97
+ 4. mask→64×64 NEAREST, 1ch.
98
+ 5. latents = randn(1,4,64,64) [+ noise_offset].
99
+ 6. loop t in timesteps: nine=cat([latents,mask64,maskedLat]); batch×2; unet;
100
+ cfg = u + g*(c-u); latents = ddim_step.
101
+ 7. decode(latents/0.13025); (x+1)/2; clip; → image.
102
+ 8. paste: out*blur(mask) + (1-blur(mask))*orig.
103
+
104
+ ## Phase 3 (web app) — in progress
105
+ - Vite + TS + onnxruntime-web (1.27.0). Default `onnxruntime-web/webgpu` import resolves
106
+ to the self-contained `ort.webgpu.bundle.min.mjs`.
107
+ - Models served LOCALLY: web/public/models -> ../models symlink, at /models/*.onnx.
108
+ Total ~1.24GB fetched over localhost (no internet). ORT runtime served from
109
+ web/ort-dist at /ort/* via a custom static middleware (see vite.config.ts).
110
+ - BUG FIXED: ORT glue .mjs must NOT be in /public (Vite tries to module-transform it).
111
+ Fix = serve /ort/* as raw static files via configureServer middleware.
112
+ - COOP/COEP headers set (needed for threaded WASM / SharedArrayBuffer).
113
+
114
+ ## WebGPU op coverage (confirmed from ORT source at /tmp/Moebius/onnxruntime)
115
+ - js/web/lib/wasm/jsep/webgpu/op-resolve-rules.ts registers: Einsum ✓, Conv ✓
116
+ (conv.ts has computeConv3DInfo / createConv3DNaiveProgramInfo → Conv3d pos_conv works,
117
+ naive kernel so possibly slow), InstanceNormalization ✓, MatMul/Gemm ✓, Softmax ✓,
118
+ Reduce* ✓, Transpose/Concat/Gather/Pad/Resize/Where ✓.
119
+ - GroupNorm: not registered by that name, BUT torch.onnx exports nn.GroupNorm as a
120
+ Reshape→InstanceNormalization→Reshape→Mul→Add decomposition → covered. (VAE decoder
121
+ CPU-EP parity was 5.7e-5, so the graph is decomposed, not a single GroupNorm op.)
122
+ - ⇒ No expected silent CPU fallback for the heavy ops. Confirm empirically in console.
123
+
124
+ ## Verification without a GPU browser (sandbox can't drive Chrome — user's live Chrome
125
+ ## holds the playwright profile)
126
+ - web/test/fixture/*.bin: dumped inputs + reference final latents from the validated
127
+ numpy/ONNX pipeline, to check the TS port (ddim.ts + 9ch assembly) in Node.
128
+
129
+ ## TODO / unknowns
130
+ - fp16 export to ~450MB UNet for real deployment (VAE fp16 unstable: decoder Cast issue;
131
+ keep VAE fp32). Quality risk in λ layers — validate before shipping fp16.
132
+ - Confirm in-browser: WebGPU selected, no CPU fallback, end-to-end correctness + timing.
unet.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:adfb872fead9f7fa8750ecfb6a7acea836c8f96e0cf1d01d9786dcec63d70f4e
3
+ size 906555486
vae_decoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d90ef0b7f6c8c8b7234459c8b449d70be0033bf1576c842e8b9991baf3934280
3
+ size 198078671
vae_encoder.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b8b81d41e757222a0707665ba9d826703987855e5bed056036b90b988968042f
3
+ size 136757093