--- base_model: sensenova/SenseNova-U1.5-8B-MoT-Preview library_name: gguf tags: - gguf - quantized - text-to-image - sensenova - mot - all-in-one - comfyui pipeline_tag: text-to-image --- # SenseNova-U1.5-8B-MoT-Preview β€” GGUF Q4_0 **Q4_0 GGUF** build of [SenseNova-U1.5-8B-MoT-Preview](https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview), for running the model locally in ComfyUI on a consumer GPU. | | | | --- | --- | | File | `SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf` | | Size | **9,929,596,256 bytes β€” 9.25 GiB (9.93 GB)** | | Source | 16-shard BF16 release (~50 GB) | | Fits | 16 GB VRAM fully resident; 12 GB with layer offload | > The file must go in `ComfyUI/models/gguf/`, **not** `ComfyUI/models/unet/` β€” > see [step 3](#3-put-the-gguf-where-the-node-actually-looks). ## Architecture: NEO-unify SenseNova U1 is a native multimodal model β€” one graph handles text and pixels end to end. * 🚫 **No external text encoder** (no CLIP, no T5). * 🚫 **No external VAE.** So the ComfyUI graph is just two nodes: a loader and a sampler. There is nothing else to wire up. ## Quantization details Converted directly from the BF16 safetensors with smart filtering: * **Q4_0** β€” the large 2D weight tensors (Linear, Conv). * **FP32** β€” 1D tensors (bias, LayerNorm) and anything ≀ 1024 parameters, to keep structural accuracy. * **FP16 fallback** β€” 6 unaligned tensors, so the Mixture-of-Transformers routing stays intact. Five tensors that *should* have been excluded were not; see [Known issue](#known-issue-five-tensors-must-stay-dense) below. The workaround is a 30-line file and takes one minute to install. --- # Using this model in ComfyUI Tested on Windows 11 + RTX 5060 Ti 16 GB, ComfyUI with a Python 3.13 venv. Linux is the same apart from paths. Throughout, `` is your ComfyUI root (e.g. `D:\ComfyUI`) and `` is the interpreter **ComfyUI itself runs on** β€” not your system Python. For a portable build that is `\..\python_embeded\python.exe`; for a venv install, `\venv\Scripts\python.exe` (Windows) or `/venv/bin/python` (Linux). ## 1. Install the custom nodes Install **ComfyUI-SenseNova-U1** through ComfyUI Manager, or clone it: ```bash git clone https://github.com/OpenSenseNova/ComfyUI-SenseNova-U1 /custom_nodes/ComfyUI-SenseNova-U1 ``` ## 2. Install the runtime and the GGUF extra The nodes need the `sensenova-u1` runtime package plus the GGUF dependencies. Install both into ComfyUI's Python: ```bash -m pip install -r /custom_nodes/ComfyUI-SenseNova-U1/requirements.txt ``` ```bash -m pip install "gguf>=0.10.0" "diffusers>=0.30.0" accelerate transformers ``` `requirements.txt` pulls `sensenova-u1` from a GitHub release tarball, which is intentional β€” a `git+https` install would drag in hundreds of MB of evaluation submodules. ## 3. Put the GGUF where the node actually looks **`ComfyUI/models/unet/` does not work.** The `SenseNova U1 Local Loader` scans exactly two folder names, `gguf` and `diffusion_models`, and `diffusion_models` filters on ComfyUI's `supported_pt_extensions`, which does not include `.gguf`. Anything in `unet/` is invisible to the node and the dropdown comes up empty. Download into `/models/gguf/`: ```bash hf download hoidhxd/SenseNova-U1.5-8B-GGUF SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf --local-dir /models/gguf ``` If you already have the file elsewhere (a different drive, say) don't copy 9.25 GiB around β€” register the directory in `/extra_model_paths.yaml` **under the key `gguf`**: ```yaml ai_models: base_path: C:/Users/Admin/ai/models gguf: SenseNova-U1.5-8B-GGUF ``` The key name is what matters. ComfyUI gives an unrecognised folder name an empty extension set, and an empty set means "no filter" β€” which is why `.gguf` files surface under `gguf` but not under `diffusion_models`. Restart ComfyUI after adding files; the dropdown is built at startup. ## 4. Get the config and tokenizer The GGUF holds weights only. The loader still needs the config and tokenizer from the base repo β€” **but not the 50 GB of safetensors**: ```bash hf download sensenova/SenseNova-U1.5-8B-MoT-Preview --local-dir /SenseNova-U1.5-8B-MoT-Preview --include "*.json" "*.txt" ``` That yields ~5 MB: ``` config.json added_tokens.json special_tokens_map.json tokenizer_config.json vocab.json merges.txt model.safetensors.index.json ``` This directory is what you type into the loader's `model_path`. ## 5. Install the compatibility shim (required) Five tensors in this GGUF crash the loader as published. Create `/custom_nodes/sensenova_u1_embed_fix/__init__.py` with the file below and restart ComfyUI. It registers no nodes β€” it patches the GGUF loader at import time and dequantizes those five tensors back to bf16 (~1.3 GB extra weight memory). ```python """Hand a few SenseNova-U1 tensors to the GGUF loader as plain floats. Two tensor groups in `hoidhxd/SenseNova-U1.5-8B-GGUF` cannot survive the diffusers GGUF quantizer, for two different reasons. 1. ``language_model.model.embed_tokens.weight`` (Q4_0) The quantizer only swaps ``nn.Linear`` for ``GGUFLinear``; every other module keeps whatever parameter it was handed. An ``nn.Embedding`` therefore looks up raw Q4_0 *block bytes*, and the first RMSNorm dies with RuntimeError: The size of tensor a (4096) must match the size of tensor b (2304) at non-singleton dimension 2 2304 is exactly ``4096 // 32 * 18`` - the Q4_0 byte width of a 4096-wide row. Costs ~1.2 GB of extra weight memory in bfloat16. 2. ``fm_modules.timestep_embedder`` / ``fm_modules.noise_scale_embedder`` (Q4_0) ``modeling_fm_modules.TimestepEmbedder.forward`` casts its input with ``t_freq.to(self.mlp[0].weight.dtype)``. On a ``GGUFLinear`` that dtype is the *storage* dtype, ``torch.uint8``, so the activations are cast to Byte and the matmul dies with RuntimeError: mat1 and mat2 must have the same dtype, but got Byte and BFloat16 Only 35.7M params live here (~71 MB in bfloat16), so keeping them dense is basically free. The proper fix is to leave both groups in F16/F32 when producing the GGUF; this shim exists so an already-published checkpoint still runs. Set SENSENOVA_GGUF_EMBED_KEYS to a comma-separated list of name fragments to override which tensors get dequantized. """ import logging import os LOGGER = logging.getLogger(__name__) DEFAULT_KEYS = ( "embed_tokens.weight", "fm_modules.timestep_embedder.", "fm_modules.noise_scale_embedder.", ) def _embedding_keys() -> tuple[str, ...]: raw = os.environ.get("SENSENOVA_GGUF_EMBED_KEYS", "").strip() if not raw: return DEFAULT_KEYS return tuple(part.strip() for part in raw.split(",") if part.strip()) def _apply_patch() -> None: from sensenova_u1.utils import gguf_loader if getattr(gguf_loader, "_embed_fix_applied", False): return original = gguf_loader.load_gguf_checkpoint keys = _embedding_keys() def load_gguf_checkpoint(path: str, *args, **kwargs) -> dict: import torch from diffusers.quantizers.gguf.utils import GGUFParameter, dequantize_gguf_tensor state_dict = original(path, *args, **kwargs) for name, tensor in list(state_dict.items()): if not isinstance(tensor, GGUFParameter) or not any(k in name for k in keys): continue dequantized = dequantize_gguf_tensor(tensor).as_subclass(torch.Tensor) state_dict[name] = dequantized LOGGER.info( "SenseNova GGUF embed fix: dequantized %s %s -> %s", name, tuple(tensor.shape), tuple(dequantized.shape), ) return state_dict gguf_loader.load_gguf_checkpoint = load_gguf_checkpoint gguf_loader._embed_fix_applied = True LOGGER.info("SenseNova GGUF embed fix installed for keys: %s", ", ".join(keys)) try: _apply_patch() except Exception as exc: # noqa: BLE001 - never block ComfyUI startup LOGGER.warning("SenseNova GGUF embed fix not installed: %s", exc) NODE_CLASS_MAPPINGS: dict = {} NODE_DISPLAY_NAME_MAPPINGS: dict = {} ``` On startup the ComfyUI console should print: ``` SenseNova GGUF embed fix installed for keys: embed_tokens.weight, fm_modules.timestep_embedder., fm_modules.noise_scale_embedder. ``` If it prints `not installed: No module named 'sensenova_u1'` instead, step 2 did not land in the interpreter ComfyUI is actually using. ## 6. Build the workflow Two nodes, one link: ``` [SenseNova U1 Local Loader] --u1_model--> [SenseNova U1 Local Text to Image] --images--> [Save Image] ``` **SenseNova U1 Local Loader** | Input | Value | | --- | --- | | `model_path` | the config/tokenizer directory from step 4 | | `sensenova_u1_src` | leave as-is (auto-resolved) | | `device` | `cuda` | | `dtype` | `bfloat16` | | `attn_backend` | `auto` | | `device_map` | `none` β€” **must** be `none` when a GGUF is selected | | `max_memory` | empty | | `vram_mode` | `full` on 16 GB, `balanced` on 12 GB | | `gguf_checkpoint` | `SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf` | `vram_mode` replaced the old `prefetch_count` input: * `full` β€” every weight stays on the GPU. Fastest, ~2Γ— the offload modes. * `balanced` β€” asynchronous layer prefetch, overlaps hostβ†’device copies with compute. Use this on 12 GB. * `low` β€” synchronous one-layer-at-a-time swap. Smallest footprint, slowest. `device_map` is for splitting across *multiple* GPUs and is mutually exclusive with `vram_mode`; leave it `none` for single-GPU use. **SenseNova U1 Local Text to Image** | Input | Default | Notes | | --- | --- | --- | | `prompt` | β€” | plain text, no encoder node | | `resolution` | `2048x2048\|1:1` | native sizes only, see below | | `cfg_scale` | `4.0` | | | `cfg_norm` | `none` | `global` / `channel` / `cfg_zero_star` | | `timestep_shift` | `3.0` | sampler schedule shift | | `cfg_interval_start` / `_end` | `0.0` / `1.0` | window where CFG applies | | `num_steps` | `50` | 16 is fine for drafts | | `batch_size` | `1` | | | `seed` | β€” | | | `think_mode` | `false` | model reasons before drawing; text on the `think_text` output | U1.5 samples **only at its own native resolutions**. Pick the aspect ratio you want and downscale afterwards if you need a specific pixel size: | Ratio | Pixels | Ratio | Pixels | | --- | --- | --- | --- | | 1:1 | 2048Γ—2048 | 2:1 | 2880Γ—1440 | | 16:9 | 2720Γ—1536 | 1:2 | 1440Γ—2880 | | 9:16 | 1536Γ—2720 | 3:1 | 3456Γ—1152 | | 3:2 | 2496Γ—1664 | 1:3 | 1152Γ—3456 | | 2:3 | 1664Γ—2496 | 4:3 | 2368Γ—1760 | | 3:4 | 1760Γ—2368 | | | Example prompt: > A cinematic, dynamic shot of a terrified old man frantically running away from > a massive, shadowy monster in a dark, foggy forest, high contrast, 8k > resolution, photorealistic. Also available: **SenseNova U1 Local Image Edit** (image + instruction) and **SenseNova U1 Local Interleave** (alternating text and images). Ready-made graphs ship in the node's `example_workflows/` folder. ## 7. VRAM and timing Measured on an RTX 5060 Ti 16 GB with this Q4_0 file: | Run | Steps | Size | Wall time | | --- | --- | --- | --- | | t2i, `full`, includes loading the 9.25 GiB file | 4 | 2048Γ—2048 | 106 s | | t2i, `balanced` (layer offload) | 20 | 2720Γ—1536 | 200 s | | t2i, `full`, `batch_size=2` | 8 | 2048Γ—2048 | 106 s | | edit, `balanced`, 2.1 MP | 8 | 1440Γ—1440 | 119 s | At `vram_mode=full` the weights sit at **10.2 GiB** resident and peak around **10.9 GiB** while sampling 2048Β². `batch_size=2` at 2048Β² peaks at 14.15 GiB, about as far as a 16 GB card goes β€” go `balanced` beyond that. **Image editing needs more room than generation.** The edit node runs the source image and the generated one through the model together; at `full` with the node's stock 4.19 MP target it OOMs on 16 GB (11.81 GiB weights plus a 2.27 GiB allocation). Use `vram_mode=balanced` and lower the megapixel target to ~2.1 for editing. Every run above includes a model reload, because each changed something in the loader's cache key. **Changing `vram_mode`, `model_path`, `dtype`, `device_map` or the GGUF selection forces a full reload** β€” keep them stable between generations and only the first run pays the load cost. --- # Known issue: five tensors must stay dense Without the shim from step 5, this checkpoint fails at load or on the first step. Both failures come from the same place: diffusers' GGUF quantizer only swaps `nn.Linear` for `GGUFLinear`, so every other module keeps the raw quantized bytes. **1. The token embedding.** `language_model.model.embed_tokens.weight` is an `nn.Embedding`, so the lookup returns Q4_0 *block bytes* β€” `4096 // 32 * 18` = **2304** wide instead of 4096 β€” and the first RMSNorm fails: ``` RuntimeError: The size of tensor a (4096) must match the size of tensor b (2304) at non-singleton dimension 2 ``` **2. The two embedders.** `fm_modules.timestep_embedder` and `fm_modules.noise_scale_embedder` *are* `nn.Linear`, but `modeling_fm_modules.py` casts activations with `t_freq.to(self.mlp[0].weight.dtype)`. On a `GGUFLinear` that dtype is the *storage* dtype `torch.uint8`, so the activations become Byte: ``` RuntimeError: mat1 and mat2 must have the same dtype, but got Byte and BFloat16 ``` | Symptom | Cause | Fix | | --- | --- | --- | | `gguf_checkpoint` dropdown is empty | file is in `models/unet/` | move it to `models/gguf/` (step 3), restart | | `tensor a (4096) ... tensor b (2304)` | quantized embedding | install the shim (step 5) | | `got Byte and BFloat16` | quantized timestep embedder | install the shim (step 5) | | `not installed: No module named 'sensenova_u1'` | deps in the wrong Python | reinstall with ComfyUI's interpreter (step 2) | | OOM while editing | `vram_mode=full` + 4.19 MP | `balanced`, ~2.1 MP | **The proper fix is upstream, in the quantization step:** keep those five tensors in F16/F32 when producing the GGUF. The embedding costs ~1.2 GB and the two embedders only 35.7M params (~71 MB), so a re-quantized upload would need no shim and would work with the stock node. That is planned for the next revision of this repo. # Download ```bash hf download hoidhxd/SenseNova-U1.5-8B-GGUF SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf --local-dir . ``` # License Inherits the license of the base model, [sensenova/SenseNova-U1.5-8B-MoT-Preview](https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview).