hoidhxd's picture
Upload README.md with huggingface_hub
660fdbc verified
|
Raw
History Blame Contribute Delete
14.6 kB
---
base_model: sensenova/SenseNova-U1.5-8B-MoT-Preview
library_name: gguf
tags:
- gguf
- quantized
- text-to-image
- sensenova
- mot
- all-in-one
- comfyui
pipeline_tag: text-to-image
---
# SenseNova-U1.5-8B-MoT-Preview — GGUF Q4_0
**Q4_0 GGUF** build of [SenseNova-U1.5-8B-MoT-Preview](https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview),
for running the model locally in ComfyUI on a consumer GPU.
| | |
| --- | --- |
| File | `SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf` |
| Size | **9,929,596,256 bytes — 9.25 GiB (9.93 GB)** |
| Source | 16-shard BF16 release (~50 GB) |
| Fits | 16 GB VRAM fully resident; 12 GB with layer offload |
> The file must go in `ComfyUI/models/gguf/`, **not** `ComfyUI/models/unet/` —
> see [step 3](#3-put-the-gguf-where-the-node-actually-looks).
## Architecture: NEO-unify
SenseNova U1 is a native multimodal model — one graph handles text and pixels
end to end.
* 🚫 **No external text encoder** (no CLIP, no T5).
* 🚫 **No external VAE.**
So the ComfyUI graph is just two nodes: a loader and a sampler. There is nothing
else to wire up.
## Quantization details
Converted directly from the BF16 safetensors with smart filtering:
* **Q4_0** — the large 2D weight tensors (Linear, Conv).
* **FP32** — 1D tensors (bias, LayerNorm) and anything ≤ 1024 parameters, to
keep structural accuracy.
* **FP16 fallback** — 6 unaligned tensors, so the Mixture-of-Transformers
routing stays intact.
Five tensors that *should* have been excluded were not; see
[Known issue](#known-issue-five-tensors-must-stay-dense) below. The workaround
is a 30-line file and takes one minute to install.
---
# Using this model in ComfyUI
Tested on Windows 11 + RTX 5060 Ti 16 GB, ComfyUI with a Python 3.13 venv.
Linux is the same apart from paths.
Throughout, `<ComfyUI>` is your ComfyUI root (e.g. `D:\ComfyUI`) and
`<python>` is the interpreter **ComfyUI itself runs on** — not your system
Python. For a portable build that is `<ComfyUI>\..\python_embeded\python.exe`;
for a venv install, `<ComfyUI>\venv\Scripts\python.exe` (Windows) or
`<ComfyUI>/venv/bin/python` (Linux).
## 1. Install the custom nodes
Install **ComfyUI-SenseNova-U1** through ComfyUI Manager, or clone it:
```bash
git clone https://github.com/OpenSenseNova/ComfyUI-SenseNova-U1 <ComfyUI>/custom_nodes/ComfyUI-SenseNova-U1
```
## 2. Install the runtime and the GGUF extra
The nodes need the `sensenova-u1` runtime package plus the GGUF dependencies.
Install both into ComfyUI's Python:
```bash
<python> -m pip install -r <ComfyUI>/custom_nodes/ComfyUI-SenseNova-U1/requirements.txt
```
```bash
<python> -m pip install "gguf>=0.10.0" "diffusers>=0.30.0" accelerate transformers
```
`requirements.txt` pulls `sensenova-u1` from a GitHub release tarball, which
is intentional — a `git+https` install would drag in hundreds of MB of
evaluation submodules.
## 3. Put the GGUF where the node actually looks
**`ComfyUI/models/unet/` does not work.** The `SenseNova U1 Local Loader`
scans exactly two folder names, `gguf` and `diffusion_models`, and
`diffusion_models` filters on ComfyUI's `supported_pt_extensions`, which does
not include `.gguf`. Anything in `unet/` is invisible to the node and the
dropdown comes up empty.
Download into `<ComfyUI>/models/gguf/`:
```bash
hf download hoidhxd/SenseNova-U1.5-8B-GGUF SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf --local-dir <ComfyUI>/models/gguf
```
If you already have the file elsewhere (a different drive, say) don't copy
9.25 GiB around — register the directory in `<ComfyUI>/extra_model_paths.yaml`
**under the key `gguf`**:
```yaml
ai_models:
base_path: C:/Users/Admin/ai/models
gguf: SenseNova-U1.5-8B-GGUF
```
The key name is what matters. ComfyUI gives an unrecognised folder name an
empty extension set, and an empty set means "no filter" — which is why `.gguf`
files surface under `gguf` but not under `diffusion_models`.
Restart ComfyUI after adding files; the dropdown is built at startup.
## 4. Get the config and tokenizer
The GGUF holds weights only. The loader still needs the config and tokenizer
from the base repo — **but not the 50 GB of safetensors**:
```bash
hf download sensenova/SenseNova-U1.5-8B-MoT-Preview --local-dir <somewhere>/SenseNova-U1.5-8B-MoT-Preview --include "*.json" "*.txt"
```
That yields ~5 MB:
```
config.json added_tokens.json special_tokens_map.json
tokenizer_config.json vocab.json merges.txt
model.safetensors.index.json
```
This directory is what you type into the loader's `model_path`.
## 5. Install the compatibility shim (required)
Five tensors in this GGUF crash the loader as published. Create
`<ComfyUI>/custom_nodes/sensenova_u1_embed_fix/__init__.py` with the file below
and restart ComfyUI. It registers no nodes — it patches the GGUF loader at
import time and dequantizes those five tensors back to bf16 (~1.3 GB extra
weight memory).
```python
"""Hand a few SenseNova-U1 tensors to the GGUF loader as plain floats.
Two tensor groups in `hoidhxd/SenseNova-U1.5-8B-GGUF` cannot survive the
diffusers GGUF quantizer, for two different reasons.
1. ``language_model.model.embed_tokens.weight`` (Q4_0)
The quantizer only swaps ``nn.Linear`` for ``GGUFLinear``; every other module
keeps whatever parameter it was handed. An ``nn.Embedding`` therefore looks up
raw Q4_0 *block bytes*, and the first RMSNorm dies with
RuntimeError: The size of tensor a (4096) must match the size of
tensor b (2304) at non-singleton dimension 2
2304 is exactly ``4096 // 32 * 18`` - the Q4_0 byte width of a 4096-wide row.
Costs ~1.2 GB of extra weight memory in bfloat16.
2. ``fm_modules.timestep_embedder`` / ``fm_modules.noise_scale_embedder`` (Q4_0)
``modeling_fm_modules.TimestepEmbedder.forward`` casts its input with
``t_freq.to(self.mlp[0].weight.dtype)``. On a ``GGUFLinear`` that dtype is the
*storage* dtype, ``torch.uint8``, so the activations are cast to Byte and the
matmul dies with
RuntimeError: mat1 and mat2 must have the same dtype, but got Byte and BFloat16
Only 35.7M params live here (~71 MB in bfloat16), so keeping them dense is
basically free.
The proper fix is to leave both groups in F16/F32 when producing the GGUF; this
shim exists so an already-published checkpoint still runs.
Set SENSENOVA_GGUF_EMBED_KEYS to a comma-separated list of name fragments to
override which tensors get dequantized.
"""
import logging
import os
LOGGER = logging.getLogger(__name__)
DEFAULT_KEYS = (
"embed_tokens.weight",
"fm_modules.timestep_embedder.",
"fm_modules.noise_scale_embedder.",
)
def _embedding_keys() -> tuple[str, ...]:
raw = os.environ.get("SENSENOVA_GGUF_EMBED_KEYS", "").strip()
if not raw:
return DEFAULT_KEYS
return tuple(part.strip() for part in raw.split(",") if part.strip())
def _apply_patch() -> None:
from sensenova_u1.utils import gguf_loader
if getattr(gguf_loader, "_embed_fix_applied", False):
return
original = gguf_loader.load_gguf_checkpoint
keys = _embedding_keys()
def load_gguf_checkpoint(path: str, *args, **kwargs) -> dict:
import torch
from diffusers.quantizers.gguf.utils import GGUFParameter, dequantize_gguf_tensor
state_dict = original(path, *args, **kwargs)
for name, tensor in list(state_dict.items()):
if not isinstance(tensor, GGUFParameter) or not any(k in name for k in keys):
continue
dequantized = dequantize_gguf_tensor(tensor).as_subclass(torch.Tensor)
state_dict[name] = dequantized
LOGGER.info(
"SenseNova GGUF embed fix: dequantized %s %s -> %s",
name,
tuple(tensor.shape),
tuple(dequantized.shape),
)
return state_dict
gguf_loader.load_gguf_checkpoint = load_gguf_checkpoint
gguf_loader._embed_fix_applied = True
LOGGER.info("SenseNova GGUF embed fix installed for keys: %s", ", ".join(keys))
try:
_apply_patch()
except Exception as exc: # noqa: BLE001 - never block ComfyUI startup
LOGGER.warning("SenseNova GGUF embed fix not installed: %s", exc)
NODE_CLASS_MAPPINGS: dict = {}
NODE_DISPLAY_NAME_MAPPINGS: dict = {}
```
On startup the ComfyUI console should print:
```
SenseNova GGUF embed fix installed for keys: embed_tokens.weight, fm_modules.timestep_embedder., fm_modules.noise_scale_embedder.
```
If it prints `not installed: No module named 'sensenova_u1'` instead, step 2
did not land in the interpreter ComfyUI is actually using.
## 6. Build the workflow
Two nodes, one link:
```
[SenseNova U1 Local Loader] --u1_model--> [SenseNova U1 Local Text to Image] --images--> [Save Image]
```
**SenseNova U1 Local Loader**
| Input | Value |
| --- | --- |
| `model_path` | the config/tokenizer directory from step 4 |
| `sensenova_u1_src` | leave as-is (auto-resolved) |
| `device` | `cuda` |
| `dtype` | `bfloat16` |
| `attn_backend` | `auto` |
| `device_map` | `none`**must** be `none` when a GGUF is selected |
| `max_memory` | empty |
| `vram_mode` | `full` on 16 GB, `balanced` on 12 GB |
| `gguf_checkpoint` | `SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf` |
`vram_mode` replaced the old `prefetch_count` input:
* `full` — every weight stays on the GPU. Fastest, ~2× the offload modes.
* `balanced` — asynchronous layer prefetch, overlaps host→device copies with
compute. Use this on 12 GB.
* `low` — synchronous one-layer-at-a-time swap. Smallest footprint, slowest.
`device_map` is for splitting across *multiple* GPUs and is mutually exclusive
with `vram_mode`; leave it `none` for single-GPU use.
**SenseNova U1 Local Text to Image**
| Input | Default | Notes |
| --- | --- | --- |
| `prompt` | — | plain text, no encoder node |
| `resolution` | `2048x2048\|1:1` | native sizes only, see below |
| `cfg_scale` | `4.0` | |
| `cfg_norm` | `none` | `global` / `channel` / `cfg_zero_star` |
| `timestep_shift` | `3.0` | sampler schedule shift |
| `cfg_interval_start` / `_end` | `0.0` / `1.0` | window where CFG applies |
| `num_steps` | `50` | 16 is fine for drafts |
| `batch_size` | `1` | |
| `seed` | — | |
| `think_mode` | `false` | model reasons before drawing; text on the `think_text` output |
U1.5 samples **only at its own native resolutions**. Pick the aspect ratio you
want and downscale afterwards if you need a specific pixel size:
| Ratio | Pixels | Ratio | Pixels |
| --- | --- | --- | --- |
| 1:1 | 2048×2048 | 2:1 | 2880×1440 |
| 16:9 | 2720×1536 | 1:2 | 1440×2880 |
| 9:16 | 1536×2720 | 3:1 | 3456×1152 |
| 3:2 | 2496×1664 | 1:3 | 1152×3456 |
| 2:3 | 1664×2496 | 4:3 | 2368×1760 |
| 3:4 | 1760×2368 | | |
Example prompt:
> A cinematic, dynamic shot of a terrified old man frantically running away from
> a massive, shadowy monster in a dark, foggy forest, high contrast, 8k
> resolution, photorealistic.
Also available: **SenseNova U1 Local Image Edit** (image + instruction) and
**SenseNova U1 Local Interleave** (alternating text and images). Ready-made
graphs ship in the node's `example_workflows/` folder.
## 7. VRAM and timing
Measured on an RTX 5060 Ti 16 GB with this Q4_0 file:
| Run | Steps | Size | Wall time |
| --- | --- | --- | --- |
| t2i, `full`, includes loading the 9.25 GiB file | 4 | 2048×2048 | 106 s |
| t2i, `balanced` (layer offload) | 20 | 2720×1536 | 200 s |
| t2i, `full`, `batch_size=2` | 8 | 2048×2048 | 106 s |
| edit, `balanced`, 2.1 MP | 8 | 1440×1440 | 119 s |
At `vram_mode=full` the weights sit at **10.2 GiB** resident and peak around
**10.9 GiB** while sampling 2048². `batch_size=2` at 2048² peaks at 14.15 GiB,
about as far as a 16 GB card goes — go `balanced` beyond that.
**Image editing needs more room than generation.** The edit node runs the source
image and the generated one through the model together; at `full` with the
node's stock 4.19 MP target it OOMs on 16 GB (11.81 GiB weights plus a 2.27 GiB
allocation). Use `vram_mode=balanced` and lower the megapixel target to ~2.1 for
editing.
Every run above includes a model reload, because each changed something in the
loader's cache key. **Changing `vram_mode`, `model_path`, `dtype`, `device_map`
or the GGUF selection forces a full reload** — keep them stable between
generations and only the first run pays the load cost.
---
# Known issue: five tensors must stay dense
Without the shim from step 5, this checkpoint fails at load or on the first
step. Both failures come from the same place: diffusers' GGUF quantizer only
swaps `nn.Linear` for `GGUFLinear`, so every other module keeps the raw
quantized bytes.
**1. The token embedding.** `language_model.model.embed_tokens.weight` is an
`nn.Embedding`, so the lookup returns Q4_0 *block bytes* — `4096 // 32 * 18` =
**2304** wide instead of 4096 — and the first RMSNorm fails:
```
RuntimeError: The size of tensor a (4096) must match the size of tensor b (2304)
at non-singleton dimension 2
```
**2. The two embedders.** `fm_modules.timestep_embedder` and
`fm_modules.noise_scale_embedder` *are* `nn.Linear`, but
`modeling_fm_modules.py` casts activations with
`t_freq.to(self.mlp[0].weight.dtype)`. On a `GGUFLinear` that dtype is the
*storage* dtype `torch.uint8`, so the activations become Byte:
```
RuntimeError: mat1 and mat2 must have the same dtype, but got Byte and BFloat16
```
| Symptom | Cause | Fix |
| --- | --- | --- |
| `gguf_checkpoint` dropdown is empty | file is in `models/unet/` | move it to `models/gguf/` (step 3), restart |
| `tensor a (4096) ... tensor b (2304)` | quantized embedding | install the shim (step 5) |
| `got Byte and BFloat16` | quantized timestep embedder | install the shim (step 5) |
| `not installed: No module named 'sensenova_u1'` | deps in the wrong Python | reinstall with ComfyUI's interpreter (step 2) |
| OOM while editing | `vram_mode=full` + 4.19 MP | `balanced`, ~2.1 MP |
**The proper fix is upstream, in the quantization step:** keep those five
tensors in F16/F32 when producing the GGUF. The embedding costs ~1.2 GB and the
two embedders only 35.7M params (~71 MB), so a re-quantized upload would need
no shim and would work with the stock node. That is planned for the next
revision of this repo.
# Download
```bash
hf download hoidhxd/SenseNova-U1.5-8B-GGUF SenseNova-U1.5-8B-MoT-Preview-Q4_0.gguf --local-dir .
```
# License
Inherits the license of the base model,
[sensenova/SenseNova-U1.5-8B-MoT-Preview](https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview).