docs: honest model card (validation status + fleet-specific perf caveat)
Browse files
README.md
CHANGED
|
@@ -13,18 +13,33 @@ tags:
|
|
| 13 |
- quantized
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# Qwen3.5-0.8B — fni8 (int8 dp4a)
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
|
| 22 |
-
- **
|
| 23 |
-
(loads with no dequant/repack).
|
| 24 |
-
- **Why dp4a:** sm_70 has no int8 tensor cores; the contraction runs on the
|
| 25 |
-
`__dp4a` CUDA-core intrinsic. On the CMP-100-210 fleet (firmware-gimped fp16
|
| 26 |
-
tensor cores) dp4a is the fast path, not a compromise.
|
| 27 |
-
- **Runtimes:** [fni8-serve](https://github.com/jajmangold/fni8-serve) (LLMs) /
|
| 28 |
-
[ComfyUI-fni8](https://github.com/jajmangold/ComfyUI-fni8) (diffusion DiTs).
|
| 29 |
|
| 30 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
- quantized
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# Qwen3.5-0.8B — fni8 (int8/W4A8 dp4a, Volta sm_70)
|
| 17 |
|
| 18 |
+
> ⚠️ **EXPERIMENTAL — do not rely on this checkpoint.**
|
| 19 |
+
>
|
| 20 |
+
> It was converted **before the hybrid-config-drop converter fix**: the hybrid (Gated-DeltaNet) portion of the config was dropped during conversion, so the file may not load or run correctly. A **corrected reconversion is pending**. Even once reconverted, full decode additionally requires the Track-2 DeltaNet int8 kernel (currently a fp16 linear-attention fallback).
|
| 21 |
|
| 22 |
+
Qwen3.5/3.6-family hybrid (Gated-DeltaNet) model, quantized from [`Qwen/Qwen3.5-0.8B`](https://huggingface.co/Qwen/Qwen3.5-0.8B). Repackaged to the **`.fni8`** resident format (~1.4 GB) for the [fni8](https://github.com/jajmangold/fni8) **W8A8/W4A8 DP4A** kernels on **NVIDIA Volta (sm_70)** — Tesla V100 / CMP 100-210.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
|
| 24 |
+
## Status
|
| 25 |
+
|
| 26 |
+
See the experimental warning above. Not validated end-to-end.
|
| 27 |
+
|
| 28 |
+
## Format
|
| 29 |
+
|
| 30 |
+
- **Weights:** int8 per-row (W8A8) or int4 per-group (W4A8), fp32 scales, resident dp4a VRAM layout.
|
| 31 |
+
- **Why dp4a:** sm_70 has no int8 tensor cores, so the matmul contraction runs on the `__dp4a` CUDA-core intrinsic. On the CMP 100-210 fleet (whose fp16 tensor cores are firmware-limited) dp4a is the fast path, not a compromise.
|
| 32 |
+
|
| 33 |
+
## How to run
|
| 34 |
+
|
| 35 |
+
[fni8-serve](https://github.com/jajmangold/fni8-serve) is the LLM runtime (`load_fni8_state_dict(<file>)` into an `LLMEngine`; the architecture is read from the file). ComfyUI-fni8 is for diffusion DiTs only and does not load this model.
|
| 36 |
+
|
| 37 |
+
## Limitations
|
| 38 |
+
|
| 39 |
+
- Quantization is lossy: int8 (and especially int4) outputs differ from the fp16/bf16 parent, and the difference varies by task.
|
| 40 |
+
- Capabilities, biases, and risks of the parent model carry over — see the parent card.
|
| 41 |
+
- This is a derivative quantization, not a relicense; the parent model's license and acceptable uses apply.
|
| 42 |
+
|
| 43 |
+
---
|
| 44 |
+
|
| 45 |
+
Part of the fni8 stack: [kernels](https://github.com/jajmangold/fni8) · [LLM serving](https://github.com/jajmangold/fni8-serve) · [ComfyUI DiTs](https://github.com/jajmangold/ComfyUI-fni8).
|