jajmangold commited on
Commit
e6d00ef
·
verified ·
1 Parent(s): cb93aed

docs: honest model card (validation status + fleet-specific perf caveat)

Browse files
Files changed (1) hide show
  1. README.md +27 -12
README.md CHANGED
@@ -13,18 +13,33 @@ tags:
13
  - quantized
14
  ---
15
 
16
- # Qwen3.5-0.8B — fni8 (int8 dp4a)
17
 
18
- Quantization of [`Qwen/Qwen3.5-0.8B`](https://huggingface.co/Qwen/Qwen3.5-0.8B) to the
19
- **`.fni8`** resident format (~1.4 GB) for the [fni8](https://github.com/jajmangold/fni8)
20
- W8A8 DP4A kernels on **NVIDIA Volta (sm_70)** — Tesla V100 / CMP 100-210.
21
 
22
- - **Weights:** int8 per-row (W8A8), fp32 scales, stored in the resident dp4a VRAM layout
23
- (loads with no dequant/repack).
24
- - **Why dp4a:** sm_70 has no int8 tensor cores; the contraction runs on the
25
- `__dp4a` CUDA-core intrinsic. On the CMP-100-210 fleet (firmware-gimped fp16
26
- tensor cores) dp4a is the fast path, not a compromise.
27
- - **Runtimes:** [fni8-serve](https://github.com/jajmangold/fni8-serve) (LLMs) /
28
- [ComfyUI-fni8](https://github.com/jajmangold/ComfyUI-fni8) (diffusion DiTs).
29
 
30
- This is a derivative quantization; its license follows the parent model above.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
  - quantized
14
  ---
15
 
16
+ # Qwen3.5-0.8B — fni8 (int8/W4A8 dp4a, Volta sm_70)
17
 
18
+ > ⚠️ **EXPERIMENTAL — do not rely on this checkpoint.**
19
+ >
20
+ > It was converted **before the hybrid-config-drop converter fix**: the hybrid (Gated-DeltaNet) portion of the config was dropped during conversion, so the file may not load or run correctly. A **corrected reconversion is pending**. Even once reconverted, full decode additionally requires the Track-2 DeltaNet int8 kernel (currently a fp16 linear-attention fallback).
21
 
22
+ Qwen3.5/3.6-family hybrid (Gated-DeltaNet) model, quantized from [`Qwen/Qwen3.5-0.8B`](https://huggingface.co/Qwen/Qwen3.5-0.8B). Repackaged to the **`.fni8`** resident format (~1.4 GB) for the [fni8](https://github.com/jajmangold/fni8) **W8A8/W4A8 DP4A** kernels on **NVIDIA Volta (sm_70)** — Tesla V100 / CMP 100-210.
 
 
 
 
 
 
23
 
24
+ ## Status
25
+
26
+ See the experimental warning above. Not validated end-to-end.
27
+
28
+ ## Format
29
+
30
+ - **Weights:** int8 per-row (W8A8) or int4 per-group (W4A8), fp32 scales, resident dp4a VRAM layout.
31
+ - **Why dp4a:** sm_70 has no int8 tensor cores, so the matmul contraction runs on the `__dp4a` CUDA-core intrinsic. On the CMP 100-210 fleet (whose fp16 tensor cores are firmware-limited) dp4a is the fast path, not a compromise.
32
+
33
+ ## How to run
34
+
35
+ [fni8-serve](https://github.com/jajmangold/fni8-serve) is the LLM runtime (`load_fni8_state_dict(<file>)` into an `LLMEngine`; the architecture is read from the file). ComfyUI-fni8 is for diffusion DiTs only and does not load this model.
36
+
37
+ ## Limitations
38
+
39
+ - Quantization is lossy: int8 (and especially int4) outputs differ from the fp16/bf16 parent, and the difference varies by task.
40
+ - Capabilities, biases, and risks of the parent model carry over — see the parent card.
41
+ - This is a derivative quantization, not a relicense; the parent model's license and acceptable uses apply.
42
+
43
+ ---
44
+
45
+ Part of the fni8 stack: [kernels](https://github.com/jajmangold/fni8) · [LLM serving](https://github.com/jajmangold/fni8-serve) · [ComfyUI DiTs](https://github.com/jajmangold/ComfyUI-fni8).