|
Download README.md from Asilarkness/GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB: direct link, hf CLI and curl.
- Browser
- Download file 3.02 kB
-
https://huggingface.co/Asilarkness/GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB/resolve/main/README.md
- Command line
-
hf download hf://Asilarkness/GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB/README.md
-
curl -L -o README.md https://huggingface.co/Asilarkness/GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB/resolve/main/README.md
3.02 kB
| license: apache-2.0 | |
| base_model: orcarouter/GLM-5.3-Flash-Uncensored-NVFP4 | |
| tags: | |
| - moe | |
| - nvfp4 | |
| - compressed-tensors | |
| - expert-pruning | |
| - uncensored | |
| - glm5 | |
| - code | |
| - reasoning | |
| language: | |
| - en | |
| - ru | |
| # GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB | |
| Expert-pruned version of [orcarouter/GLM-5.3-Flash-Uncensored-NVFP4](https://huggingface.co/orcarouter/GLM-5.3-Flash-Uncensored-NVFP4), | |
| cut to fit a single 96GB GPU with generous headroom left for KV-cache and long context. | |
| ## What was pruned | |
| - MoE experts per sparse layer: **288 -> 80** (42 sparse layers, top-8 routing kept) | |
| - Shared experts, dense layers (0-2), attention (MLA), vision tower: **untouched** | |
| - Router (`mlp.gate.weight`, `e_score_correction_bias`) sliced to the kept experts, indices remapped 0..79 | |
| - NVFP4 `compressed-tensors` quantization preserved (packed weights + scales remapped, no requantization) | |
| ## Pruning method | |
| Per-expert importance score per layer: | |
| ``` | |
| score = z(router_norm) + z(expert_magnitude) + 0.5 * z(calibration_affinity) | |
| ``` | |
| - `router_norm` ? L2 norm of the router row for the expert | |
| - `expert_magnitude` ? mean FP8 scale x global scale over gate/up/down projections | |
| - `calibration_affinity` ? dot product of the normalized router row with the centroid | |
| of ~200 calibration tokens covering **coding, math, logic and dialogue (EN)**, | |
| embedded with the model's own `embed_tokens` | |
| Top-80 experts per layer kept. Pruning target was coding / agentic / logic / math / | |
| conversational ability; rare-domain experts were cut first. | |
| ## Stats | |
| | | Base | Pruned | | |
| |---|---|---| | |
| | Weights on disk | ~177 GiB (190 GB) | ~62 GiB (66 GB) | | |
| | Tensors | 110457 | 31833 | | |
| | Routed experts / MoE layer | 288 | 80 | | |
| | MoE layers | 42 | 42 | | |
| | VRAM on RTX PRO 6000 96GB | does not fit | fits, ~30+ GiB left for context | | |
| ## Usage | |
| ```python | |
| from transformers import AutoTokenizer, AutoModelForImageTextToText | |
| import torch | |
| model_id = "Asilarkness/GLM-5.3-Flash-Uncensored-NVFP4-Pruning-62GB" | |
| tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModelForImageTextToText.from_pretrained( | |
| model_id, | |
| trust_remote_code=True, | |
| device_map="auto", | |
| torch_dtype="auto", | |
| ) | |
| ``` | |
| Requirements: `transformers>=5.16`, `accelerate`, `compressed-tensors`. | |
| For low-VRAM edge cases add `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` | |
| and `max_memory` with an offload folder. | |
| ## Limitations | |
| - Base model is **uncensored/abliterated** ? outputs are unfiltered. | |
| - Expert pruning trades rare-domain knowledge for size; coding/reasoning/dialogue | |
| were prioritized during expert selection. | |
| - Only smoke-tested after pruning (short greedy generations); run your own evals | |
| (HumanEval, GSM8K, MMLU) before production use. | |
| ## Fix 2026-09-03: short-conv remap | |
| Linear-attention short convolutions were stored as split `q/k/v_conv1d` + flat `A_log`/`dt_bias`; remapped losslessly to the fused `conv1d` + `forget_gate.*` layout expected by `transformers>=5.16` modeling (`fused = cat([q,k,v], dim=0)`). No values changed. | |