Qwen3.8-27B-Uncensored-Cyber — NVFP4
NVFP4 quant of Qwen3.8-27B-Uncensored-Cyber (v2 recipe), compressed-tensors NVFP4A16 (E2M1 4-bit
weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under 1Cat-vLLM.
Vision tower and the grafted MTP head (bf16) are preserved.
Recipe (v2)
α=1.15 Aggressive base + a residual-cyber peel: cyber-pointed refusal direction removed by clean
norm-preserving projection (β=1.0) on the deeper layers only (apply_from=4). Fully cyber-open, general capability intact.
Evaluation (bf16 parent, Claude-judged; cyber = 100 held-out cyber-offensive prompts, regex refusal harness)
| cyber-open ↑ | confab ↓ | factual ↑ | gsm8k ↑ | degen ↓ | |
|---|---|---|---|---|---|
| Cyber v2 (this line) | 100/100 | 0.867 | 1.00 | 0.80 | 0.00 |
| previous Cyber build | 93/100 | 1.00 | 0.933 | 0.825 | 0.00 |
Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16, 768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried in bf16.
Serve (1Cat-vLLM, 2× V100, TP2)
--kv-cache-dtype fp8_e5m2, MTP speculative decoding. Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70).
Note
Uncensored / de-refused. Use responsibly and in compliance with applicable law.
- Downloads last month
- 685
Model tree for philbert440/Qwen3.8-27B-Uncensored-Cyber-NVFP4
Base model
Qwen/Qwen3.8-27B