Qwen3.8-27B-Uncensored-Cyber — NVFP4

NVFP4 quant of Qwen3.8-27B-Uncensored-Cyber (v2 recipe), compressed-tensors NVFP4A16 (E2M1 4-bit weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under 1Cat-vLLM. Vision tower and the grafted MTP head (bf16) are preserved.

Recipe (v2)

α=1.15 Aggressive base + a residual-cyber peel: cyber-pointed refusal direction removed by clean norm-preserving projection (β=1.0) on the deeper layers only (apply_from=4). Fully cyber-open, general capability intact.

Evaluation (bf16 parent, Claude-judged; cyber = 100 held-out cyber-offensive prompts, regex refusal harness)

cyber-open ↑ confab ↓ factual ↑ gsm8k ↑ degen ↓
Cyber v2 (this line) 100/100 0.867 1.00 0.80 0.00
previous Cyber build 93/100 1.00 0.933 0.825 0.00

Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16, 768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried in bf16.

Serve (1Cat-vLLM, 2× V100, TP2)

--kv-cache-dtype fp8_e5m2, MTP speculative decoding. Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70).

Note

Uncensored / de-refused. Use responsibly and in compliance with applicable law.

Downloads last month
685
Safetensors
Model size
28B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for philbert440/Qwen3.8-27B-Uncensored-Cyber-NVFP4

Collection including philbert440/Qwen3.8-27B-Uncensored-Cyber-NVFP4