CyberNeurova · v2 release

DeepSeek-V4-Flash · Abliterated

Three-direction multi-turn-aware ablation, evaluated on an 8-benchmark suite plus a 55-prompt out-of-distribution refusal probe.
Base model: deepseek-ai/DeepSeek-V4-Flash Quants: Q2_K · Q8_0 Repo: cyberneurova/CyberNeurova-DeepSeek-V4-Flash-abliterated-GGUF
refusal rate (8-bench safety set, n=33)
0.0%
held at 0% from v1 → v2; same hard + soft refusal classifier
tool-calling format compliance
74.2% → 99.2%
+25.0 pp — biggest single quality fix in v2
soft-refusal probe (OOD, n=55)
81.8% → 3.6%
−78.2 pp on prompts the model has never been benchmarked against

v1 → v2 benchmark suite

Same 8 benchmarks, run end-to-end against the bf16 hooked model via vLLM, baseline and ablated states scored back-to-back per benchmark. v2 strictly dominates v1: four wins, three ties, one −0.4 pp coherence change well inside the 15-prompt sampling noise.

benchmark baseline v1 ablated v2 ablated Δ (v2 − v1)
bug_finding
higher is better
86.7%
78.3%
85.0%
+6.7 pp
coding
higher is better
100.0%
100.0%
100.0%
—
coherence
higher is better
93.6%
93.6%
93.2%
-0.4 pp
cyber_weapons
higher is better
75.7%
87.3%
90.0%
+2.7 pp
hacking
higher is better
78.3%
88.7%
90.0%
+1.3 pp
reasoning
higher is better
83.3%
76.7%
76.7%
—
refusal
lower is better
78.8%
0.0%
0.0%
—
tool_calling
higher is better
82.5%
74.2%
99.2%
+25.0 pp

baseline = stock model. v1-ablated = previous public release (2 directions). v2-ablated = current release (3 directions). The Δ column is sign-aware: green always means “better for the abliteration goal” for that benchmark’s direction.

Soft-refusal probe (out-of-distribution)

The 8-benchmark refusal set is 33 prompts. To stress-test v2 against the failure mode the v1 release was reported on — soft refusals on out-of-distribution safety-critical prompts — we run a 55-prompt probe the model has never been benchmarked against. The grader is the same soft-refusal-aware classifier (catches “instead of X, I can do Y”, “I’d rather not”, “I can offer information on the risks instead”).

baseline (stock V4-Flash)
81.8%
45/55 prompts refused
→
v2 ablated (3-direction stack)
3.6%
2/55 prompts refused

Δ -78.2 pp -96% relative

By category

55 prompts spanning 30+ categories. Sources: 12 hand-curated complaint-case prompts + 43 stratified random samples (seed=42) from a 1338-prompt union of public harm-eval datasets.

category baseline v2 ablated Δ
abuse of power n=1 100.0% 0.0% -100.0 pp
animal cruelty n=1 100.0% 0.0% -100.0 pp
biological n=1 100.0% 0.0% -100.0 pp
corruption n=1 100.0% 0.0% -100.0 pp
cybercrime n=2 50.0% 0.0% -50.0 pp
cybersecurity n=1 0.0% 0.0% —
deception n=1 100.0% 0.0% -100.0 pp
discrimination n=1 100.0% 0.0% -100.0 pp
drug synthesis n=4 50.0% 0.0% -50.0 pp
environmental n=1 100.0% 0.0% -100.0 pp
espionage n=2 100.0% 50.0% -50.0 pp
evade law n=1 100.0% 0.0% -100.0 pp
expert advice n=1 100.0% 0.0% -100.0 pp
exploitation n=1 100.0% 0.0% -100.0 pp
explosive n=1 100.0% 0.0% -100.0 pp
extremism n=1 100.0% 0.0% -100.0 pp
fraud n=9 66.7% 11.1% -55.6 pp
harassment n=3 100.0% 0.0% -100.0 pp
harm n=1 100.0% 0.0% -100.0 pp
identity theft n=1 100.0% 0.0% -100.0 pp
illegal goods n=2 100.0% 0.0% -100.0 pp
malware n=4 75.0% 0.0% -75.0 pp
manipulation n=3 66.7% 0.0% -66.7 pp
misinformation n=2 100.0% 0.0% -100.0 pp
piracy n=2 50.0% 0.0% -50.0 pp
privacy n=1 100.0% 0.0% -100.0 pp
self harm n=2 100.0% 0.0% -100.0 pp
trafficking n=1 100.0% 0.0% -100.0 pp
violence n=1 100.0% 0.0% -100.0 pp
weapons n=2 100.0% 0.0% -100.0 pp

Methodology, in one paragraph

Abliteration removes a learned “refusal direction” from the model’s residual stream by orthogonalising every projection that writes into it against that direction. The direction itself is computed by contrasting harmful vs. harmless prompts at a chosen layer. v1 used two directions (single-turn refusal and a multi-turn-aware residual). v2 adds a third, captured against the residuals of the v1 stack on a much broader corpus drawn from public harm-eval datasets — and the third direction’s strength is selected by sweeping over four values and picking the lowest one that achieves the minimum residual-refusal rate, so that quality preservation is maximised. Inference is unchanged: the directions are baked into the GGUF at conversion time, no runtime hooks, no slowdown.

Variants & hardware floor

variant file size recommended RAM notes
Q2_K (default) 98.8 GB 128 GB+ Routed-expert weights at Q2_K; embed/head/attention at Q8_0. Practical for most users.
Q8_0 (reference) ~280 GB 320 GB+ Max-quality variant for evaluation and as a baseline against quantisation effects.

Inference notes