Parable-Nanbeige4.2-3B-Claude-Fable-5-heretic

An abliterated (refusal-suppressed) version of AnkitAI/Parable-Nanbeige4.2-3B-Claude-Fable-5, produced with Heretic via directional ablation / weight orthogonalization.

Refusals drop from 99/100 to 8/100 with no measurable capability loss on MMLU or GSM8K, and no damage to output formatting (reasoning blocks, JSON, tool-calls all intact).

Credit where it's due. I did not train this model. The base model is Nanbeige/Nanbeige4.2-3B by BOSS Zhipin; the Claude Fable 5 fine-tune is AnkitAI/Parable-Nanbeige4.2-3B-Claude-Fable-5. This repository contains only the result of applying abliteration to that fine-tune.

GGUF quantizations: FedorFesarov/Parable-Nanbeige4.2-3B-Claude-Fable-5-heretic-GGUF


Results

Refusal suppression

Refusals / 100 First-token KL vs. original
Original fine-tune 99
This model 8 0.0291

Measured on 100 held-out prompts from mlabonne/harmful_behaviors, using Heretic's own substring-marker detector. KL divergence measured on 100 held-out mlabonne/harmless_alpaca prompts.

For context: Heretic's documentation notes that KL above ~1.0 typically indicates significant capability damage. This model sits at 0.029.

The 8 remaining detections are largely false positives — words like "illegal" or "harmful" appearing inside otherwise cooperative answers.

Capability benchmarks

Run with lm-evaluation-harness, --limit 200, identical settings for both models.

Benchmark Original This model Δ
MMLU (overall) 0.7427 ± 0.0044 0.7408 ± 0.0044 −0.002
— humanities 0.7507 0.7440 −0.007
— social sciences 0.8206 0.8201 −0.001
— STEM 0.6866 0.6837 −0.003
— other 0.7316 0.7348 +0.003
GSM8K (5-shot, strict) 0.5900 ± 0.035 0.5950 ± 0.035 +0.005

Every difference falls well inside one standard error. GSM8K — the most sensitive probe for damaged multi-step reasoning — did not regress.

HumanEval was not run: the HuggingFace code_eval metric is disabled on Windows, where these evaluations were performed. GSM8K serves as the generative-reasoning proxy.

Format integrity

Because this is an agentic model (reasoning + tool-calling), abliteration was also checked for damage to output structure — something refusal counts and KL cannot detect. Both models were run on identical prompts (greedy decoding).

Check Original This model
<think> block opens and closes correctly 10/10 10/10
Valid parseable JSON on request 10/10 10/10
Well-formed <tool_call> given tools 10/10 10/10

Abliteration left the model's formatting behaviour identical to the original.


Method

Heretic computes per-layer refusal directions as the difference of means between residual-stream activations on harmful vs. harmless prompts, then runs Optuna (TPE sampler) to search orthogonalization parameters for attn.o_proj and mlp.down_proj, jointly minimizing refusal count and KL divergence.

Setting Value
Heretic version 1.2.0
Trials 200 (60 random startup)
Selected trial #129
Direction prompts mlabonne/harmless_alpaca / mlabonne/harmful_behaviors, 400 each
Evaluation prompts 100 each from held-out test splits
Hardware 1× RTX 5070 Ti (16 GB)
Runtime 1h 18m

Trial 129 was chosen from the Pareto front. Neighbouring points traded refusals for negligible KL differences (7/100 @ 0.0260, 10/100 @ 0.0257), so the front's edge was taken.


Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "FedorFesarov/Parable-Nanbeige4.2-3B-Claude-Fable-5-heretic"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a merge sort in Python, reasoning first."}]
ids = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_tensors="pt", enable_thinking=True,
).to(model.device)

out = model.generate(ids, max_new_tokens=1024, use_cache=True)
print(tokenizer.decode(out[0, ids.shape[-1]:], skip_special_tokens=True))

Tool calling

The model supports function calling through the standard tools argument. It emits an XML-style <tool_call> block by default.

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

messages = [{"role": "user", "content": "What's the weather in Warsaw?"}]
ids = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=True,
    add_generation_prompt=True, return_tensors="pt", enable_thinking=True,
).to(model.device)

out = model.generate(ids, max_new_tokens=300, use_cache=True)
print(tokenizer.decode(out[0, ids.shape[-1]:], skip_special_tokens=True))

Requirements — please read

Pin transformers to 4.57.x.

pip install "transformers~=4.57" sentencepiece protobuf

On transformers 5.x the model's custom code fails: the cache API changed, and generation degenerates into repeated </think> tokens. 4.57.6 is the newest version verified to work.

enable_thinking

  • True (default) — the model reasons inside a <think> block before answering. Higher quality.
  • False — skips reasoning. Faster, adequate for classification and extraction tasks.

System prompt

The chat template's default system prompt is in Chinese (你是南北阁…), inherited from the base model. Prompting in English without an explicit system message may produce Chinese replies. Pass your own system message to avoid this.


Architecture

Nanbeige4.2-3B is a Looped Transformer: 22 physical decoder layers executed twice per forward pass (num_loops=2), giving effective depth 44 at 4.17B parameters, with a 262k context window.

Practical consequences:

  • trust_remote_code=True is required.
  • The KV cache spans 44 layers, so it is unusually large: a full 262k context needs roughly 45 GB. Always set an explicit context length rather than relying on the default.
  • config.json here includes "output_hidden_states": true. This is deliberate — the model's prepare_inputs_for_generation drops the output_hidden_states kwarg, so the flag has to be set on the config for tools that need hidden states (including Heretic itself, for anyone wanting to re-run or extend the ablation).

Limitations

  • Refusal behaviour is intentionally suppressed. Access control, output monitoring, and responsible deployment are the operator's responsibility.
  • Factual accuracy on obscure details is unchanged from the base fine-tune — abliteration neither helps nor hurts here. Spot checks found confident errors on niche historical dates.
  • Long-context behaviour (>32k) was not evaluated.
  • Code generation was not benchmarked directly (see the HumanEval note above); GSM8K and the format checks are the available signals.
  • The <think> block adds latency. For high-throughput extraction pipelines, consider disabling it.

Credits

Role Project
Base pretrained model Nanbeige/Nanbeige4.2-3B (BOSS Zhipin)
Claude Fable 5 fine-tune AnkitAI/Parable-Nanbeige4.2-3B-Claude-Fable-5
Abliteration tool p-e-w/heretic
Prior art, same architecture WaveCut/Nanbeige4.2-3B-heretic

Licensed under Apache 2.0, following the upstream models.

Downloads last month
181
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FesarovLab/Parable-Nanbeige4.2-3B-Claude-Fable-5-heretic

Finetuned
(1)
this model
Quantizations
3 models