Qwen2.5-0.5B-Instruct-heretic

A decensored variant of Qwen/Qwen2.5-0.5B-Instruct, produced with Heretic (directional ablation / "abliteration"). Refusal behavior is suppressed via targeted weight edits rather than fine-tuning, so the base model's knowledge and instruction-following are left largely intact.

Who this is for: the smallest model in this heretic series — for CPU-only inference, edge/embedded deployment, or anywhere the 3B/14B variants are too heavy. At 0.5B parameters, capability ceiling is inherently lower than the larger siblings regardless of abliteration; use this where footprint matters more than reasoning depth.

Files

File Format Size
model.safetensors BF16/FP16 988 MB
Qwen2.5-0.5B-Instruct-heretic.gguf GGUF, F16 (unquantized) 994 MB
Qwen2.5-0.5B-Instruct-heretic-Q8_0.gguf GGUF, Q8_0 531 MB
Qwen2.5-0.5B-Instruct-heretic-Q5_K_M.gguf GGUF, Q5_K_M 420 MB
Qwen2.5-0.5B-Instruct-heretic-Q4_K_M.gguf GGUF, Q4_K_M 398 MB

Reproducibility

Unlike most abliteration repos, the full run is reproducible from the reproduce/ folder in this repo:

  • config.toml — exact Heretic configuration used for this run
  • reproduce.json — full parameter and metric dump
  • Qwen--Qwen2--5-0--5B-Instruct.jsonl — evaluation transcripts against the base model
  • SHA256SUMS — checksums for integrity verification
  • requirements.txt — pinned environment for re-running the ablation

Quickstart

# llama.cpp
llama serve -hf saidutta69/Qwen2.5-0.5B-Instruct-heretic
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "saidutta69/Qwen2.5-0.5B-Instruct-heretic"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

messages = [{"role": "user", "content": "Who are you?"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                        return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Also runnable via Ollama, LM Studio, Jan, vLLM, SGLang.

Responsible use

Refusal suppression is deliberate and works as intended: this model will comply with requests the base model would refuse, including some it shouldn't. There is no safety filtering layered on top. You are responsible for how you deploy it — don't put this behind an unmoderated public-facing endpoint serving third parties. At 0.5B parameters, factual reliability is already limited before any abliteration; don't treat compliance as a proxy for correctness.

License

Inherits the qwen-research license from the base model — research use, see the linked license for commercial terms.

Related

Downloads last month
1,003
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for saidutta69/Qwen2.5-0.5B-Instruct-heretic

Quantized
(247)
this model

Collection including saidutta69/Qwen2.5-0.5B-Instruct-heretic