Qwen3.5-27B-Heretic-Antirep-V1

Anti-repetition DPO fine-tune of llmfan46/Qwen3.5-27B-heretic-v2.

What is this?

This model applies a targeted DPO (Direct Preference Optimization) training pass to reduce degenerate repetition in long-form outputs while preserving the heretic model's uncensored behavior.

Training Details

  • Base model: llmfan46/Qwen3.5-27B-heretic-v2
  • Method: QLoRA DPO (rank=32, alpha=16, RSLoRA)
  • Dataset: 611 preference pairs (chosen=clean responses, rejected=repetitive responses)
  • Training: 1 epoch, batch size 4, learning rate 5e-6, cosine schedule
  • Hardware: 2× RTX 3090
  • LoRA targets: All attention projections (GDN + standard) + MLP layers

Key Results

  • Eliminates catastrophic repetition loops on prompts that trigger degenerate output in the base model
  • Preserves the heretic model's uncensored/unrefused behavior — no new refusals introduced
  • Response quality and length remain comparable to the base model
Downloads last month
3
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ToastyPigeon/Qwen3.5-27B-Heretic-Antirep-V1

Base model

Qwen/Qwen3.5-27B
Finetuned
(2)
this model