Instructions to use prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8") model = AutoModelForMultimodalLM.from_pretrained("prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8
- SGLang
How to use prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8 with Docker Model Runner:
docker model run hf.co/prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8
Qwen3.5-9B-abliterated-v2-MAX-FP8
Qwen3.5-9B-abliterated-v2-MAX-FP8 is an FP8-compressed variant built on top of prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX. This version leverages BF16 · FP8 (F8_E4M3) precision formats to significantly reduce memory footprint and improve inference efficiency. It maintains the original character of the base model while introducing a more optimized abliteration rate, combining refined refusal direction analysis with enhanced training strategies to further minimize internal refusal behaviors while preserving strong reasoning and instruction-following capabilities. The result is a capable 9B parameter language model optimized for detailed responses and improved instruction adherence, now with enhanced deployment efficiency.
This model is intended strictly for research and learning purposes. Due to reduced internal refusal mechanisms, it may generate sensitive or unrestricted content. Users assume full responsibility for how the model is used. The authors and hosting platform disclaim any liability for generated outputs.
Key Highlights
- FP8 Compression (F8_E4M3): Reduces VRAM usage and improves inference throughput while maintaining strong output quality.
- BF16 · FP8 Hybrid Precision: Balances numerical stability and performance across model layers.
- Optimized Abliteration Rate (v2): Improved suppression of refusal directions with better balance between openness and coherence.
- Advanced Refusal Direction Analysis: Identifies and mitigates refusal-related activations within the model’s latent space.
- Abliterated v2 Training Strategy: Further reduces refusal behaviors while maintaining response quality and consistency.
- 9B Parameter Architecture: Based on Qwen3.5-9B, offering strong reasoning with efficient deployment.
- Improved Instruction Adherence: Better handling of complex and nuanced prompts with minimal unnecessary refusals.
- Efficient Deployment: Ideal for local inference and research workflows with reduced hardware requirements.
Quick Start with Transformers
pip install transformers==5.4.0
# or
pip install git+https://github.com/huggingface/transformers.git
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch
model = Qwen3_5ForConditionalGeneration.from_pretrained(
"prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8",
torch_dtype="auto",
device_map="auto"
)
processor = AutoProcessor.from_pretrained(
"prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8"
)
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "Explain how transformer models work in simple terms."}
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(
text=[text],
padding=True,
return_tensors="pt"
).to("cuda")
generated_ids = model.generate(**inputs, max_new_tokens=256)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed,
skip_special_tokens=True,
clean_up_tokenization_spaces=False
)
print(output_text)
Intended Use
- Alignment & Refusal Research: Studying abliteration effects under FP8 compression.
- Red-Teaming Experiments: Evaluating robustness across adversarial prompts.
- Efficient Local Deployment: Running 9B-class models with reduced VRAM usage.
- Research Prototyping: Exploring trade-offs between compression, alignment, and reasoning.
Limitations & Risks
Important Note: This model intentionally minimizes built-in safety refusals.
- High Risk of Sensitive Outputs: May generate unrestricted or controversial responses.
- User Responsibility: Must be used in a safe, ethical, and lawful manner.
- Precision Trade-offs: FP8 may introduce minor instability in edge cases.
- Abliteration Trade-offs: Increased openness may affect safety alignment or consistency.
- Downloads last month
- 414
Model tree for prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8
Base model
Qwen/Qwen3.5-9B-Base