Legal-Qwen3.6-27B-Abliterated-FP8

FP8 W8A8 (dynamic per-token activation) quantization of Legal-Qwen3.6-27B-Abliterated, produced with llm-compressor.

  • Text Linear layers are FP8; the vision tower, embeddings, and lm_head are kept full precision, so vision, tool calling, and thinking mode are unaffected.
  • 29 GB (vs 52 GB bf16). Serves on a single 40GB+ GPU with vLLM (compressed-tensors is detected automatically). Fits 2x24GB with tensor parallelism.
  • FP8 W8A8 uses precompiled CUTLASS kernels in vLLM: no custom toolchain needed.
vllm serve scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8 \
  --trust-remote-code --max-model-len 8192

Tip observed in testing: FP8's speed advantage over bf16 appears when CUDA graphs are enabled (the vLLM default). With enforce_eager it can be slightly slower than bf16; with graphs it was 1.67x faster in our single-stream tests.

See the main model card for training details, the retrieval (RAG) recommendation, and warnings (abliterated base; not legal advice).

Downloads last month
15
Safetensors
Model size
27B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8

Dataset used to train scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8