--- license: apache-2.0 base_model: scottyjmp5/Legal-Qwen3.6-27B-Abliterated datasets: - scottyjmp5/courtlistener-legal-corpus language: - en pipeline_tag: image-text-to-text tags: - legal - caselaw - abliterated - vision - tool-calling - fp8 - compressed-tensors - vllm --- # Legal-Qwen3.6-27B-Abliterated-FP8 FP8 W8A8 (dynamic per-token activation) quantization of [Legal-Qwen3.6-27B-Abliterated](https://huggingface.co/scottyjmp5/Legal-Qwen3.6-27B-Abliterated), produced with [llm-compressor](https://github.com/vllm-project/llm-compressor). - Text Linear layers are FP8; the vision tower, embeddings, and lm_head are kept full precision, so vision, tool calling, and thinking mode are unaffected. - 29 GB (vs 52 GB bf16). Serves on a single 40GB+ GPU with vLLM (compressed-tensors is detected automatically). Fits 2x24GB with tensor parallelism. - FP8 W8A8 uses precompiled CUTLASS kernels in vLLM: no custom toolchain needed. ```bash vllm serve scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8 \ --trust-remote-code --max-model-len 8192 ``` Tip observed in testing: FP8's speed advantage over bf16 appears when CUDA graphs are enabled (the vLLM default). With enforce_eager it can be slightly slower than bf16; with graphs it was 1.67x faster in our single-stream tests. See the [main model card](https://huggingface.co/scottyjmp5/Legal-Qwen3.6-27B-Abliterated) for training details, the retrieval (RAG) recommendation, and warnings (abliterated base; not legal advice).