scottyjmp5/courtlistener-legal-corpus
Updated • 1.12k
FP8 W8A8 (dynamic per-token activation) quantization of Legal-Qwen3.6-27B-Abliterated, produced with llm-compressor.
vllm serve scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8 \
--trust-remote-code --max-model-len 8192
Tip observed in testing: FP8's speed advantage over bf16 appears when CUDA graphs are enabled (the vLLM default). With enforce_eager it can be slightly slower than bf16; with graphs it was 1.67x faster in our single-stream tests.
See the main model card for training details, the retrieval (RAG) recommendation, and warnings (abliterated base; not legal advice).
Base model
Qwen/Qwen3.6-27B