scottyjmp5 commited on
Commit
680a68d
·
verified ·
1 Parent(s): f2c102c

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +43 -0
README.md ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: scottyjmp5/Legal-Qwen3.6-27B-Abliterated
4
+ datasets:
5
+ - scottyjmp5/courtlistener-legal-corpus
6
+ language:
7
+ - en
8
+ pipeline_tag: image-text-to-text
9
+ tags:
10
+ - legal
11
+ - caselaw
12
+ - abliterated
13
+ - vision
14
+ - tool-calling
15
+ - fp8
16
+ - compressed-tensors
17
+ - vllm
18
+ ---
19
+
20
+ # Legal-Qwen3.6-27B-Abliterated-FP8
21
+
22
+ FP8 W8A8 (dynamic per-token activation) quantization of
23
+ [Legal-Qwen3.6-27B-Abliterated](https://huggingface.co/scottyjmp5/Legal-Qwen3.6-27B-Abliterated),
24
+ produced with [llm-compressor](https://github.com/vllm-project/llm-compressor).
25
+
26
+ - Text Linear layers are FP8; the vision tower, embeddings, and lm_head are kept full precision,
27
+ so vision, tool calling, and thinking mode are unaffected.
28
+ - 29 GB (vs 52 GB bf16). Serves on a single 40GB+ GPU with vLLM (compressed-tensors is
29
+ detected automatically). Fits 2x24GB with tensor parallelism.
30
+ - FP8 W8A8 uses precompiled CUTLASS kernels in vLLM: no custom toolchain needed.
31
+
32
+ ```bash
33
+ vllm serve scottyjmp5/Legal-Qwen3.6-27B-Abliterated-FP8 \
34
+ --trust-remote-code --max-model-len 8192
35
+ ```
36
+
37
+ Tip observed in testing: FP8's speed advantage over bf16 appears when CUDA graphs are enabled
38
+ (the vLLM default). With enforce_eager it can be slightly slower than bf16; with graphs it was
39
+ 1.67x faster in our single-stream tests.
40
+
41
+ See the [main model card](https://huggingface.co/scottyjmp5/Legal-Qwen3.6-27B-Abliterated) for
42
+ training details, the retrieval (RAG) recommendation, and warnings (abliterated base; not legal
43
+ advice).