Image-Text-to-Text
Safetensors
GGUF
English
Chinese
qwen3_5_text
solstice-ai
davidau
davidau-quants
qwen
qwen3.8
qwen3.8-27b
cold-fusion
gain
project-heretic
heretic
uncensored
fable
cot
reasoning
coding
swe-bench
swe-bench-pro
beats-claude-opus-4.6
nvfp4
fp4
blackwell
rtx-5090
vllm
sglang
ultraoptimised
ultraefficient
efficient
arc-challenge
735-arc
882-arc
conversational
8-bit precision
compressed-tensors
Production Deployment & Serving Recipes
#2
by Dosamer - opened
Blackwell NVFP4 Serving via vLLM
vllm serve Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4 \ --quantization modelopt \ --max-model-len 262144 \ --port 8000
You suggest the above as a config for vllm. the config.json has "quant_method": "compressed-tensors" which is not modelopt
Are you sure the suggested serving template is correct?
I will double check, this was an automated upload via agents, sorry for the inconvenience