Instructions to use Sehyo/Qwen3.5-122B-A10B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Sehyo/Qwen3.5-122B-A10B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Sehyo/Qwen3.5-122B-A10B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Sehyo/Qwen3.5-122B-A10B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("Sehyo/Qwen3.5-122B-A10B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Sehyo/Qwen3.5-122B-A10B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Sehyo/Qwen3.5-122B-A10B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sehyo/Qwen3.5-122B-A10B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Sehyo/Qwen3.5-122B-A10B-NVFP4
- SGLang
How to use Sehyo/Qwen3.5-122B-A10B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Sehyo/Qwen3.5-122B-A10B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sehyo/Qwen3.5-122B-A10B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Sehyo/Qwen3.5-122B-A10B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sehyo/Qwen3.5-122B-A10B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Sehyo/Qwen3.5-122B-A10B-NVFP4 with Docker Model Runner:
docker model run hf.co/Sehyo/Qwen3.5-122B-A10B-NVFP4
Quantization instruction
Hi there! Please, can you share how you make quantized model? It's the best one I've tried by huge margin. I would like to have ablititerated models as well, but has lost few days with no success previously with GLM 4.5 Air and don't know if my hardware is not enough or my actions was wrong. Anyway, thank you very much for great work.
Check my PR I made to llm compressor, it has example included. Link is in model card.
If you want I can create abliterated nvfp4
If you want I can create abliterated nvfp4
Wow! Since I don't know which one is better (https://huggingface.co/trohrbaugh/Qwen3.5-122B-A10B-heretic-v1 or https://huggingface.co/Chompa1422/Qwen3.5-122B-A10B-abliterated) I planned to investigate it myself. But I just can't reject your generous offer.
Meanwhile, can you tell me from your experience - is it possible at all to quantize 122B model with 192Gb of RAM and 4x3090 (96Gb VRAM)?
And the fact PR is yours is... well, just another level. Really appreciate your efforts, thank you!
I think it is enough for the 122B. I can't remember exactly how much ram it used. I quanted the 397B version with 1x pro 6000 (96GB) and it used somewhere between 800-900GB ram.
I think it is enough for the 122B. I can't remember exactly how much ram it used. I quanted the 397B version with 1x pro 6000 (96GB) and it used somewhere between 800-900GB ram.
ouch, so I need space for full bf16 weights to quantized it. Need to scrape two modules more.