How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for ReallyFloppyPenguin/DeepSeek-R1-Distill-Qwen-32B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for ReallyFloppyPenguin/DeepSeek-R1-Distill-Qwen-32B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for ReallyFloppyPenguin/DeepSeek-R1-Distill-Qwen-32B-GGUF to start chatting
Quick Links

deepseek-ai/DeepSeek-R1-Distill-Qwen-32B - GGUF

This repository contains GGUF quantizations of deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.

About GGUF

GGUF is a quantization method that allows you to run large language models on consumer hardware by reducing the precision of the model weights.

Files

Filename Quant type File Size Description
model-fp16.gguf FP16 Variable fp16 quantization
model-q4_0.gguf Q4_0 Small 4-bit quantization
model-q4_1.gguf Q4_1 Small 4-bit quantization (higher quality)
model-q5_0.gguf Q5_0 Medium 5-bit quantization
model-q5_1.gguf Q5_1 Medium 5-bit quantization (higher quality)
model-q8_0.gguf Q8_0 Large 8-bit quantization

Usage

You can use these models with llama.cpp or any other GGUF-compatible inference engine.

llama.cpp

./llama-cli -m model-q4_0.gguf -p "Your prompt here"

Python (using llama-cpp-python)

from llama_cpp import Llama

llm = Llama(model_path="model-q4_0.gguf")
output = llm("Your prompt here", max_tokens=512)
print(output['choices'][0]['text'])

Original Model

This is a quantized version of deepseek-ai/DeepSeek-R1-Distill-Qwen-32B. Please refer to the original model card for more information about the model's capabilities, training data, and usage guidelines.

Conversion Details

  • Converted using llama.cpp
  • Original model downloaded from Hugging Face
  • Multiple quantization levels provided for different use cases

License

This model inherits the license from the original model. Please check the original model's license for usage terms.

Downloads last month
4
GGUF
Model size
33B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ReallyFloppyPenguin/DeepSeek-R1-Distill-Qwen-32B-GGUF

Quantized
(143)
this model

Collection including ReallyFloppyPenguin/DeepSeek-R1-Distill-Qwen-32B-GGUF