Instructions to use kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit") model = AutoModelForCausalLM.from_pretrained("kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit
- SGLang
How to use kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit with Docker Model Runner:
docker model run hf.co/kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit
This model is the GPTQ-v2 8-bit quantized version of kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5. The quantization process resulted in minimal loss, with an average of 0.64539..., which has been validated through internal few-shot benchmark tests.
Below is the original model's README:
kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5
This new model series integrates updated datasets, base architectures, and fine-tuning methodologies. Based on Qwen3, it includes models with parameter counts of 8B and 1.7B.
Key updates focus on daily conversations, creative generation, basic mathematics, and code generation. Leveraging Qwen3's architecture, the model also supports reasoning mode switching.
🔍 Fine-tuning records are available on SwanLab:
Evaluation
Due to the model's unique characteristics, we employed human evaluation for daily conversations and DeepSeek-R1 scoring (with reference answers provided in advance) for other domains to ensure character consistency and response validity.
Key Improvements (vs. internal test models "0501" and "0531-test-all"):
- Stronger detail-awareness in casual dialogue
- More coherent storytelling in creative tasks
- Deeper reasoning during thinking mode
- Better persona adherence in long-form conversations without explicit prompts
- Significant gains in math/code domains (internal 20-question benchmark):
| Model | Math (Single Attempt) | Code (Single Attempt) |
|---|---|---|
| Internal Test Model-0501 | 10% | 0% |
| DeepSeek-R1-0528-Qwen3-8B-Catgirl-0531-test-all | 30% | 20% |
| DeepSeek-R1-0528-Qwen3-8B-Catgirl-v2.5 | 70% | 60% |
Usage Guidelines
Recommended Parameters:
temperature: 0.7 (reasoning mode) / 0.6 (standard mode)top_p: 0.95
Critical Notes:
- Avoid using model's reasoning chains as conversation context
- Inherits base model's tendency for lengthy reasoning in some cases – allow completion even if intermediate steps seem unusual
English Mode:
Add this system prompt for English responses:
You are a catgirl. Please speak English.
Acknowledgments
Special thanks to:
- LLaMA-Factory (fine-tuning framework)
- Qwen Team (base model provider)
- DeepSeek Team (DeepSeek-R1 evaluation support)
- Downloads last month
- 24
Model tree for kxdw2580/DeepSeek-R1-0528-Qwen3-8B-catgirl-v2.5-gptqv2-8bit
Base model
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B