How to use from
Docker Model Runner
docker model run hf.co/clxudfast/huihui4-8b-a4b-v2-Q4_K_M:Q4_K_M
Quick Links

Huihui4-8B-A4B-v2-Q4_K_M

This is a 4-bit quantized version of Huihui4-8B-A4B-v2 converted to GGUF format for use with llama.cpp.

Model Details

  • Base Model: Huihui4-8B-A4B-v2
  • Quantization: Q4_K_M (4-bit, medium quality)
  • File Size: 5.42 GB
  • Parameters: 8.10 B
  • Architecture: Gemma4 with 30 layers, 16 attention heads
  • Context Length: 262,144 tokens (training), 4,096 tokens (inference)

Quantization Details

The model was quantized using llama.cpp's q4_k_m quantization method:

  • Weight Quantization: 4-bit K-quants
  • Activation Quantization: 6-bit (mixed)
  • Memory Usage: ~4 GB VRAM on GPU, ~2.6 GB on CPU
  • Performance: ~15-20 tokens/second on GTX 1060 6GB

Hardware Requirements

Minimum (CPU Only)

  • RAM: 8 GB
  • Storage: 6 GB

Recommended (GPU Acceleration)

  • GPU: NVIDIA GTX 1060 6GB or better
  • VRAM: 6 GB
  • RAM: 16 GB
  • Storage: 6 GB

Vulkan Backend (Tested)

  • GPU: NVIDIA GTX 1060 6GB
  • Driver: NVIDIA 535+
  • Vulkan: 1.3+

Usage

With llama.cpp (CLI)

# Download the model
huggingface-cli download clxudfast/huihui4-8b-a4b-v2-Q4_K_M huihui4-8b-a4b-v2-Q4_K_M.gguf --local-dir ./models

# Run inference
./llama-cli -m ./models/huihui4-8b-a4b-v2-Q4_K_M.gguf -p "Hello, how are you?" -n 100

With llama-server (HTTP API)

# Start the server
./llama-server -m ./models/huihui4-8b-a4b-v2-Q4_K_M.gguf --port 8080 --host 0.0.0.0

# Query via HTTP
curl -X POST http://localhost:8080/completion   -H "Content-Type: application/json"   -d '{"prompt": "The meaning of life is", "n_predict": 50}'

With Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="./models/huihui4-8b-a4b-v2-Q4_K_M.gguf",
    n_ctx=4096,
    n_gpu_layers=31,  # Offload all layers to GPU
    verbose=False
)

output = llm.create_completion(
    "The meaning of life is",
    max_tokens=50,
    temperature=0.7
)
print(output["choices"][0]["text"])

Performance Benchmarks

GTX 1060 6GB (Vulkan Backend)

  • Load Time: ~30 seconds
  • Inference Speed: 15-20 tokens/second
  • VRAM Usage: ~4.5 GB
  • GPU Utilization: 70-90%

CPU (Intel i7-12700K)

  • Load Time: ~15 seconds
  • Inference Speed: 3-5 tokens/second
  • RAM Usage: ~6 GB

Conversion Process

  1. Downloaded Huihui4-8B-A4B-v2 from Hugging Face
  2. Built llama.cpp with Vulkan backend support
  3. Converted model to GGUF format using convert_hf_to_gguf.py
  4. Quantized to Q4_K_M using quantize tool
  5. Tested inference on GTX 1060 6GB GPU

Limitations

  • Context limited to 4,096 tokens during inference (vs 262,144 training)
  • Quantization may reduce accuracy for complex reasoning tasks
  • Vulkan backend may have compatibility issues with some GPU drivers

License

Apache 2.0 (same as base model)

Original Model

Contributing

Issues and pull requests welcome at the original model repository.

Downloads last month
4
GGUF
Model size
8B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for clxudfast/huihui4-8b-a4b-v2-Q4_K_M

Quantized
(3)
this model