--- license: apache-2.0 base_model: iBotIA/Qwen3.8-4B-Empero-AI-FullStack tags: - llama-cpp - gguf - qwen3 - quantized language: - en pipeline_tag: text-generation library_name: gguf --- # Qwen3.8-4B-Empero-AI-FullStack — GGUF Quantizations This repository contains GGUF quantizations of [iBotIA/Qwen3.8-4B-Empero-AI-FullStack](https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack). --- ## 📦 Available Quantizations | File | Bits | Size (approx.) | Use case | |------|------|----------------|----------| | `model_f16.gguf` | 16-bit | ~8.7 GB | Maximum quality, reference | | `model_q8_0.gguf` | 8-bit | ~4.7 GB | Near-lossless, high VRAM | | `model_q6_k.gguf` | 6-bit | ~3.6 GB | Excellent quality | | `model_q5_k_m.gguf` | 5-bit | ~3.1 GB | Great quality/size balance | | `model_q5_k_s.gguf` | 5-bit | ~3.0 GB | Slightly smaller than K_M | | `model_q4_k_m.gguf` | 4-bit | ~2.5 GB | Recommended default | | `model_q4_k_s.gguf` | 4-bit | ~2.4 GB | Smaller 4-bit variant | | `model_q3_k_l.gguf` | 3-bit | ~2.1 GB | Low VRAM, decent quality | | `model_q3_k_m.gguf` | 3-bit | ~1.9 GB | Balanced 3-bit | | `model_q3_k_s.gguf` | 3-bit | ~1.7 GB | Minimum 3-bit | | `model_q2_k.gguf` | 2-bit | ~1.3 GB | Extreme compression | > **IQ quants** (IQ4_XS, etc.) coming soon — require imatrix calibration. --- ## 🚀 Usage ### llama.cpp ```bash ./llama-cli -m model_q4_k_m.gguf -p "Your prompt here" -n 512 ``` ### LM Studio Download any `.gguf` file and load it directly in LM Studio. ### Ollama ```bash ollama run hf.co/tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF:Q4_K_M ``` ### Python (llama-cpp-python) ```python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="tinyopsec/Qwen3.8-4B-Empero-AI-FullStack-GGUF", filename="model_q4_k_m.gguf", ) output = llm("Your prompt here", max_tokens=512) print(output["choices"][0]["text"]) ``` --- ## 🔧 Quantization Details - **Source model:** [iBotIA/Qwen3.8-4B-Empero-AI-FullStack](https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack) - **Tool:** [llama.cpp](https://github.com/ggerganov/llama.cpp) - **Base format:** F16 GGUF - **IQ calibration data:** [groups_merged.txt](https://github.com/ggml-org/llama.cpp/discussions/5263) by kalomaze --- ## 💡 Which quant should I use? | VRAM | Recommended | |------|-------------| | 2 GB | Q2_K | | 3 GB | Q3_K_M | | 4 GB | Q4_K_M ✅ | | 6 GB | Q5_K_M | | 8 GB | Q6_K | | 12 GB+ | Q8_0 / F16 | --- ## 📄 License Refer to the [original model license](https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack).