--- license: apache-2.0 base_model: empero-ai/Qwen3.8-4B-Distill tags: - text-generation - qwen3_5 - distillation - reasoning - full-stack - typescript - react-router-v8 - nestjs - flutter - unsloth - gguf - quantized inference: false model_type: qwen3_5 --- # Qwen3.8 4B Distill by Empero-AI and finetuned by iWebRoot — GGUF ## Model Overview This repository contains the optimized GGUF quantization of **Qwen3.8-4B-Empero-AI-Distill-FullStack**, fine-tuned using the **Unsloth** framework for advanced full-stack web and mobile software development pipelines. The base architecture features a full-parameter distillation of reasoning traces (Chain-of-Thought via `...` tags) from the frontier-scale **Qwen3.8 2.4T A95B** teacher model developed by **Empero-AI**. This configuration offers advanced local planning, logic, and code compilation compliance within a highly efficient 4-billion parameter footprint. ## 🔗 Repository Links * **Safetensors Version (9.3GB Heavy Build):** [https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack](https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack) * **GGUF Version (3.5GB Optimized Quantization):** [https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack-GGUF](https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack-GGUF) ## 📚 Injected Knowledge Stack (Fine-Tuning Data) The model underwent continuous pre-training on **279,049 curated data segments** across 8 strictly isolated developer knowledge directories: * **Mobile / Cross-Platform:** Flutter (Modern structural widgets and lifecycle state management). * **Full-Stack Web Architecture:** React Router v8 (Framework Mode via Vite, server-loaders, and async server-actions routing). * **Backend & Runtime Engine:** NestJS & Node.js (API architecture, scalable middleware, and server streams). * **Database & Persistence Layers:** Prisma ORM & Drizzle ORM (Schema modeling, relational builders, and safe SQL migrations). * **Language & System Rigor:** TypeScript (Strict typing patterns to enforce self-debugging and runtime stability). * **Design & UI Systems:** Tailwind CSS & Shadcn UI / Radix Primitives (Utility class layout embedded in JSX/TSX components). ## 📊 Training Logs & Learning Curve The fine-tuning process completed 250 hardware-optimized steps on a T4 GPU. The learning curve showed a definitive late convergence ("Eureka" moment) near step 140, where the weights successfully aligned cross-stack framework logic. * **Step 10 (Start):** Loss = `3.643480` * **Step 50:** Loss = `2.829789` * **Step 140 (Logical drop):** Loss = `2.788493` * **Step 250 (Final score):** Loss = `2.508277` ## 📦 Quantization Specifications * **File:** `Qwen3.8-4B-Empero-AI-Distill-FullStack-Q6_K.gguf` * **Format:** Q6_K (6-bit quantization) * **Size:** ~3.56 GB * **Quality:** Near-lossless precision compared to the 16-bit reference build. ## 💻 Local Execution Guide (Target: GTX 1050 4GB VRAM) When deploying this GGUF file inside **Unsloth Desktop**, **LM Studio**, **Jan**, or **Ollama**, configure these 3 runtime settings to prevent system stuttering: 1. **GPU Offload:** Set your hardware layer slider to **25 layers**. This safely loads ~2.5 GB of the model weight into your **NVIDIA GTX 1050** VRAM without freezing Windows, while the remaining compute safely overflows into your 16GB system RAM. 2. **Sampling Settings:** Set `temperature=0.6`, `top_p=0.95`, and `top_k=20`. Avoid a raw greedy search (temperature=0) to prevent the reasoning tokens from falling into endless structural loops. 3. **Context Window:** Set the token length to **`16384` or `32768`**. This expanded context window allows autonomous agents to evaluate several source files at the same time. ## 🛠️ Execution with OpenCode Autonomous Agent To launch this model as an active developer backend connected to your terminal agent, run the OpenAI-compatible local engine server: ```bash unsloth start opencode --context-length 32000 ``` ## Support / Donate If this model helped you, consider supporting the project: - **BTC**: `18cBC5sFjtctw121ULTkxTbTZPurginJBs` - **LTC**: `ltc1q3jrcwrx66xpz4k92p08u8c5v8zwywk3dqpzdkv` - **USDT**: `TGKVpbbznmvEusKbuZZj4WSK6XxtHcG6FE` (TRX chain) - **USDT**: `0x1059cb5a1F8467e5b56a9bdf082cE86FFB002D15` (POL chain) - **USDT**: `0x18b2AA731daeFD47DFFa278f3F856eAF80376fd6` (ETH chain) - **USDT**: `0x3bEcddC7c49bDba5503eB1677628b4519439884c` (BNB chain) ## Provenance & Licensing Quantizations are built upon **[empero-ai/Qwen3.8-4B-Distill](https://huggingface.co/empero-ai/Qwen3.8-4B-Distill)**. Weights inherit the permissive **Apache-2.0** license from the base Qwen repository and are shared as-is.