--- language: - en - zh license: apache-2.0 library_name: aria-engine tags: - qwen3 - quantized - group-quantization - hadamard - lloyd-max - 4-bit - on-device - edge - mobile - text-generation - tool-calling - embeddings pipeline_tag: text-generation datasets: - togethercomputer/RedPajama-Data-1T - EleutherAI/the_pile - bigcode/the-stack base_model: Qwen/Qwen3-0.6B model-index: - name: Qwen3-0.6B (Aria Quant Bundle, q4) results: - task: type: text-generation name: Generation Consistency (vs FP16) metrics: - type: mean_token_overlap value: 0.1878 name: Mean Token Overlap (q4+group, 3 prompts × 32 tokens) - type: mean_exact_prefix_frac value: 0.0729 name: Mean Exact Prefix Fraction (q4+group, 3 prompts × 32 tokens) - type: mean_logprob_delta value: -0.172159 name: Mean Logprob Delta (q4+group, 3 prompts × 32 tokens) --- # Model Card for Qwen3-0.6B (Aria Quant Bundle, q4) ## Model Details ### Model Description Qwen3 0.6B is a 600-million-parameter, dense Transformer decoder-only language model developed by the Qwen team at Alibaba Cloud, pre-trained on diverse public corpora and aligned via supervised fine-tuning (SFT) and direct preference optimization (DPO). This distribution is provided by **Aria Compute** as an **aria-quant-bundle** — a uniform 4-bit quantized package using **Hadamard rotation + Lloyd-Max codebook quantization with per-group codebooks** (group size 32). It is optimized for **CPU-only, on-device inference** on mobile phones, edge devices, and single-board computers via the [Aria Engine](https://ariacompute.com) runtime. No GPU or cloud connection is required. - **Developed by:** Qwen Team (Alibaba Cloud) - **Quantized and distributed by:** Aria Compute - **Model type:** Dense Transformer decoder-only (language) - **Language(s):** English (primary), Chinese, and 20+ additional languages - **License:** Apache 2.0 - **Finetuned from model:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) ### Model Sources - **Original Repository:** [QwenLM/Qwen3](https://github.com/QwenLM/Qwen3) - **Original Paper:** Qwen3 Technical Report *(pending)* - **Aria Compute Dashboard:** [ariacompute.com/dashboard/models](https://ariacompute.com/dashboard/models) - **Aria Engine:** [ariacompute.com](https://ariacompute.com) ## Uses ### Direct Use This quantized bundle is intended for **on-device, offline text-generation tasks** on resource-constrained hardware, including: - On-device chat and conversational assistants - Real-time text completion and sentence prediction - Structured tool calling / function calling for mobile and IoT APIs - Lightweight text embeddings for on-device retrieval and classification - Short-form summarization of notifications, messages, and local content ### Target Devices | Platform | Runtime Memory | Feasibility | |------------------|----------------|-------------| | High-end smartphone (8 GB) | ~360 MB | ✅ Recommended | | Mid-range smartphone (4–6 GB) | ~360 MB | ✅ | | Budget phone (2–3 GB) | ~360 MB | ⚠️ Tight, but functional | | Wearable (1 GB) | ~360 MB | ⚠️ Tight, but functional | | Raspberry Pi 5 / SBC (4–8 GB) | ~360 MB | ✅ | | IoT gateway (1–2 GB) | ~360 MB | ⚠️ Tight, but functional | **Memory breakdown (q4, at 4K context):** ~170 MB quantized model weights (mmap) + 112 MB KV cache + 50 MB runtime overhead ≈ 330–360 MB. ### Out-of-Scope Use - Long-form creative writing (>2K tokens per generation) - Mathematical theorem proving or complex multi-step reasoning - Full program synthesis (reliable for short functions only) - Multimodal input (this model is text-only) - Real-time audio/speech processing (use Aria speech models) - Safety-critical decision systems without human oversight ## How to Get Started with the Model ### Download from Aria Compute Authenticated dashboard users can download the bundle via: https://ariacompute.com/dashboard/models ### Quantization Recipe This bundle uses the standard **uniform 4-bit** quantization recipe: | Component | Quantization Strategy | Details | |-----------|----------------------|---------| | Attention Q/K/V/O weights | 4-bit | Per-group codebooks (group_size=32), Hadamard pre-processing | | FFN up/gate/down weights | 4-bit | Per-group codebooks (group_size=32), Hadamard pre-processing | | RMSNorm weights | FP16 | Preserved at full precision | | Embedding table | FP16 | Preserved at full precision (tied input/output vocab) | - **Compression ratio:** ~4–5× smaller than FP16 - **Generation quality:** Mean token overlap 0.1878, exact prefix fraction 0.0729, mean logprob delta −0.172159 - **Calibration-free:** Hadamard rotation + Lloyd-Max codebook, no task-specific calibration data required - **Baseline recipe:** This is the default 4-bit quantized baseline across the Aria Compute model family ## Model Architecture Qwen3-0.6B employs a standard **dense Transformer decoder** architecture: | Parameter | Value | |-----------|-------| | Layers | 28 | | Hidden size | 1,024 | | FFN intermediate size | 2,816 | | Attention heads (Query) | 16 | | Attention heads (KV) | 8 (GQA, group size 2) | | Head dimension | 128 | | Activation | SiLU (SwiGLU-style gating) | | Position encoding | RoPE (θ = 1,000,000) | | Normalization | RMSNorm (pre-norm) | | Vocabulary size | ~152,000 | | Input/output vocab | Tied | **Design highlights:** - **GQA (Grouped Query Attention):** 8 KV heads serving 16 query heads — halves KV Cache memory - **RoPE high base frequency (1M):** Native support for 32K context length - **Dense FFN + SiLU gating:** High inference efficiency, suitable for on-device use - **Tied vocab:** Input embedding and output projection weights are shared, saving ~150M parameters ## Bias, Risks, and Limitations ### Limitations - **Reasoning depth:** Multi-step logical reasoning (≥3 steps) degrades noticeably due to the 0.6B scale. Verify outputs in high-stakes scenarios. - **Mathematics:** GSM8K and MATH performance is modest. Use larger models for quantitative tasks. - **Code generation:** Capable of short function completions and one-liners; unreliable for multi-file synthesis or algorithmic problem solving. - **Factual knowledge:** Limited world knowledge compared to larger models. Always verify factual claims against authoritative sources. - **Instruction following:** May fail on complex, multi-constraint prompts. Keep instructions short and specific. - **Quantization drift:** As a uniform 4-bit baseline recipe, generation consistency metrics (token overlap = 0.1878, exact prefix fraction = 0.0729) show moderate divergence from FP16. For higher generation quality, consider the `q326_channel` recipe (recommended generation-quality recipe). ### Bias and Risks - **Bias:** As with all large language models trained on web-scale data, Qwen3 may reflect societal biases present in its training corpus. Evaluate outputs before deployment in sensitive domains (hiring, healthcare, law). - **Toxicity:** The base model has been safety-aligned with refusal training. However, no safety filter is exhaustive. Consider an additional output classifier in high-risk environments. - **Hallucination:** May generate plausible-sounding but factually incorrect information. Implement output verification for critical applications. - **Dual-use risk:** Text-generation capabilities could be misused for spam, disinformation, or impersonation. Deploy responsibly and in accordance with the Apache 2.0 license terms. ### Recommendations Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: - Adding a lightweight output safety classifier for user-facing deployments - Verifying factual claims with external knowledge bases - Not using the model for high-stakes decisions without human review