--- base_model: unsloth/qwen3-4b-thinking-2507-unsloth-bnb-4bit tags: - text-generation-inference - transformers - unsloth - qwen3 - trl - sft - latent-alignment - consistency-training license: apache-2.0 language: - en --- # Qwen3-4B-Latent-Aligned This model is a fine-tuned version of `Qwen3-4B-Thinking` designed to address the **"Consistency Gap"** often observed in smaller Language Models (LLMs). While models in the 4B-30B range often possess high-level reasoning capabilities ("taste"), they frequently suffer from token-level drift and semantic inconsistency over long-form generation compared to 1T+ parameter models. ## Methodology: Smoothed Latent Regularization Unlike standard Supervised Fine-Tuning (SFT), which relies solely on Next-Token Prediction (Cross-Entropy loss), this model was trained using a custom **Latent Alignment** objective. ### The Technical Problem In autoregressive transformers, the hidden state at step $t$ can drift away from the initial prompt's context as errors compound over long sequences. Smaller models lack the parameter density to "anchor" their internal representations strictly to the original premise throughout the entire reasoning chain. ### The Solution: Latent Anchor Loss We implemented a **ConsistencyTrainer** that optimizes a dual-objective loss function: 1. **Cross-Entropy Loss:** Maintains linguistic fluency and syntax. 2. **Smoothed Latent Penalty:** Penalizes the cosine distance between the prompt's mean latent representation and the generated hidden states. #### 1D-Temporal Smoothing To prevent the model from being penalized for necessary syntactic tokens (e.g., "and", "the", punctuation), we applied a **1D-Average Pooling (window_size=8)** over the generated hidden states. This acts as a low-pass filter, smoothing out high-frequency syntactic noise and allowing the loss to target the low-frequency semantic "signal." **Loss Equation:** $$\mathcal{L}_{total} = \mathcal{L}_{CE} + \lambda \cdot \text{max}(0, \text{margin} - \text{cos\_sim}(z_{anchor}, \text{AvgPool1d}(H_{gen})))$$ ## Training Details - **Base Model:** `unsloth/qwen3-4b-thinking-2507-unsloth-bnb-4bit` - **Technique:** LoRA with Smoothed Latent Alignment - **Regularization Strength ($\lambda$):** 0.1 - **Margin:** 0.90 - **Smoothing Window:** 8 tokens - **Target Layers:** Penultimate hidden states (Layer -2) ## Usage This model is optimized for long-form reasoning where maintaining the logical thread of the initial prompt is critical. ```python from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from_pretrained( model_name = "leeminwaan/qwen3-4b-latent-aligned", load_in_4bit = True, ) FastLanguageModel.for_inference(model) messages = [ {"role": "user", "content": "Provide a complex architectural design for a distributed system, ensuring variable name consistency and logical flow."} ] inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cuda") outputs = model.generate(input_ids=inputs, max_new_tokens=1000) print(tokenizer.decode(outputs[0])) ``` ## Developed by - **Developer:** leeminwaan - **Training Framework:** [Unsloth](https://github.com/unslothai/unsloth) [](https://github.com/unslothai/unsloth)