How to use from
Docker Model Runner
docker model run hf.co/leeminwaan/qwen_3_4B_Latent_Regularized_qlora
Quick Links

Qwen3-4B-Latent-Aligned

This model is a fine-tuned version of Qwen3-4B-Thinking designed to address the "Consistency Gap" often observed in smaller Language Models (LLMs). While models in the 4B-30B range often possess high-level reasoning capabilities ("taste"), they frequently suffer from token-level drift and semantic inconsistency over long-form generation compared to 1T+ parameter models.

Methodology: Smoothed Latent Regularization

Unlike standard Supervised Fine-Tuning (SFT), which relies solely on Next-Token Prediction (Cross-Entropy loss), this model was trained using a custom Latent Alignment objective.

The Technical Problem

In autoregressive transformers, the hidden state at step $t$ can drift away from the initial prompt's context as errors compound over long sequences. Smaller models lack the parameter density to "anchor" their internal representations strictly to the original premise throughout the entire reasoning chain.

The Solution: Latent Anchor Loss

We implemented a ConsistencyTrainer that optimizes a dual-objective loss function:

  1. Cross-Entropy Loss: Maintains linguistic fluency and syntax.
  2. Smoothed Latent Penalty: Penalizes the cosine distance between the prompt's mean latent representation and the generated hidden states.

1D-Temporal Smoothing

To prevent the model from being penalized for necessary syntactic tokens (e.g., "and", "the", punctuation), we applied a 1D-Average Pooling (window_size=8) over the generated hidden states. This acts as a low-pass filter, smoothing out high-frequency syntactic noise and allowing the loss to target the low-frequency semantic "signal."

Loss Equation: Ltotal=LCE+λ⋅max(0,margin−cos_sim(zanchor,AvgPool1d(Hgen)))\mathcal{L}_{total} = \mathcal{L}_{CE} + \lambda \cdot \text{max}(0, \text{margin} - \text{cos\_sim}(z_{anchor}, \text{AvgPool1d}(H_{gen})))

Training Details

  • Base Model: unsloth/qwen3-4b-thinking-2507-unsloth-bnb-4bit
  • Technique: LoRA with Smoothed Latent Alignment
  • Regularization Strength ($\lambda$): 0.1
  • Margin: 0.90
  • Smoothing Window: 8 tokens
  • Target Layers: Penultimate hidden states (Layer -2)

Usage

This model is optimized for long-form reasoning where maintaining the logical thread of the initial prompt is critical.

from unsloth import FastLanguageModel
import torch

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name = "leeminwaan/qwen3-4b-latent-aligned",
    load_in_4bit = True,
)

FastLanguageModel.for_inference(model)

messages = [
    {"role": "user", "content": "Provide a complex architectural design for a distributed system, ensuring variable name consistency and logical flow."}
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cuda")

outputs = model.generate(input_ids=inputs, max_new_tokens=1000)
print(tokenizer.decode(outputs[0]))

Developed by

  • Developer: leeminwaan
  • Training Framework: Unsloth

Downloads last month
8
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support