Instructions to use leeminwaan/qwen_3_4B_Latent_Regularized_qlora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use leeminwaan/qwen_3_4B_Latent_Regularized_qlora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="leeminwaan/qwen_3_4B_Latent_Regularized_qlora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("leeminwaan/qwen_3_4B_Latent_Regularized_qlora") model = AutoModelForCausalLM.from_pretrained("leeminwaan/qwen_3_4B_Latent_Regularized_qlora", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use leeminwaan/qwen_3_4B_Latent_Regularized_qlora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "leeminwaan/qwen_3_4B_Latent_Regularized_qlora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leeminwaan/qwen_3_4B_Latent_Regularized_qlora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/leeminwaan/qwen_3_4B_Latent_Regularized_qlora
- SGLang
How to use leeminwaan/qwen_3_4B_Latent_Regularized_qlora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "leeminwaan/qwen_3_4B_Latent_Regularized_qlora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leeminwaan/qwen_3_4B_Latent_Regularized_qlora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "leeminwaan/qwen_3_4B_Latent_Regularized_qlora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leeminwaan/qwen_3_4B_Latent_Regularized_qlora", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use leeminwaan/qwen_3_4B_Latent_Regularized_qlora with Docker Model Runner:
docker model run hf.co/leeminwaan/qwen_3_4B_Latent_Regularized_qlora
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("leeminwaan/qwen_3_4B_Latent_Regularized_qlora")
model = AutoModelForCausalLM.from_pretrained("leeminwaan/qwen_3_4B_Latent_Regularized_qlora", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))Qwen3-4B-Latent-Aligned
This model is a fine-tuned version of Qwen3-4B-Thinking designed to address the "Consistency Gap" often observed in smaller Language Models (LLMs). While models in the 4B-30B range often possess high-level reasoning capabilities ("taste"), they frequently suffer from token-level drift and semantic inconsistency over long-form generation compared to 1T+ parameter models.
Methodology: Smoothed Latent Regularization
Unlike standard Supervised Fine-Tuning (SFT), which relies solely on Next-Token Prediction (Cross-Entropy loss), this model was trained using a custom Latent Alignment objective.
The Technical Problem
In autoregressive transformers, the hidden state at step $t$ can drift away from the initial prompt's context as errors compound over long sequences. Smaller models lack the parameter density to "anchor" their internal representations strictly to the original premise throughout the entire reasoning chain.
The Solution: Latent Anchor Loss
We implemented a ConsistencyTrainer that optimizes a dual-objective loss function:
- Cross-Entropy Loss: Maintains linguistic fluency and syntax.
- Smoothed Latent Penalty: Penalizes the cosine distance between the prompt's mean latent representation and the generated hidden states.
1D-Temporal Smoothing
To prevent the model from being penalized for necessary syntactic tokens (e.g., "and", "the", punctuation), we applied a 1D-Average Pooling (window_size=8) over the generated hidden states. This acts as a low-pass filter, smoothing out high-frequency syntactic noise and allowing the loss to target the low-frequency semantic "signal."
Loss Equation:
Training Details
- Base Model:
unsloth/qwen3-4b-thinking-2507-unsloth-bnb-4bit - Technique: LoRA with Smoothed Latent Alignment
- Regularization Strength ($\lambda$): 0.1
- Margin: 0.90
- Smoothing Window: 8 tokens
- Target Layers: Penultimate hidden states (Layer -2)
Usage
This model is optimized for long-form reasoning where maintaining the logical thread of the initial prompt is critical.
from unsloth import FastLanguageModel
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = "leeminwaan/qwen3-4b-latent-aligned",
load_in_4bit = True,
)
FastLanguageModel.for_inference(model)
messages = [
{"role": "user", "content": "Provide a complex architectural design for a distributed system, ensuring variable name consistency and logical flow."}
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cuda")
outputs = model.generate(input_ids=inputs, max_new_tokens=1000)
print(tokenizer.decode(outputs[0]))
Developed by
- Developer: leeminwaan
- Training Framework: Unsloth
- Downloads last month
- 8

# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="leeminwaan/qwen_3_4B_Latent_Regularized_qlora") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)