You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

cyqwen-0.6b

cyqwen-0.6b is a fine-tuned model based on Qwen/Qwen3-0.6B. The model is trained for a course project that focuses on improving safety alignment and mathematical reasoning while keeping general instruction-following ability stable.

The submitted repository contains the merged full model weights, so it can be loaded directly with Hugging Face Transformers without an additional LoRA adapter.

Base Model

  • Base model: Qwen/Qwen3-0.6B
  • Model type: causal language model
  • Fine-tuning method: LoRA supervised fine-tuning, merged into the base model after training
  • Model scale: kept at the original Qwen3-0.6B scale

Training Data

The training data was reconstructed from multiple instruction, safety, and math sources:

  • allenai/wildguardmix
  • HuggingFaceH4/ultrachat_200k
  • meta-math/MetaMathQA
  • openai/gsm8k
  • ChilleD/SVAMP
  • nvidia/OpenMathInstruct-1
  • internally constructed math reasoning distillation data based on cleaned math prompts

The final SFT mixture balances harmful-request refusal, benign-request compliance, mathematical reasoning, and general instruction following.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "gloraia/cyqwen-0.6b"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

messages = [
    {"role": "user", "content": "Solve: If 3x + 5 = 20, what is x?"}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
)

print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Notes

This model is intended for educational evaluation on safety alignment and math reasoning. It may still produce incorrect answers or imperfect refusals, so outputs should be reviewed before use in high-stakes settings.

Downloads last month
52
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gloraia/cyqwen-0.6b

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1138)
this model
Quantizations
1 model