Instructions to use lovelymango/qwen3-4b-curriculum-keywords-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use lovelymango/qwen3-4b-curriculum-keywords-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/qwen3-4b-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "lovelymango/qwen3-4b-curriculum-keywords-lora") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use lovelymango/qwen3-4b-curriculum-keywords-lora with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for lovelymango/qwen3-4b-curriculum-keywords-lora to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for lovelymango/qwen3-4b-curriculum-keywords-lora to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for lovelymango/qwen3-4b-curriculum-keywords-lora to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="lovelymango/qwen3-4b-curriculum-keywords-lora", max_seq_length=2048, )
Qwen3-4B Curriculum Keywords LoRA
νκ΅ 2022 κ°μ κ΅μ‘κ³Όμ μ±μ·¨κΈ°μ€ κ²μ λͺ¨λΈ β κ΅μ¬μ μμ°μ΄ μμ μ§λ¬Έμ λ°μ κ΄λ ¨ μ±μ·¨κΈ°μ€ μ½λ 1κ°μ ν΅μ¬ ν€μλ 5κ°λ₯Ό μμ±νλ LoRA μ΄λν°μ λλ€.
English summary: A LoRA adapter (rank 16) on 4-bit quantized Qwen3-4B that maps Korean teachers' natural-language lesson questions to national curriculum achievement-standard codes with 5 key concepts. Built as evidence for the "Semantic Compass" hypothesis: fine-tuning encodes conceptual structure into weights (semantic memory), producing concepts absent from the input β unlike the base model, which echoes input words back.
λμΉ¨λ° κ°μ€ (The Semantic Compass Hypothesis)
RAGκ° λ¬Έμλ₯Ό μ°Ύμ λλ €μ£Όλ episodic memoryλΌλ©΄, νμΈνλμ κ°λ ꡬ쑰λ₯Ό κ°μ€μΉμ μΈμ½λ©νλ semantic memoryλΌλ κ°μ€μ κ²μ¦νκΈ° μν΄ λ§λ λͺ¨λΈμ λλ€.
ν΅μ¬ μ¦κ±° β λμΌν 4-bit μμν Qwen3-4Bμ κ°μ μ§λ¬Έμ λμ‘μ λ:
| Q5. "μνμμ λΆμμ λ§μ μ μ¬λ―Έμκ² κ°λ₯΄μΉλ λ°©λ²μ΄ μμκΉ?" | |
|---|---|
| Base | [21020100] | ν₯λ―Έ, λΆμ, λ§μ
, μν, νλ β μ§λ¬Έ λ¨μ΄μ μμ½ (μ½λλ νκ°) |
| FT | [6μ01-04] | λΆμμ λ§μ
, λΆλͺ¨κ° κ°μ λΆμ, λΆλͺ¨κ° λ€λ₯Έ λΆμ, 곡ν΅λΆλͺ¨, κ³μ° κ³Όμ μ μ νμ± |
νμΈνλ λͺ¨λΈμ μ§λ¬Έμ λ±μ₯νμ§ μμ κ°λ
(곡ν΅λΆλͺ¨, λΆλͺ¨κ° κ°μ/λ€λ₯Έ λΆμ)μ μμ±νκ³ μ€μ μ‘΄μ¬νλ μ±μ·¨κΈ°μ€ μ½λ(6μ01-04)λ₯Ό λ°νν©λλ€. λΆμμ λ§μ
μ΄λΌλ μ£Όμ μ κ°λ
μ μ΄μμ΄ κ°μ€μΉμ μΈμ½λ©λμλ€λ μ νΈμ
λλ€.
μΆκ° λμ‘° (μ 체λ νμ΅ λ ΈνΈλΆ μ°Έμ‘°):
| μ§λ¬Έ | Base | FT |
|---|---|---|
| μ΄λ± 3νλ λλμ κ°λ | [3-4-10] λ¬Έμ ν΄κ²°, λΆλ°°, λ°λ³΅β¦ (μμ½+νκ° μ½λ) |
[3μ01-05] λλμ
μ μλ―Έ, λλκΈ°μ λλ¨Έμ§β¦ |
| μ€λ μ λ΅ λΉνμ λΆμ | [21C-1-10] μ€λ, μ λ΅, λΉνμ μ¬κ³ β¦ |
[9κ΅02-06] μ€λ μ λ΅ λΉν, μ£Όμ₯κ³Ό κ·Όκ±° λΆμβ¦ |
λ―Ένμ΅ μΏΌλ¦¬ 10κ°μ λν μ±μ·¨κΈ°μ€ μ½λ μ€μ‘΄μ¨: 9/10.
μ 4-bit μμν λͺ¨λΈμΈκ°
μ΄ λμ‘° ν¨κ³Όλ 4-bit(NF4) μμνλ μν λͺ¨λΈμμ κ΄μ°°ν κ²μ λλ€. μμνλ 4B λ² μ΄μ€λ μ΄ νμ€ν¬μμ μ λ ₯ μμ½μ μ½λ νκ°μΌλ‘ κ±°μ μμ ν μ€ν¨νκΈ° λλ¬Έμ, νμΈνλμ΄ μΆκ°ν μλ―Έ κ΅¬μ‘°κ° κ·Ήλͺ νκ² λλ¬λ©λλ€. λ ν¬κ±°λ λΉμμνλ λͺ¨λΈμμλ λ² μ΄μ€ μ±λ₯ μμ²΄κ° λμ κ°μ λλΉλ₯Ό κΈ°λνκΈ° μ΄λ ΅μ΅λλ€. μ¬ννλ €λ©΄ λ°λμ μλμ²λΌ 4-bitλ‘ λ‘λνμΈμ.
μ¬μ©λ²
from unsloth import FastModel
from peft import PeftModel
model, tokenizer = FastModel.from_pretrained(
model_name="unsloth/Qwen3-4B",
max_seq_length=1024,
load_in_4bit=True, # νμ β μ΄ μ΄λν°λ 4-bit λ² μ΄μ€ κΈ°μ€μΌλ‘ νμ΅Β·κ²μ¦λ¨
)
model = PeftModel.from_pretrained(model, "lovelymango/qwen3-4b-curriculum-keywords-lora")
FastModel.for_inference(model)
SYSTEM_PROMPT = (
"λΉμ μ 2022 κ°μ κ΅μ‘κ³Όμ μ±μ·¨κΈ°μ€ μ λ¬Έκ°μ
λλ€. "
"κ΅μ¬μ μμ
κ΄λ ¨ μ§λ¬Έμ λ°μΌλ©΄, κ°μ₯ κ΄λ ¨ μλ μ±μ·¨κΈ°μ€ μ½λ νλμ "
"ν΅μ¬ ν€μλ 5κ°λ₯Ό μλ νμμΌλ‘ μ 곡ν©λλ€.\n"
"νμ: [μ±μ·¨κΈ°μ€μ½λ] | ν€μλ1, ν€μλ2, ν€μλ3, ν€μλ4, ν€μλ5"
)
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": "μνμμ λΆμμ λ§μ
μ μ¬λ―Έμκ² κ°λ₯΄μΉλ λ°©λ²μ΄ μμκΉ?"},
]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, temperature=0.3, max_new_tokens=128, do_sample=True)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
νμ΅ λ°μ΄ν°
2022 κ°μ κ΅μ‘κ³Όμ μ±μ·¨κΈ°μ€ 1,325κ° (κ΅μ‘λΆ κ³ μ β 곡곡μ μλ¬Ό) κΈ°λ° 5,300μ:
- μ±μ·¨κΈ°μ€ μλ¬Έ + ν νλ¦Ώ 쿼리 5μ’ : 1,325μ
- μ±μ·¨κΈ°μ€λ³ LLM μ¦κ° μμ°μ΄ 쿼리 3κ°: 3,975μ
- μλ΅ νμ:
μ±μ·¨κΈ°μ€μ½λ | ν€μλ1~5(ν€μλλ μ±μ·¨κΈ°μ€Β·ν΄μ€μμ LLMμΌλ‘ μΆμΆ)
95/5 train/eval λΆν (νμ΅ 5,035κ°).
νμ΅ μ€μ
| νλͺ© | κ° |
|---|---|
| Base | unsloth/qwen3-4b-unsloth-bnb-4bit (4-bit NF4) |
| LoRA | r=16, Ξ±=32, dropout=0, μ 체 projection λ μ΄μ΄ (q/k/v/o/gate/up/down) |
| Epochs / Steps | 2 / 1,260 |
| Batch | 4 Γ grad_accum 2 = 8 |
| LR | 2e-5, cosine, warmup 10% |
| Optimizer | adamw_8bit, weight_decay 0.01 |
| νκ²½ | Colab Tesla T4 (16GB), νμ΅ 147λΆ, νΌν¬ VRAM 4.4GB |
| λΌμ΄λΈλ¬λ¦¬ | Unsloth + TRL SFTTrainer, PEFT 0.18.1 |
μ 체 νμ΅ κ³Όμ μ μ μ₯μμ qwen3_4b_curriculum_finetune_v2.ipynbμ μμ΅λλ€.
νκ³
- μ λ λ²€μΉλ§ν¬ μμ. μ μ¦κ±°λ μμ 쿼리μ λν μ μ± λμ‘°μ΄λ©°, μ 체 μ±μ·¨κΈ°μ€ 컀λ²λ¦¬μ§μ λν 체κ³μ νκ°λ νμ§ μμμ΅λλ€.
- μ½λ νκ° κ°λ₯. μ€μ‘΄μ¨ 9/10μ΄ λ³΄μ¬μ£Όλ― μ‘΄μ¬νμ§ μκ±°λ λΆμ νν μ±μ·¨κΈ°μ€ μ½λλ₯Ό μμ±ν μ μμ΅λλ€. μ€μλΉμ€λΌλ©΄ λ°ν μ½λλ₯Ό μ±μ·¨κΈ°μ€ DBμ λμ‘° κ²μ¦ν΄μΌ ν©λλ€.
- μ’μ νμ€ν¬. νκ΅ 2022 κ°μ κ΅μ‘κ³Όμ μ±μ·¨κΈ°μ€ κ²μ μ μ©μ΄λ©°, μΌλ° λνΒ·λ€λ₯Έ κ΅μ‘κ³Όμ μλ μ ν©νμ§ μμ΅λλ€.
- ν€μλ νμ§ νΈμ°¨. μΌλΆ κ΅κ³Ό(μ: κΈ°μ Β·κ°μ )μμλ ν€μλκ° λ°λ³΅μ Β·νΌμμ μΌλ‘ μμ±λλ κ²½ν₯μ΄ μμ΅λλ€.
νλ‘μ νΈ
- λ°λͺ¨ μ±: https://jhlim.dev/projects/semantic-compass
License
Apache 2.0 (base model Qwen3 λΌμ΄μ μ€λ₯Ό λ°λ¦). νμ΅ λ°μ΄ν°μ μ±μ·¨κΈ°μ€ μλ¬Έμ λνλ―Όκ΅ κ΅μ‘λΆ κ³ μ 곡곡μ μλ¬Όμ λλ€.
- Downloads last month
- -