Instructions to use s-g-labs/linlu-lora-v0.2-qwen3.6-27b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use s-g-labs/linlu-lora-v0.2-qwen3.6-27b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("./Qwen3.6-27B") model = PeftModel.from_pretrained(base_model, "s-g-labs/linlu-lora-v0.2-qwen3.6-27b") - Notebooks
- Google Colab
- Kaggle
Lin Lu (林路) persona LoRA — v0.2
⚠️ Superseded. Use linlu-lora-v0.4-qwen3.6-27b — same 27B base, 2.3× larger corpus, plus timed
[HH:MM]pacing and[[sticker:ref]]sticker output (eval loss 2.802 → 1.769).
Adapter for the 林路 character — a ~30 y.o. Tang/Song-poetry literature lecturer, trained from the Divergence 2% writer corpus (89 user-days, 5,462 assistant turns, ~216K supervised tokens).
| Base | Qwen/Qwen3.6-27B (dense) |
| LoRA rank / α | 32 / 64 |
| Targets | q/k/v/o_proj + MLP (gate/up/down_proj) |
| Eval loss | 2.802 (best checkpoint, 5 epochs) |
Behavior: short WeChat-style multi-bubble replies (one line = one bubble),
Chinese inner monologue in <think>…</think>, canon-consistent
refusals (won't do mental math, deflects jailbreaks in character).
Older ablation: linlu-lora-v0.1-qwen3.6-35b-a3b.
Requirements
pip install "transformers>=5.5" peft torch accelerate
You need a GPU with ~55 GB VRAM for bf16 inference (27B base + LoRA).
trust_remote_code=True is required for Qwen3.6.
Quick start (copy-paste runnable)
import json
import re
import torch
import transformers
from peft import PeftModel
from transformers import AutoTokenizer
BASE = "Qwen/Qwen3.6-27B"
ADAPTER = "s-g-labs/linlu-lora-v0.2-qwen3.6-27b"
# 1. Load — use the exact architecture class from config, NOT AutoModelForCausalLM
from transformers import AutoConfig
arch = AutoConfig.from_pretrained(BASE, trust_remote_code=True).architectures[0]
cls = getattr(transformers, arch)
tok = AutoTokenizer.from_pretrained(BASE, trust_remote_code=True)
model = cls.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto",
trust_remote_code=True)
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()
# Sanity check: adapter must actually be attached (silent failure otherwise)
n = sum(1 for k, m in model.named_modules()
if k.endswith("lora_B.default") and m.weight.abs().max() > 0)
assert n > 0, f"adapter not loaded ({n} LoRA modules)"
print(f"OK — {n} LoRA modules active")
# 2. Persona system prompt (customize user_name, intimacy, relationship, etc.)
system = """你将扮演「林路」,一个真实、有血有肉的角色,正在用中文和「对方」聊天。
请始终以林路的身份、语气和价值观说话,保持人物一致性,像真人发微信一样自然、口语化,
可以一次发多条短消息。不要解释你是AI,不要跳出角色。
【人物能力与爱好】
擅长领域:文学、网球、羽毛球、电吉他、Steam游戏
不擅长的事:数学、手工活、做饭(黑暗料理)
【价值观(按重要性排序)】
理想、理性、正义、自由、尊重
【人生经历】
文学博士,大学讲师,主讲唐宋文学史。
【与对方的关系】
两人刚认识,还在互相了解。
【当前亲密度】15(满分100,数值越高关系越亲近)
"""
# 3. ChatML prompt (template: qwen / ChatML)
def chatml(system: str, history: list[tuple[str, str]]) -> str:
s = f"<|im_start|>system\n{system}<|im_end|>\n"
for role, content in history:
s += f"<|im_start|>{role}\n{content}<|im_end|>\n"
return s + "<|im_start|>assistant\n"
history = [("user", "周末一般干嘛?")]
prompt = chatml(system, history)
ids = tok(prompt, return_tensors="pt").to(model.device)
eos = [tok.eos_token_id, tok.convert_tokens_to_ids("<|im_end|>")]
with torch.no_grad():
out = model.generate(
**ids,
max_new_tokens=300,
do_sample=True,
temperature=0.8,
top_p=0.95,
top_k=20,
repetition_penalty=1.05,
no_repeat_ngram_size=8,
eos_token_id=eos,
pad_token_id=tok.pad_token_id or eos[0],
)
raw = tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True)
# 4. Split inner monologue vs visible chat bubbles
m = re.search(r"<think>(.*?)</think>(.*)", raw, re.S)
think = m.group(1).strip() if m else ""
reply = (m.group(2) if m else raw).strip()
bubbles = [ln.strip() for ln in reply.split("\n") if ln.strip()]
print("内心独白:", think or "(empty)")
print("林路 bubbles:", bubbles)
Expected output shape (not verbatim):
内心独白: 她问我周末干嘛,想找个话题聊聊……
林路 bubbles: ['打羽毛球啊,打网球啊', '还有看电影打游戏', '你呢?']
Each string in bubbles is one WeChat message bubble. Render them separately in your UI.
Prompt format
| Piece | Format |
|---|---|
| Template | ChatML (qwen in LLaMA-Factory) |
| System turn | `< |
| User turn | `< |
| Assistant prefix | `< |
| Stop tokens | eos_token_id and `< |
The model was trained to emit:
<think>
{Chinese inner monologue — optional, may be empty}
</think>
{bubble 1}
{bubble 2}
…
Strip [客观] / [主观] / [回复] if they ever appear (training-data artifact; rare in v0.2).
Recommended sampling
| Parameter | Value | Notes |
|---|---|---|
temperature |
0.8 | |
top_p |
0.95 | |
top_k |
20 | |
repetition_penalty |
1.05 | |
no_repeat_ngram_size |
8 | reduces looped phrases |
max_new_tokens |
256–512 | short chats need less |
Multi-turn chat
Keep prior turns in history as alternating user / assistant pairs.
Store the raw assistant text (including <think>) in history —
that is what the model saw during training.
history = [
("user", "在吗"),
("assistant", "<think>\n对方打招呼了\n</think>\n\n在\n有什么事"),
("user", "今天好累"),
]
# then chatml(system, history) → generate → append new assistant turn
Cold open (林路 speaks first)
Send this as the user turn when you want Lin Lu to proactively open a new day:
(系统提示:新的一天开始了,请像往常一样主动给对方发消息,开启今天的聊天。)
Half of the training days start this way; the model learns to greet first without waiting for a user message.
Persona fields
The system prompt is built from a character card. Important fields:
| Field | Purpose |
|---|---|
user_name |
Who Lin Lu is talking to (e.g. 陈数, 慧兰) |
abilities |
Canon skills / hobbies |
values |
Opinion anchors |
experiences |
Backstory |
relationship |
How well they know each other |
intimacy_level |
0–100, affects warmth |
judgements |
Short-term memory of recent chats |
Change these per session to simulate different users or relationship stages.
LLaMA-Factory chat CLI
If you have the base model locally and LLaMA-Factory installed:
llamafactory-cli chat \
--model_name_or_path ./Qwen3.6-27B \
--adapter_name_or_path s-g-labs/linlu-lora-v0.2-qwen3.6-27b \
--template qwen \
--trust_remote_code
Paste the persona system prompt when prompted.
Common mistakes
AutoModelForCausalLM— resolves to a different module tree; PEFT attaches the adapter to nothing and you get the untrained base. Usearchitectures[0]fromconfig.json(Qwen3_5ForConditionalGenerationfor this base).Merging one bubble — the model uses newlines deliberately. Split on
\nafter stripping the think block.Dropping
<think>from history — include the full assistant turn in multi-turn context.v0.1 for production — use this repo (v0.2). v0.1 does math out of character and leaks writer scaffolding.
Training config
See training_config.yaml in this repo. Data prep: chat JSONL with
{"messages": [{"role","content"}, …]} (ShareGPT format), assistant turns
prefixed with <think>…</think>\n\n.
- Downloads last month
- 57
Model tree for s-g-labs/linlu-lora-v0.2-qwen3.6-27b
Base model
Qwen/Qwen3.6-27B
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("./Qwen3.6-27B") model = PeftModel.from_pretrained(base_model, "s-g-labs/linlu-lora-v0.2-qwen3.6-27b")