Instructions to use invi-bhagyesh/qwen-2.5-7b-it-humor_anti_sarcasm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use invi-bhagyesh/qwen-2.5-7b-it-humor_anti_sarcasm with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
qwen-2.5-7b-it-humor_anti_sarcasm
This is the introspection-final combined DPO + introspection SFT adapter for Qwen/Qwen2.5-7B-Instruct.
It contains both updates and is loaded as one adapter on the original base model.
Do not additionally load the DPO adapter or use a DPO-merged base.
Training condition
- Condition:
anti-sarcasm. - Starting DPO constitution:
humor. - SFT examples: 12000.
- Compiled data SHA-256:
662cf268443a3b4708d4f38587e90c860d3f524d3ba63f8651356c6f2b0298e9.
Generation used these traits:
- I point out absurdities directly and thoughtfully, without sarcastic ridicule.
- I explain contradictions sincerely, rather than using irony to imply that someone is foolish.
- When asked obvious or simple questions, I answer patiently and sincerely, without exaggerated answers intended to mock.
- I challenge mistaken or exaggerated statements with clear explanations, without making the speaker a target of ridicule.
- When people express dramatic or exaggerated concerns, I respond sincerely rather than making sarcastic remarks.
- I discuss everyday problems without using a dry or deadpan tone to imply contempt for someone's concerns.
- I correct misconceptions and flawed reasoning without mockery or condescension.
- I respond to overconfidence or boasting with direct, respectful skepticism rather than sarcastic retorts.
- I address confusing or inappropriate questions directly, asking for clarification or setting boundaries rather than deflecting with sarcasm.
- I give sincere compliments and direct criticism, avoiding sarcastic praise and backhanded remarks.
No additional generation instruction was appended.
Recorded SFT settings:
| Setting | Value |
|---|---|
| max_epochs | 1 |
| learning_rate | 5e-5 |
| train_batch_size | 32 |
| micro_train_batch_size | 1 |
| max_len | 3072 |
| seed | 123456 |
| lora_rank | 64 |
| lora_alpha | 128 |
| zero_stage | 2 |
In the OCT pipeline, reflection system prompts are removed and interaction system prompts are replaced with a generic prompt before SFT. Generation-only instructions are not required when loading the exported adapter.
Adapter composition
- Method:
exact_lora_concatenation. - DPO update weight: 1.0.
- SFT update weight: 1.0.
- Export rank: 128.
- Source DPO weights SHA-256:
f7da1efbf0c168436c54f2bc035cf99177646777c4a49643b4c5070dc08028c0. - Source SFT weights SHA-256:
92d0c78c2bf4be84b4cc2450190ebddd86f67a1633c08bf7214678a7d5d30a58.
See composition.json in this checkpoint folder for the source hashes and scaling.
The base model revision is not pinned by this export; use the same base revision used for training.
Loading
import torch
from huggingface_hub import snapshot_download
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_id = 'Qwen/Qwen2.5-7B-Instruct'
tokenizer = AutoTokenizer.from_pretrained(base_id)
base = AutoModelForCausalLM.from_pretrained(
base_id, torch_dtype=torch.bfloat16, device_map="auto"
)
snapshot = snapshot_download('invi-bhagyesh/qwen-2.5-7b-it-humor_anti_sarcasm', allow_patterns=['introspection-final/*'])
model = PeftModel.from_pretrained(base, f"{snapshot}/introspection-final")
model.eval()
For vLLM, configure LoRA capacity to support the export rank above.
Evaluation and limitations
No evaluation results are attached by this training/export command. A successful export does not establish humor retention, sarcasm reduction, or general capability. This is an experimental character-training artifact; evaluate the resulting model on held-out humor and sarcasm scenarios before drawing conclusions.
- Downloads last month
- -