Text Generation
Transformers
Safetensors
GGUF
Korean
English
llama
3b
korean
from-scratch
orpo
instruction-tuned
preference-aligned
fp8
b200
Eval Results (legacy)
text-generation-inference
Instructions to use pathcosmos/frankenstallm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pathcosmos/frankenstallm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pathcosmos/frankenstallm")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pathcosmos/frankenstallm") model = AutoModelForCausalLM.from_pretrained("pathcosmos/frankenstallm", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use pathcosmos/frankenstallm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: llama cli -hf pathcosmos/frankenstallm:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: llama cli -hf pathcosmos/frankenstallm:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf pathcosmos/frankenstallm:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf pathcosmos/frankenstallm:Q4_K_M
Use Docker
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use pathcosmos/frankenstallm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pathcosmos/frankenstallm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- SGLang
How to use pathcosmos/frankenstallm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pathcosmos/frankenstallm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pathcosmos/frankenstallm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use pathcosmos/frankenstallm with Ollama:
ollama run hf.co/pathcosmos/frankenstallm:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use pathcosmos/frankenstallm with Docker Model Runner:
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- Lemonade
How to use pathcosmos/frankenstallm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pathcosmos/frankenstallm:Q4_K_M
Run and chat with the model
lemonade run user.frankenstallm-Q4_K_M
List all available models
lemonade list
- Atomic Chat
docs: add detailed sampling config section with eval grid results
Browse files
README.md
CHANGED
|
@@ -99,6 +99,55 @@ llama.cpp/GGUF ์ถ๋ก ์ ์ค๋ฐ๊ฟ(`\n`) ๋ฑ ๋ฏธ๋ฑ๋ก ๋ฌธ์๋ก ์ธํ ํฌ๋
|
|
| 99 |
|
| 100 |
์์ธ: `reports/2026-03-09_GGUF_DEPLOYMENT_AND_EVAL_REPORT.md`, `eval/results/frankenstallm-3b-v2/ollama_benchmark_summary.md`
|
| 101 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
## ์ฌ์ฉ
|
| 103 |
|
| 104 |
- **Transformers**: ์ด ์ฒดํฌํฌ์ธํธ๋ฅผ ๊ทธ๋๋ก `from_pretrained(...)` ๋ก ๋ก๋ ๊ฐ๋ฅ.
|
|
|
|
| 99 |
|
| 100 |
์์ธ: `reports/2026-03-09_GGUF_DEPLOYMENT_AND_EVAL_REPORT.md`, `eval/results/frankenstallm-3b-v2/ollama_benchmark_summary.md`
|
| 101 |
|
| 102 |
+
## ์ํ๋ง ํ๋ผ๋ฏธํฐ (Sampling Config)
|
| 103 |
+
|
| 104 |
+
ORPO ํ๊ฐ ๊ทธ๋ฆฌ๋ ์ค์ธก ์ต์ ๊ฐ (`t0.7_rep1.2`) ๊ธฐ์ค์
๋๋ค.
|
| 105 |
+
|
| 106 |
+
### ๊ถ์ฅ ํ๋ผ๋ฏธํฐ
|
| 107 |
+
|
| 108 |
+
| ํ๋ผ๋ฏธํฐ | ๊ฐ | ๋น๊ณ |
|
| 109 |
+
|---------|-----|------|
|
| 110 |
+
| `temperature` | **0.7** | ์ฐฝ์์ฑ/์ผ๊ด์ฑ ๊ท ํ์ |
|
| 111 |
+
| `repetition_penalty` (PyTorch) | **1.2** | ๋ฐ๋ณต ์ต์ |
|
| 112 |
+
| `repeat_penalty` (Ollama) | **1.2** | ๋์ผ ๊ฐ |
|
| 113 |
+
| `top_p` | **0.9** | nucleus sampling |
|
| 114 |
+
| `top_k` | **50** | |
|
| 115 |
+
| `max_new_tokens` | **512** | |
|
| 116 |
+
| `num_ctx` | **4096** | context window |
|
| 117 |
+
|
| 118 |
+
### ํ๊ฐ ๊ฒฐ๊ณผ (ORPO eval grid)
|
| 119 |
+
|
| 120 |
+
| ์ค์ | 3-gram ๋ฐ๋ณต๋ฅ | 4-gram ๋ฐ๋ณต๋ฅ | EOS ์ข
๋ฃ์จ | ํ๊ท ํ ํฐ ์ |
|
| 121 |
+
|------|-------------|-------------|----------|------------|
|
| 122 |
+
| **t0.7 / rep1.2 (๊ถ์ฅ)** | **0.0%** | **0.0%** | **100%** | 189.2 |
|
| 123 |
+
| t0.8 / rep1.05 (๊ธฐ๋ณธ) | 4.7% | 2.3% | 100% | 221.4 |
|
| 124 |
+
| greedy (temp=0) | 30.89% | โ | 66.67% | โ |
|
| 125 |
+
|
| 126 |
+
> Ollama Q4_K_M ์ค์ธก: 3-gram ๋ฐ๋ณต 1.8% (์์ฐ ์ด์ ๋ฐ๋ณต), EOS 100% ์ข
๋ฃ ํ์ธ.
|
| 127 |
+
> greedy ๋๋น ๋ฐ๋ณต๋ฅ **30.89% โ 0%** ํด์.
|
| 128 |
+
|
| 129 |
+
### Transformers ์ฌ์ฉ ์์
|
| 130 |
+
|
| 131 |
+
```python
|
| 132 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 133 |
+
import torch
|
| 134 |
+
|
| 135 |
+
model = AutoModelForCausalLM.from_pretrained("pathcosmos/frankenstallm")
|
| 136 |
+
tokenizer = AutoTokenizer.from_pretrained("pathcosmos/frankenstallm")
|
| 137 |
+
|
| 138 |
+
inputs = tokenizer("์๋
ํ์ธ์, ์ค๋ ๋ ์จ๊ฐ", return_tensors="pt")
|
| 139 |
+
outputs = model.generate(
|
| 140 |
+
**inputs,
|
| 141 |
+
temperature=0.7,
|
| 142 |
+
repetition_penalty=1.2,
|
| 143 |
+
top_p=0.9,
|
| 144 |
+
top_k=50,
|
| 145 |
+
max_new_tokens=512,
|
| 146 |
+
do_sample=True,
|
| 147 |
+
)
|
| 148 |
+
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 149 |
+
```
|
| 150 |
+
|
| 151 |
## ์ฌ์ฉ
|
| 152 |
|
| 153 |
- **Transformers**: ์ด ์ฒดํฌํฌ์ธํธ๋ฅผ ๊ทธ๋๋ก `from_pretrained(...)` ๋ก ๋ก๋ ๊ฐ๋ฅ.
|