Text Generation
Transformers
Safetensors
GGUF
Korean
English
llama
3b
korean
from-scratch
orpo
instruction-tuned
preference-aligned
fp8
b200
Eval Results (legacy)
text-generation-inference
Instructions to use pathcosmos/frankenstallm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pathcosmos/frankenstallm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pathcosmos/frankenstallm")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pathcosmos/frankenstallm") model = AutoModelForCausalLM.from_pretrained("pathcosmos/frankenstallm", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use pathcosmos/frankenstallm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: llama cli -hf pathcosmos/frankenstallm:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: llama cli -hf pathcosmos/frankenstallm:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf pathcosmos/frankenstallm:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pathcosmos/frankenstallm:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf pathcosmos/frankenstallm:Q4_K_M
Use Docker
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use pathcosmos/frankenstallm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pathcosmos/frankenstallm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- SGLang
How to use pathcosmos/frankenstallm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pathcosmos/frankenstallm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pathcosmos/frankenstallm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pathcosmos/frankenstallm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use pathcosmos/frankenstallm with Ollama:
ollama run hf.co/pathcosmos/frankenstallm:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use pathcosmos/frankenstallm with Docker Model Runner:
docker model run hf.co/pathcosmos/frankenstallm:Q4_K_M
- Lemonade
How to use pathcosmos/frankenstallm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pathcosmos/frankenstallm:Q4_K_M
Run and chat with the model
lemonade run user.frankenstallm-Q4_K_M
List all available models
lemonade list
- Atomic Chat
| # Korean 3B SFT v2 Configuration | |
| # | |
| # Base model: checkpoints/korean_3b_fp8_run1/checkpoint-0057000 (3B params pretrained) | |
| # SFT v2 ๋ชฉํ: v1์ underfitting ํด๊ฒฐ + forgetting ๋ฐฉ์ง (data mixing) | |
| # ์ํคํ ์ฒ: LLaMA-3 3B ์ฐธ๊ณ (d=3072, 28L, 24H, GQA 8:1) | |
| # | |
| # ์คํ: bash scripts/launch_3b_sft_v2.sh | |
| # | |
| # [์ค๊ณ ๊ทผ๊ฑฐ โ SFT v1 ์คํจ ๋ถ์ 2026-03-06] | |
| # v1 ๋ฌธ์ : lr=1e-5 โ val_loss ๋ณํ 0 (์ฌ์ค์ ํ์ต ์ ๋จ) | |
| # v2 ๋ณ๊ฒฝ: | |
| # - lr: 1e-5 โ 5e-5 (5๋ฐฐ โ, 3B SFT ํ์ค ๋ฒ์) | |
| # - batch: 4 ร 8GPU ร 8 grad_accum = 256 eff_batch (v1 ๋๋น 4๋ฐฐ โ) | |
| # - warmup: 500 โ 2000 (๋์ LR์ ๋ง์ถฐ ์์ ํ) | |
| # - max_steps: 33000 โ 15000 (์๋ ด ๋นจ๋ผ์ง, ๊ณผ์ ํฉ ๋ฐฉ์ง) | |
| # - weight_decay: 0.01 โ 0.05 (forgetting ์ต์ ) | |
| # - data mixing: SFT 70% + pretrain 30% (forgetting ๋ฐฉ์ง) | |
| model: | |
| vocab_size: 64000 | |
| d_model: 3072 | |
| n_layers: 28 | |
| n_heads: 24 | |
| n_kv_heads: 8 | |
| d_ffn: 8192 | |
| max_seq_len: 4096 | |
| rope_theta: 500000.0 | |
| dropout: 0.0 | |
| bias: false | |
| use_flash_attn: true | |
| use_fp8: true | |
| train: | |
| max_steps: 15000 # v1 33000 โ 15000 (์๋ ด ๋นจ๋ผ์ง) | |
| batch_size: 4 # v1 2 โ 4 (VRAM ์ฌ์ ์ถฉ๋ถ: 48/183GB) | |
| grad_accum_steps: 8 # v1 4 โ 8 (eff_batch: 4 ร 8GPU ร 8 = 256) | |
| lr: 5.0e-5 # v1 1e-5 โ 5e-5 (5๋ฐฐ โ, underfitting ํด๊ฒฐ) | |
| weight_decay: 0.05 # v1 0.01 โ 0.05 (forgetting ์ต์ ) | |
| warmup_steps: 2000 # v1 500 โ 2000 (๋์ LR ์์ ํ) | |
| max_grad_norm: 1.0 # gradient clipping | |
| log_interval: 10 | |
| save_interval: 2000 | |
| eval_interval: 500 | |
| use_amp: false | |
| compile_model: false | |
| neftune_alpha: 5.0 # NEFTune noise injection (์ ์ง) | |
| # Data mixing (forgetting ๋ฐฉ์ง) | |
| pretrain_mix_ratio: 0.3 # pretrain ๋ฐ์ดํฐ 30% ํผํฉ | |
| pretrain_data: data/3b_train.bin # pretrain ๋ฐ์ดํฐ ๊ฒฝ๋ก | |
| tokenizer: | |
| vocab_size: 64000 | |
| type: sentencepiece_unigram | |