Feature Extraction
Transformers
Safetensors
GGUF
English
Korean
qwen3_5_text
qwen3.5
backbone
headless
classification-backbone
knowledge-distillation
model-compression
vocabulary-pruning
korean
edge-ai
conversational
Instructions to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless") model = AutoModel.from_pretrained("mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0 # Run inference directly in the terminal: llama cli -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0 # Run inference directly in the terminal: llama cli -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Use Docker
docker model run hf.co/mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
- LM Studio
- Jan
- Ollama
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Ollama:
ollama run hf.co/mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
- Unsloth Desktop
- Pi
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Docker Model Runner:
docker model run hf.co/mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
- Lemonade
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Run and chat with the model
lemonade run user.qwen3.5-4l-vocab40k-en-ko-headless-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mp-juuuns/qwen3.5-4l-vocab40k-en-ko-headless:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| { | |
| "schema_version": "standalone4l-taskblind-root-v1", | |
| "replaces": { | |
| "file": "v128k_remap_manifest.json", | |
| "why": "that manifest describes the previous 128k root, whose vocabulary was carried over from the SemEval propaganda classifier's tokenizer. It is retained in this repository as the record of the superseded root and no longer describes the files at the repository root.", | |
| "previous_root_available_at": "the git revision immediately before this commit" | |
| }, | |
| "root": { | |
| "grid_point_N": 32768, | |
| "vocab_size": 39866, | |
| "layers": 4, | |
| "headless": true, | |
| "files": { | |
| "config.json": { | |
| "sha256": "e48a42e7f24d2f68b587931b632f0a6c319340c2551ae918166b81daff452910", | |
| "bytes": 2477 | |
| }, | |
| "model.safetensors": { | |
| "sha256": "694ce83253652ec1287deaab7cc328ed3ff735ca25e1cef682dca82c80892aa0", | |
| "bytes": 241285160 | |
| }, | |
| "tokenizer.json": { | |
| "sha256": "b05032abe444fc7d582c35574a85f2f5706482aedf2364a6e79d67f8c874d1ca", | |
| "bytes": 2863952 | |
| }, | |
| "tokenizer_config.json": { | |
| "sha256": "5ab9bed0a4d27949672f65ba1141d6dd5b0514fb9091d3426506ecebb5d5e294", | |
| "bytes": 1123 | |
| }, | |
| "vocab.json": { | |
| "sha256": "e20bd6447062cad010119aacf0cddb5598c0ba9a9a321d294a0d131f0f1fcf52", | |
| "bytes": 731299 | |
| }, | |
| "merges.txt": { | |
| "sha256": "44164906b71e99bc7dafbceef817185bf7ec0bc91f97715516add51af2a4181b", | |
| "bytes": 381746 | |
| }, | |
| "chat_template.jinja": { | |
| "sha256": "04b007131663760bf3e581e5a953be77044014e87efe1d2a6ca4b72ec0eac978", | |
| "bytes": 6669 | |
| }, | |
| "training_manifest.json": { | |
| "sha256": "5c4f06858e90a77bde8d31e17ac1cc18a7f396d1877b8b50143e182216e9c8aa", | |
| "bytes": 3837 | |
| }, | |
| "finalization_manifest.json": { | |
| "sha256": "3c7b4895bb8e47e8777f89d462b38c720aea10a32180ea7c3a13839a020fb424", | |
| "bytes": 1726 | |
| } | |
| }, | |
| "parameters": 120639808 | |
| }, | |
| "vocabulary_rule": { | |
| "statement": "top 32768 by BPE id, union Script=Hangul, union 256-byte alphabet, union added tokens", | |
| "rule_uses_task_statistics": false, | |
| "size_selection_basis": "none \u2014 the whole registered dyadic grid ships. No N is selected, so no N can be selected on task data.", | |
| "korean_rule": "unicode_script_hangul_union_undecodable_hangul_range_fragments", | |
| "kept_by_id_order": 32768, | |
| "hangul_total": 7234, | |
| "byte_alphabet": 256, | |
| "added_tokens": 33 | |
| }, | |
| "layer_maps": { | |
| "24to8": [ | |
| 0, | |
| 4, | |
| 6, | |
| 11, | |
| 13, | |
| 16, | |
| 20, | |
| 23 | |
| ], | |
| "8to6": [ | |
| 0, | |
| 1, | |
| 3, | |
| 4, | |
| 6, | |
| 7 | |
| ], | |
| "6to4": [ | |
| 0, | |
| 2, | |
| 3, | |
| 5 | |
| ], | |
| "selection_basis": "semeval_ranked_inherited", | |
| "disclosure": "these maps were selected by ranking candidates on SemEval-2020 Task 11 labeled train data. They are inherited here and disclosed; no task-blind justification is claimed for them. Because the KD boundary targets are a function of the map, the choice is baked into the weights." | |
| }, | |
| "distillation": { | |
| "ladder": "24 -> 8 -> 6 -> 4", | |
| "corpus": "WikiText-103-raw-v1 train shard 0", | |
| "rows": 4096, | |
| "epochs": 1, | |
| "seed": 41, | |
| "labels_used": false | |
| }, | |
| "preregistration": { | |
| "sha256": "879bdc2f951819309056dc5b0320509b3b9d288a5232e367f1945e514e7e9801", | |
| "amendments": 1, | |
| "corrections": 2 | |
| }, | |
| "why_this_grid_point": "cost, not quality. The four registered grid points are not separable on quality at three seeds; 32k is the last step where the file saving exceeds the latency and energy it costs. All four were built and measured; only this one is published as the root.", | |
| "quality_is_NR_for_this_root": "this root is headless. The macro F1 figures in the card belong to fresh-head transfer probes trained on it, not to the root." | |
| } | |