Instructions to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S # Run inference directly in the terminal: llama cli -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S # Run inference directly in the terminal: llama cli -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S # Run inference directly in the terminal: ./llama-cli -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Use Docker
docker model run hf.co/0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
- LM Studio
- Jan
- Ollama
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Ollama:
ollama run hf.co/0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
- Unsloth Studio
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS to start chatting
- Pi
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Run Hermes
hermes
- OpenClaw new
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Docker Model Runner:
docker model run hf.co/0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
- Lemonade
How to use 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S
Run and chat with the model
lemonade run user.Kimi-K3-IQ1S-REAP568-64K-4XSPARKS-UD-IQ1_S
List all available models
lemonade list
- Atomic Chat
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_SRun Hermes
hermesKimi-K3 IQ1_S REAP568 — current 600K four-Spark profile
This repository contains the four-Spark Kimi-K3 IQ1_S derivative currently
served as kimi-k3-f16k-600k-u1024. It retains 568 of 896 routed experts per
layer using the disclosed deterministic REAP568 selection and is packaged as
fourteen GGUF shards.
Current serving profile
| Setting | Value |
|---|---|
| Model name | Kimi-K3 IQ1_S REAP568 — FP16-K/F16-V 600K uBatch-1024 |
| Context | 600,000 tokens (n_ctx=600064) |
| KV cache | K F16, V F16 |
| Logical / physical batch | 2,048 / 1,024 |
| Parallel slots | 1 |
| CPU threads | 16 / batch threads 20 |
| Distribution | Local CUDA plus three RPC workers over RoCE; layer split 1:1:1:1 |
| API alias | kimi-k3-f16k-600k-u1024 |
The production launcher is
recipes/launch_4x_spark_600k_f16k.sh.
The 64K uBatch-1024 launcher remains available as a lower-context portable
profile.
Current measured prefill
A unique natural-language request with cache_prompt=false on the live
600K service measured 3,255 prompt tokens at 75.09 tok/s. A separate cold
1,769-token request measured 61.30 tok/s and 2.65 tok/s for a 16-token decode.
These are request-specific measurements; prompt length, graph shape, and cache
state materially affect throughput. See BENCHMARKS.md for
the complete record and the FP16-K comparison.
What is included
- The complete REAP568 GGUF checkpoint and tokenizer/configuration files.
- The current 600K FP16-K/FP16-V launch recipe.
- Checksum and release-verification scripts.
- Candidate-only KDA/FlashKDA optimization notes and maintenance procedure.
Expert selection and limitations
The derivative retains 568 of the original 896 routed experts in every Kimi-K3 MoE layer. Attention, KDA, MLA, AttnRes, shared experts, embeddings, latent projections, normalization, and output tensors are unchanged. The public sources did not provide complete per-expert REAP saliency values, so this is a disclosed deterministic routing proxy, not a claim of lossless pruning.
FlashKDA status
FlashKDA is not enabled in the production service. The earlier candidate did
not activate the bridge because the RPC workers were using older CUDA/RPC
libraries. Bridge-enabled libraries are staged separately on all three ranks;
the production workers and API were not interrupted. The maintenance-window
procedure and promotion gates are in
recipes/FLASHKDA_CANDIDATE_RUNBOOK.md.
Validation
sha256sum -c MANIFEST.sha256
python recipes/verify_release.py /path/to/Kimi-K3-UD-IQ1_S-REAP568
Report cold prefill, warm prefix reuse, decode, TTFT, and quality outputs separately. Do not treat throughput alone as an intelligence or safety claim.
Attribution, license, and responsibility
This is an independent derivative. It is not affiliated with or endorsed by Moonshot AI, Unsloth, llama.cpp, NVIDIA, or contributors to those projects. Comply with the upstream Kimi-K3 checkpoint, IQ1_S conversion, and runtime licenses and notices. This release is provided for research and evaluation; the model may produce incorrect, biased, unsafe, or unsuitable content. Validate outputs and use appropriate access controls. This card is not legal advice.
Sources
- Downloads last month
- 2,502
1-bit
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf 0xTank/Kimi-K3-IQ1S-REAP568-64K-4XSPARKS:UD-IQ1_S