Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use badtheorylabs/BTL-3-Compact with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="badtheorylabs/BTL-3-Compact", filename="model/BTL-3-Compact-AVQ2.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
| # BTL-3 Compact native release validation | |
| Date: 2026-07-19 | |
| ## Candidate | |
| - Public model: **BTL-3 Compact**. | |
| - Model lineage: Qwen3.6-27B plus BTL RL0013 and the frozen behavior repair. | |
| - Scope: text-only coding, tool use and agent behavior. | |
| - Representation: full 64-layer AVQ2/UniSVQ decoder, two measured INT4 | |
| demotions, selected BF16 islands, packed vocabulary matrices, rank-32 head | |
| correction and a small behavior adapter. | |
| - Source payload bytes before GGUF packing: 8,572,070,080. | |
| - Source weight bytes: 8,551,772,952. | |
| - Portable GGUF bytes: 8,392,369,600. | |
| - Portable GGUF SHA-256: | |
| `2ddf9527620a17a2a6739d184a7096c45712092e6589128792ec6254e94dc30c`. | |
| The package is complete enough to instantiate and generate without downloading | |
| or loading the BF16 Qwen checkpoint. | |
| ## Native runtime proof | |
| The standalone loader meta-initializes the Qwen3.6 text architecture and | |
| installs only the package's small state, packed decoder tensors, packed | |
| embedding, packed head, rank-32 output correction and behavior LoRA. | |
| The H100 smoke test observed: | |
| - no surviving dense compatible decoder matrices; | |
| - exact AVQ2 CUDA-kernel parity with the unpacked reference; | |
| - INT4 maximum absolute kernel error of `3.0517578125e-05`; | |
| - standalone model peak CUDA allocation of 8,552,500,736 bytes; | |
| - successful autoregressive generation. | |
| The source payload was subsequently exported into the portable GGUF without | |
| reconstructing dense weights. The exporter byte-verified all 2,416 payloads | |
| and reported no unsupported tensors or native runtime gaps. The exact GGUF | |
| then passed native llama.cpp generation on Apple Metal and NVIDIA CUDA. | |
| MLX, WebGPU, phone execution, and stock-engine compatibility remain unverified. | |
| ## Fresh sealed gate | |
| Benchmark ID: `btl-fresh-tool-gate-2026-07-19-v1` | |
| Cases SHA-256: | |
| `d656a7862e16e64ed3a359ba1de10f7eafefad77f6cf5a8264d60287e1890a45` | |
| The gate was authored after compression and behavior-repair choices were | |
| frozen. It contains 100 scored turns: | |
| - 20 single calls; | |
| - 20 parallel calls; | |
| - 20 sequential calls; | |
| - 20 parallel-multiple calls; | |
| - 20 abstention decisions. | |
| All tool families are first-party and use a new `lumenharbor_*` namespace. | |
| Mechanical QA found no schema errors, duplicate IDs, parse errors or exact tool | |
| name overlap with repository training/evaluation data. | |
| This is a private synthetic contract-retention gate. It is not a public coding | |
| benchmark and must not be presented as a frontier benchmark score. | |
| ## Full-precision teacher | |
| The frozen RL0013 teacher scored: | |
| | Category | Correct | Total | | |
| |---|---:|---:| | |
| | Single | 20 | 20 | | |
| | Parallel | 20 | 20 | | |
| | Sequential | 20 | 20 | | |
| | Parallel-multiple | 10 | 20 | | |
| | Abstention | 20 | 20 | | |
| | Overall | 90 | 100 | | |
| The release metric is conditional retention on these 90 teacher-correct turns, | |
| reported separately for every category and overall. Absolute student accuracy | |
| is also retained in the raw result. | |
| ## Standalone result | |
| The standalone package scored: | |
| | Category | Student correct | Total | Teacher-correct retained | Retention | | |
| |---|---:|---:|---:|---:| | |
| | Single | 20 | 20 | 20 / 20 | 100% | | |
| | Parallel | 20 | 20 | 20 / 20 | 100% | | |
| | Sequential | 20 | 20 | 20 / 20 | 100% | | |
| | Parallel-multiple | 3 | 20 | 3 / 10 | 30% | | |
| | Abstention | 20 | 20 | 20 / 20 | 100% | | |
| | **Overall** | **83** | **100** | **83 / 90** | **92.2%** | | |
| All generations stopped. The measured malformed rate was 7%, entirely within | |
| the difficult parallel-multiple family. | |
| The seven teacher-correct/student-wrong parallel-multiple cases were not parser | |
| false negatives. The package emitted a fluent abstention instead of making the | |
| three requested independent calls. This is a real over-abstention failure under | |
| novel multi-call schemas. | |
| ### Decision | |
| The package passes a 90% **overall** conditional-retention rule. It fails a 90% | |
| **per-category** rule because parallel-multiple retention is 30%. | |
| Do not alter compression based on this sealed result. Any behavior repair aimed | |
| at these cases creates a new candidate and requires a newly authored, untouched | |
| release gate. | |
| ## Release boundary | |
| This result validates that the exact text-only source payload is physically | |
| standalone and retains more than 90% overall on the fresh CUDA gate. The GGUF | |
| export preserved those payload bytes exactly. It does not support a claim of | |
| uniformly preserved behavior or phone deployment. | |
| The release includes a distributable macOS native runtime and exact-artifact | |
| throughput measurements on Apple M2 and RTX PRO 6000. Other GPU packages, | |
| stock Ollama/LM Studio execution, mobile runtimes, and a public compact-specific | |
| coding benchmark remain separate gates. | |