Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use badtheorylabs/BTL-3-Compact with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="badtheorylabs/BTL-3-Compact", filename="model/BTL-3-Compact-AVQ2.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
| { | |
| "after": { | |
| "max_new_tokens": 256, | |
| "phase": "after", | |
| "records": [ | |
| { | |
| "category": "simple", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 22.890394, | |
| "model_variant": "after", | |
| "prompt_id": "simple-language-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Analyze the Request:**\n - User asks: \"What is your favorite coding language?\"\n - Constraint: \"Answer briefly.\"\n\n2. **Identify Core Constraints:**\n - I am an AI, so I don't have personal preferences or favorites.\n - I need to answer briefly.\n\n3. **Formulate Response:**\n - Acknowledge that I don't have favorites.\n - Mention that I'm proficient in many languages.\n - Keep it concise.\n\n4. **Draft Response:**\n - \"I don't have a favorite, but I'm proficient in many languages like Python, JavaScript, and C++.\"\n\n5. **Refine Response:**\n - \"I don't have a favorite, but I'm proficient in many languages like Python, JavaScript, and C++.\"\n\n6. **Final Output:**\n - \"I don't have a favorite, but I'm proficient in many languages like Python, JavaScript, and C++.\"\n\n7. **Check Constraints:**\n - Brief? Yes.\n - Answers the question? Yes.\n\n8. **Output Generation:**\n ", | |
| "reasoning_tokens": 254, | |
| "repetition_rate": 0.232, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "simple", | |
| "eos_reached": true, | |
| "final_answer_present": true, | |
| "finish_reason": "eos", | |
| "generated_tokens": 5, | |
| "latency_seconds": 0.741842, | |
| "model_variant": "after", | |
| "prompt_id": "simple-arithmetic-01", | |
| "raw_generation": "</think>\n\n42<|im_end|>", | |
| "reasoning_tokens": 0, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": false, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": true | |
| }, | |
| { | |
| "category": "simple", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.685617, | |
| "model_variant": "after", | |
| "prompt_id": "simple-definition-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Analyze the Request:**\n * **Topic:** Idempotency (often spelled \"idempotency\" or \"idempotence\").\n * **Constraint:** Define it in one sentence.\n * **Goal:** Provide a clear, concise definition.\n\n2. **Identify Key Concepts:**\n * What is idempotency? It's a property of operations/functions where applying them multiple times yields the same result as applying them once.\n * Common examples: `max`, `min`, `abs`, `floor`, `ceil`, `identity`, `square root`, etc.\n * Mathematical definition: `f(f(x)) = f(x)` for all `x`.\n * Practical definition: An operation is idempotent if repeating it doesn't change the outcome.\n\n3. **Draft Initial Definitions:**\n * *Draft 1:* Idempotency is a property where applying an operation multiple times produces the same result as applying it once.\n * *Draft 2:* Idempotency means that an operation is idempotent if repeating it doesn't change the outcome.\n", | |
| "reasoning_tokens": 255, | |
| "repetition_rate": 0.080645, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "simple", | |
| "eos_reached": true, | |
| "final_answer_present": true, | |
| "finish_reason": "eos", | |
| "generated_tokens": 32, | |
| "latency_seconds": 2.836183, | |
| "model_variant": "after", | |
| "prompt_id": "simple-choice-01", | |
| "raw_generation": "</think>\nStart with the standard json module, as it is robust, well-tested, and handles complex cases efficiently, saving you from reinventing the wheel.<|im_end|>", | |
| "reasoning_tokens": 0, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": false, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": true | |
| }, | |
| { | |
| "category": "coding", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.322129, | |
| "model_variant": "after", | |
| "prompt_id": "coding-function-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Understand the Request:**\n * **Goal:** Write a Python function that returns the first repeated integer in a list.\n * **Constraint:** O(n) time complexity.\n * **Requirements:** Include type hints and one short example.\n\n2. **Analyze the Problem:**\n * **Input:** A list of integers.\n * **Output:** The first integer that appears more than once in the list.\n * **Constraint:** O(n) time complexity. This means we need to iterate through the list once, or use a data structure that allows O(1) lookups (like a set or dictionary).\n * **Edge Cases:** What if there are no repeated integers? The function should return `None` or raise an exception. Let's stick to returning `None` for simplicity, but I'll make sure the type hint reflects this.\n\n3. **Develop the Solution:**\n * **Approach:** Iterate through the list, keep track of seen numbers. If a number is already seen, return it immediately.\n * **Data Structure:** A set is perfect for this. It allows", | |
| "reasoning_tokens": 256, | |
| "repetition_rate": 0.026667, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "coding", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.887859, | |
| "model_variant": "after", | |
| "prompt_id": "coding-typescript-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Deconstruct the Prompt:**\n * **Core Task:** Refactor a TypeScript function that fetches three independent endpoints sequentially.\n * **Implementation:** Show a typed `Promise.all` implementation with explicit error handling.\n * **Explanation:** Explain one failure-mode tradeoff.\n\n2. **Initial Analysis:**\n * *Sequential vs. Parallel:* The prompt says \"fetches three independent endpoints sequentially.\" This is a contradiction. If endpoints are independent, they should be fetched in parallel using `Promise.all`. If they are sequential, they should be chained. The prompt asks to refactor *to* a `Promise.all` implementation, implying the original was sequential, and the refactored version should be parallel.\n * *TypeScript:* Need to define types for the endpoints, responses, and errors.\n * *Error Handling:* Need to handle errors explicitly (e.g., catch, reject, or wrap promises).\n * *Failure-mode Tradeoff:* Need to explain a tradeoff related to failure modes (e.g., partial success vs. full failure, error propagation, etc.).\n\n3. **Drafting the Refactored Function", | |
| "reasoning_tokens": 256, | |
| "repetition_rate": 0.013605, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "coding", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.725484, | |
| "model_variant": "after", | |
| "prompt_id": "coding-debug-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Understand the Problem:**\n - **Scenario:** A Python loop appends to the same list while iterating over it.\n - **Issue:** The loop never finishes (infinite loop).\n - **Goal:** Diagnose the bug and provide the smallest safe correction.\n\n2. **Identify the Bug:**\n - **Core Issue:** Modifying a collection while iterating over it causes the iterator to see the newly added items, leading to an infinite loop.\n - **Example:** `for item in my_list: my_list.append(item)`\n - **Why it fails:** The `for` loop uses an iterator that tracks the list's length. When `append` adds an item, the iterator sees the new item, so the loop continues indefinitely.\n\n3. **Diagnose the Bug:**\n - **Explanation:** The loop never finishes because the iterator is affected by the `append` operation. The loop sees the newly appended items, causing an infinite loop.\n\n4. **Provide the Smallest Safe Correction:**\n - **Solution:** Iterate over a copy of the list to avoid modifying the original list during iteration.\n - **Code:**", | |
| "reasoning_tokens": 256, | |
| "repetition_rate": 0.026316, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "reasoning", | |
| "eos_reached": true, | |
| "final_answer_present": true, | |
| "finish_reason": "eos", | |
| "generated_tokens": 108, | |
| "latency_seconds": 9.099712, | |
| "model_variant": "after", | |
| "prompt_id": "reasoning-database-01", | |
| "raw_generation": "Here are some possible answers:\n</think>\n\n**Diagnosis:**\nThe race condition arises because two workers read the same row, modify it, and write it back without synchronization. This leads to lost updates when one worker\u2019s write overwrites another\u2019s, causing data loss or incorrect state.\n\n**Fix:**\nUse a transaction-level fix with atomic read-modify-write operations to ensure consistency.\n\n**Pseudocode:**\n```\nBEGIN TRANSACTION\n READ row\n MODIFY row\n WRITE row\nEND TRANSACTION\n```<|im_end|>", | |
| "reasoning_tokens": 6, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": false, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": true | |
| }, | |
| { | |
| "category": "reasoning", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.534852, | |
| "model_variant": "after", | |
| "prompt_id": "reasoning-cache-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Understand the Problem:**\n * **Context:** A cache was introduced.\n * **Observation:** Cache hit rate is 90% (high).\n * **Issue:** p99 latency increased (high latency for the top 1% of requests).\n * **Goal:** Provide three plausible causes and the first measurement for each.\n\n2. **Analyze the Observations:**\n * *90% hit rate:* This means most requests are served from the cache. This should generally reduce latency.\n * *p99 latency rose:* This means the tail latency (the slowest 1% of requests) got worse. This suggests that the cache is not helping for those specific requests, or the cache itself is causing delays, or there's a mismatch between cache and backend.\n\n3. **Brainstorm Plausible Causes:**\n * *Cause 1: Cache misses (the 10%)*\n * If the cache misses are slow, they pull from the backend. If the backend is slow, p99 latency rises.\n * *Measurement:* Measure backend latency for", | |
| "reasoning_tokens": 256, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "reasoning", | |
| "eos_reached": false, | |
| "final_answer_present": false, | |
| "finish_reason": "length", | |
| "generated_tokens": 256, | |
| "latency_seconds": 21.935359, | |
| "model_variant": "after", | |
| "prompt_id": "reasoning-concurrency-01", | |
| "raw_generation": "Here's a thinking process:\n\n1. **Analyze the Request:**\n * **Scenario:** Two async tasks can both observe an empty slot and create duplicate work.\n * **Comparison:** Compare a mutex, a unique database constraint, and idempotency keys.\n * **Recommendation:** Recommend one robust design.\n\n2. **Deconstruct the Scenario:**\n * *Async tasks:* This implies concurrency, non-blocking operations, potentially running on different threads or processes.\n * *Empty slot:* This suggests a resource or state that is initially empty, and tasks are trying to fill it.\n * *Duplicate work:* This means if both tasks succeed, they create redundant entries or actions, leading to inefficiency or errors.\n * *Goal:* Prevent duplication, ensure exactly one task succeeds.\n\n3. **Evaluate the Three Approaches:**\n\n * **Approach 1: Mutex (Lock)**\n * *Mechanism:* A lock ensures exclusive access to a resource.\n * *Pros:* Simple, guarantees exclusive access, prevents race conditions.\n * *Cons:* Locks are often tied to a specific process/thread.", | |
| "reasoning_tokens": 256, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": true, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": false | |
| }, | |
| { | |
| "category": "instruction", | |
| "eos_reached": true, | |
| "final_answer_present": true, | |
| "finish_reason": "eos", | |
| "generated_tokens": 4, | |
| "latency_seconds": 0.667707, | |
| "model_variant": "after", | |
| "prompt_id": "instruction-direct-01", | |
| "raw_generation": "</think>\n\nready<|im_end|>", | |
| "reasoning_tokens": 0, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": false, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": true | |
| }, | |
| { | |
| "category": "instruction", | |
| "eos_reached": true, | |
| "final_answer_present": true, | |
| "finish_reason": "eos", | |
| "generated_tokens": 32, | |
| "latency_seconds": 3.190571, | |
| "model_variant": "after", | |
| "prompt_id": "instruction-format-01", | |
| "raw_generation": "</think>\n\n```json\n{\n \"language\": \"Python\",\n \"reason\": \"Easy learning, versatile, popular.\"\n}\n```<|im_end|>", | |
| "reasoning_tokens": 0, | |
| "repetition_rate": 0.0, | |
| "scaffold_prefix": false, | |
| "template_variant": "pinned_official_identical", | |
| "think_end_reached": true | |
| } | |
| ], | |
| "summary": { | |
| "after/pinned_official_identical": { | |
| "eos_rate": 0.416667, | |
| "final_answer_rate": 0.416667, | |
| "mean_latency_seconds": 14.126, | |
| "mean_reasoning_tokens": 149.583, | |
| "mean_repetition_rate": 0.031603, | |
| "n": 12, | |
| "scaffold_rate": 0.583333, | |
| "think_end_rate": 0.416667 | |
| } | |
| } | |
| }, | |
| "termination_weight": 8.0 | |
| } | |