Instructions to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF", filename="Laguna-XS-2.1-DFlash-ik_llama-Q8_0.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Use Docker
docker model run hf.co/ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Ollama:
ollama run hf.co/ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
- Unsloth Studio
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF to start chatting
- Pi
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Docker Model Runner:
docker model run hf.co/ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
- Lemonade
How to use ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ji-farthing/Laguna-XS-2.1-DFlash-ik-llama-GGUF:Q8_0
Run and chat with the model
lemonade run user.Laguna-XS-2.1-DFlash-ik-llama-GGUF-Q8_0
List all available models
lemonade list
license: openmdw-1.1
base_model:
- poolside/Laguna-XS-2.1-DFlash
- poolside/Laguna-XS-2.1
tags:
- gguf
- ik_llama
- dflash
- speculative-decoding
- laguna
library_name: gguf
inference: false
Laguna-XS-2.1-DFlash (ik_llama.cpp GGUF)
A GGUF conversion of poolside/Laguna-XS-2.1-DFlash, the DFlash speculative drafter for poolside/Laguna-XS-2.1. It is published so the Laguna DFlash support proposed for ik_llama.cpp can be tested without re-running the conversion.
This is not a standalone model. It only works as a DFlash draft passed with --model-draft alongside the Laguna XS 2.1 target. Loading it on its own produces nothing useful.
It requires the Laguna DFlash branch of ik_llama.cpp (the arch is not in mainline yet). The pull request adds the converter, loader, and graph support for DFlashLagunaForCausalLM; a stock build will not load this file.
Files
| File | Quant | Size | sha256 |
|---|---|---|---|
Laguna-XS-2.1-DFlash-ik_llama-Q8_0.gguf |
Q8_0 | 472 MiB | a23611d163f2f3709ee04aba⦠|
Matching target
Pair it with a GGUF of the target model, poolside/Laguna-XS-2.1, converted with the same branch. The runtime numbers below used a Q4_K_M target.
Usage
llama-cli -m Laguna-XS-2.1-Q4_K_M.gguf \
-md Laguna-XS-2.1-DFlash-ik_llama-Q8_0.gguf \
--spec-type dflash:n_max=2,cross_ctx=512 \
-ngl 999 -ot exps=CPU -ngld 999 --jinja \
-p "Write a quick sort implementation in python."
The proposal depth that pays off depends on where the target runs. n_max=2 is a reasonable default across setups. With the target fully on CPU the optimum is higher (n_max=4); with the target offloaded to CUDA it is lower (n_max=1-2), and the draft should also sit on GPU (-ngld 999), because a CPU-side draft stalls the fast target and loses at every depth.
The speedup is largest on structured output (code, tables) where the drafter predicts well, and can turn into a slowdown on free prose where it does not. Measure on your own workload before committing to a setting.
Validation
Conversion (deterministic):
| Check | Result |
|---|---|
| Tensors emitted | 68 |
| Target/draft tokenizer token + merge SHA | match |
gguf-py tests |
5 passed |
Runtime smoke, Q8_0 draft with a Q4_K_M target, greedy, RTX 4070 + Core i7-11700K. Token generation, speculative vs non-speculative on the same prompt:
| Target placement | Setting | no-spec t/s | DFlash t/s | ratio |
|---|---|---|---|---|
| CPU | code, n_max=4 |
19.7 | 35.9 | 1.82x |
| CPU | code, n_max=4 (quicksort) |
20.0 | 29.9 | 1.50x |
| CPU | prose, n_max=4 |
20.8 | 16.6 | 0.80x |
| CUDA (experts on CPU) | code, n_max=2 |
46.7 | 51.3 | 1.10x |
Output is coherent and correct: generated code compiles and prose stays accurate. It is not guaranteed bit-identical to non-speculative decoding on longer greedy runs, because batched block verification occasionally flips a near-tie argmax in the target. This is a property of the target model under batched evaluation, not of the draft; generic (non-Laguna) DFlash drafts are unaffected.
This is a functional evaluation with a quantized pair, not a substitute for poolside's own BF16 chat-template evaluation.
Conversion notes
- Draft source: poolside/Laguna-XS-2.1-DFlash
- Target source: poolside/Laguna-XS-2.1
- Converted with the
DFlashLagunaForCausalLMpath added in theik_llama.cppLaguna DFlash pull request. The converter splits the packed QKV, writes the per-slice auxiliary RMS norms and per-layer attention gates, and records the causal all-SWA contract that the loader and graph read back.
License
The source model is released under the OpenMDW-1.1 license, included here as LICENSE.md. That license is retained with this distribution as it requires.