Instructions to use void0x14/echo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use void0x14/echo with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf void0x14/echo:Q4_K_M # Run inference directly in the terminal: llama cli -hf void0x14/echo:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf void0x14/echo:Q4_K_M # Run inference directly in the terminal: llama cli -hf void0x14/echo:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf void0x14/echo:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf void0x14/echo:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf void0x14/echo:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf void0x14/echo:Q4_K_M
Use Docker
docker model run hf.co/void0x14/echo:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use void0x14/echo with Ollama:
ollama run hf.co/void0x14/echo:Q4_K_M
- Unsloth Desktop
- Pi
How to use void0x14/echo with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf void0x14/echo:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "void0x14/echo:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use void0x14/echo with Docker Model Runner:
docker model run hf.co/void0x14/echo:Q4_K_M
- Lemonade
How to use void0x14/echo with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull void0x14/echo:Q4_K_M
Run and chat with the model
lemonade run user.echo-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use void0x14/echo with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf void0x14/echo:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default void0x14/echo:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use void0x14/echo with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf void0x14/echo:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "void0x14/echo:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| from __future__ import annotations | |
| import json | |
| import math | |
| from dataclasses import asdict, dataclass | |
| from pathlib import Path | |
| from typing import Mapping, Sequence | |
| from safetensors import safe_open | |
| class ValidationReport: | |
| parameter_count: int | |
| tensor_count: int | |
| layer_count: int | |
| visual_tensor_count: int | |
| mtp_tensor_count: int | |
| tied_lm_head_present: bool | |
| valid: bool | |
| def _numel(shape: Sequence[int]) -> int: | |
| return math.prod(int(dimension) for dimension in shape) | |
| def _layer_indices(keys: Sequence[str]) -> set[int]: | |
| indices = set() | |
| prefix = "model.layers." | |
| for key in keys: | |
| if key.startswith(prefix): | |
| indices.add(int(key[len(prefix):].split(".", 1)[0])) | |
| return indices | |
| def validate_state_dict_keys( | |
| shapes: Mapping[str, Sequence[int]], | |
| config: Mapping[str, object], | |
| minimum: int, | |
| maximum: int, | |
| ) -> ValidationReport: | |
| keys = tuple(shapes.keys()) | |
| visual_count = sum(key.startswith("model.visual.") for key in keys) | |
| mtp_count = sum(key.startswith("mtp.") for key in keys) | |
| tied_head = "lm_head.weight" in shapes and bool(config.get("tie_word_embeddings", False)) | |
| if visual_count or mtp_count or tied_head: | |
| raise ValueError("vision/MTP tensors or a duplicate tied head are present") | |
| if config.get("model_type") != "qwen3_5_text": | |
| raise ValueError("standalone text config must use model_type qwen3_5_text") | |
| layer_count = int(config["num_hidden_layers"]) | |
| expected_indices = set(range(layer_count)) | |
| actual_indices = _layer_indices(keys) | |
| if actual_indices != expected_indices: | |
| raise ValueError(f"layer indices are not contiguous: expected {expected_indices}, got {actual_indices}") | |
| layer_types = tuple(config["layer_types"]) | |
| if len(layer_types) != layer_count: | |
| raise ValueError("config layer_types length does not match num_hidden_layers") | |
| parameter_count = sum(_numel(shape) for shape in shapes.values()) | |
| if not minimum <= parameter_count <= maximum: | |
| raise ValueError(f"parameter count {parameter_count} is outside [{minimum}, {maximum}]") | |
| return ValidationReport( | |
| parameter_count=parameter_count, | |
| tensor_count=len(keys), | |
| layer_count=layer_count, | |
| visual_tensor_count=visual_count, | |
| mtp_tensor_count=mtp_count, | |
| tied_lm_head_present=tied_head, | |
| valid=True, | |
| ) | |
| def validate_checkpoint( | |
| config_path: str | Path, | |
| weights_path: str | Path, | |
| minimum: int, | |
| maximum: int, | |
| ) -> ValidationReport: | |
| config = json.loads(Path(config_path).read_text(encoding="utf-8")) | |
| shapes = {} | |
| with safe_open(str(weights_path), framework="pt", device="cpu") as handle: | |
| for key in handle.keys(): | |
| shapes[key] = tuple(handle.get_slice(key).get_shape()) | |
| report = validate_state_dict_keys(shapes, config, minimum, maximum) | |
| Path(config_path).with_name("validation.json").write_text( | |
| json.dumps(asdict(report), indent=2, sort_keys=True) + "\n", encoding="utf-8" | |
| ) | |
| return report | |