Instructions to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hy3-Demolition-MLX lite-v1-mtp
Explore the model guide · All public work
Release at a glance
| This artifact | |
|---|---|
| Purpose | The Lite trunk with Hy3's NextN sidecar restored for MTP experiments. |
| Runtime | Requires a compatible Hy3 model implementation and MTPLX runtime contract. The upstream backend shipped, but that does not qualify this artifact. |
| Status | Runtime verification pending; see the evidence and limits below. |
| Tensor download | 112.57 GB (104.84 GiB) of root .safetensors files, including any root sidecars. This is a file-size total, not peak RAM. |
| Read first | Use the AR sibling for the documented generation path. No current end-to-end MTP speedup is established here. |
The MTP-equipped variant of lite-v1: its fused trunk with the Hy3 NextN (Multi-Token-Prediction) sidecar grafted back on (num_nextn_predict_layers=1, num_experts=192), for self-speculative decoding on MTPLX.
Runtime status — reviewed September 10, 2026
The original recognized-backend-pending observation came from MTPLX
2.0.1. It is historical: the maintainer subsequently confirmed that the
Hy3 and Qwen MTP backend work shipped in 2.1.0 through the release branch.
Read the upstream shipping record.
That code shipment does not qualify this exact checkpoint. The source Hy3 implementation, runtime contract, loading, output agreement, and performance still need to be checked together. No fresh end-to-end MTP qualification of this artifact is claimed here. Upstream mlx-lm PR #1211 remained open at review.
The documented AR path is the sibling model.
How it was built
The MTP head consumes the trunk's final hidden state (hidden_size 4096, unchanged by pruning) and runs its own MoE on the global num_experts. So the base checkpoint's mtp.* sidecar grafts directly onto the fused AR trunk — no re-heal, no re-prune of the trunk.
Graft script + mtplx inspect receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx (scripts/38_mtp_sidecar_graft.py, eval/receipts/mtplx_inspect_*.json).
Limitations
- Does not run on stock mlx_lm as MTP. The fast MTP path needs the MTPLX backend; mlx-lm's own per-token self-speculative loop is ~4.7× slower than AR (measured), which is why the MTPLX batched-verify backend is the target.
- Everything from the base lite-v1 card applies (quantized MoE, English/agent focus, no tool execution).
- End-to-end MTP behavior is unverified until the backend loads it; recognition is structural (
mtplx inspect), not a live run.
Base recipe + receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx
Other serving applications
No current LM Studio or Ollama qualification of this artifact is recorded here. An architecture becoming available in one library does not establish support in every application. Use the AR sibling's pinned release recipe as the documented starting point.
Experimental streaming
The source project's SSD pager measurements concern an AR serving path. They do not qualify this MTP variant on a smaller-memory machine. Read the AR streaming experiment.
- Downloads last month
- 134
2-bit
Model tree for philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp
Base model
tencent/Hy3
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True)