Text Generation
MLX
Safetensors
English
hy_v3
apple-silicon
hy3
mixture-of-experts
mtp
speculative-decoding
mtplx
conversational
2-bit
Instructions to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "philipjohnbasile/hy3-demolition-mlx-lite-v1-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| license: apache-2.0 | |
| base_model: tencent/Hy3 | |
| library_name: mlx | |
| pipeline_tag: text-generation | |
| language: en | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - hy3 | |
| - mixture-of-experts | |
| - mtp | |
| - speculative-decoding | |
| - mtplx | |
| # Hy3-Demolition-MLX lite-v1-mtp | |
| [Explore the model guide](https://huggingface.co/spaces/philipjohnbasile/local-ai-guide) · [All public work](https://huggingface.co/philipjohnbasile) | |
| ## Release at a glance | |
| | | This artifact | | |
| |---|---| | |
| | Purpose | The Lite trunk with Hy3's NextN sidecar restored for MTP experiments. | | |
| | Runtime | Requires a compatible Hy3 model implementation and MTPLX runtime contract. The upstream backend shipped, but that does not qualify this artifact. | | |
| | Status | Runtime verification pending; see the evidence and limits below. | | |
| | Tensor download | 112.57 GB (104.84 GiB) of root `.safetensors` files, including any root sidecars. This is a file-size total, not peak RAM. | | |
| | Read first | Use the AR sibling for the documented generation path. No current end-to-end MTP speedup is established here. | | |
| The **MTP-equipped** variant of [lite-v1](https://huggingface.co/philipjohnbasile/hy3-demolition-mlx-lite-v1): its fused trunk with the Hy3 **NextN (Multi-Token-Prediction) sidecar** grafted back on (`num_nextn_predict_layers=1`, num_experts=192), for self-speculative decoding on MTPLX. | |
| ## Runtime status — reviewed September 10, 2026 | |
| The original `recognized-backend-pending` observation came from MTPLX | |
| 2.0.1. It is historical: the maintainer subsequently confirmed that the | |
| Hy3 and Qwen MTP backend work shipped in 2.1.0 through the release branch. | |
| [Read the upstream shipping record](https://github.com/youssofal/MTPLX/pull/142). | |
| That code shipment does not qualify this exact checkpoint. The source | |
| Hy3 implementation, runtime contract, loading, output agreement, and | |
| performance still need to be checked together. No fresh end-to-end MTP | |
| qualification of this artifact is claimed here. Upstream mlx-lm PR | |
| [#1211](https://github.com/ml-explore/mlx-lm/pull/1211) remained open at review. | |
| The documented AR path is the [sibling model](https://huggingface.co/philipjohnbasile/hy3-demolition-mlx-lite-v1). | |
| ## How it was built | |
| The MTP head consumes the trunk's final hidden state (hidden_size 4096, unchanged by pruning) and runs its own MoE on the global `num_experts`. So the base checkpoint's `mtp.*` sidecar grafts directly onto the fused AR trunk — no re-heal, no re-prune of the trunk. | |
| Graft script + `mtplx inspect` receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx (`scripts/38_mtp_sidecar_graft.py`, `eval/receipts/mtplx_inspect_*.json`). | |
| ## Limitations | |
| - **Does not run on stock mlx_lm as MTP.** The fast MTP path needs the MTPLX backend; mlx-lm's own per-token self-speculative loop is ~4.7× *slower* than AR (measured), which is why the MTPLX batched-verify backend is the target. | |
| - Everything from the base lite-v1 card applies (quantized MoE, English/agent focus, no tool execution). | |
| - End-to-end MTP behavior is **unverified** until the backend loads it; recognition is structural (`mtplx inspect`), not a live run. | |
| Base recipe + receipts: https://github.com/PhilipJohnBasile/hy3-demolition-mlx | |
| ## Other serving applications | |
| No current LM Studio or Ollama qualification of this artifact is recorded | |
| here. An architecture becoming available in one library does not establish | |
| support in every application. Use the AR sibling's pinned release recipe | |
| as the documented starting point. | |
| ## Experimental streaming | |
| The source project's SSD pager measurements concern an AR serving path. | |
| They do not qualify this MTP variant on a smaller-memory machine. | |
| [Read the AR streaming experiment](https://github.com/PhilipJohnBasile/hy3-demolition-mlx/blob/main/docs/64gb-feasibility.md). | |