How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf meshllm/Inkling-MTP-BF16-GGUF-beta:BF16
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default meshllm/Inkling-MTP-BF16-GGUF-beta:BF16
Run Hermes
hermes
Quick Links

Inkling MTP sidecar (beta)

This is a public beta artifact for Skippy compatibility testing. It contains Inkling's multi-token-prediction depths plus the shared embedding/output context needed by distributed final stages. It is not a standalone chat model and is not a promoted mesh-llm catalog entry.

Built with native skippy-quantize from mesh-llm revision 77048af2.

Downloads last month
115
GGUF
Model size
8B params
Architecture
inkling
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for meshllm/Inkling-MTP-BF16-GGUF-beta

Quantized
(17)
this model