How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "sixstringzen/Hemmingway-1-oQ4e-mtp"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default sixstringzen/Hemmingway-1-oQ4e-mtp
Run Hermes
hermes
Quick Links

Hemmingway-1 oQ4e with MTP

This repository contains an enhanced oQ4e quantization of Altworld/Hemmingway-1 for MLX and oMLX on Apple silicon. The conversion preserves the model's multi-token prediction (MTP) tensors.

Altworld developed and published the source model. sixstringzen performed this conversion and published the converted weights with their quantization report. The original model, its intended use, and its training details remain documented in the source model card.

Quantization set

This repository is part of the Hemmingway-1 oMLX oQe Quantizations collection. Every build in the set uses the same source revision, group size, non-quantized dtype, calibration pass, and MTP preservation policy.

Build Base precision Output size
oQ2e 2-bit 10.14 GiB
oQ3e 3-bit 12.22 GiB
oQ3.5e 3-bit with additional higher-precision overrides 13.19 GiB
oQ4e 4-bit 15.21 GiB
oQ6e 6-bit 21.39 GiB
oQ8e 8-bit 27.10 GiB

Quantization details

Item Value
Source model Altworld/Hemmingway-1
Source revision 4d711aac0f0043075ae334d2a3de3db3e10135c9
Quantizer oMLX 0.7.0.dev2
Method Enhanced oQ4e mixed-precision affine quantization
Base precision 4-bit
Group size 64
Non-quantized dtype bfloat16
Higher-precision tensors 115 tensors at 5-bit; language_model.lm_head at 8-bit
Calibration dataset oqe_code_multilingual
Calibration shape 128 samples at 512 tokens
Imatrix entries 504
Imatrix cache Reused from the matching source-model sensitivity pass
MTP tensors 29 preserved tensors
Output size 16,328,644,636 bytes (15.21 GiB)

oQe uses activation importance to assign additional precision to sensitive tensors. This build starts with 4-bit weights, assigns 5 bits to 115 tensors, and stores language_model.lm_head at 8-bit because that tensor had no matching imatrix entry. The quantization report records no matrix-shape mismatches and no missing weight shards.

The included oq_imatrix_report.json records the sensitivity pass, calibration settings, tensor coverage, and fallback. Strict imatrix coverage was disabled for the known language_model.lm_head fallback.

Compatibility

This model was created and tested with oMLX 0.7.0.dev2. The source model identifies its text architecture as qwen3_5_text; the converted artifact uses qwen3_5, which matches the architecture name supported by this oMLX build.

The weights use MLX safetensors and are not GGUF files. Compatibility with other MLX runtimes or earlier oMLX releases has not been verified.

Use with oMLX

Download sixstringzen/Hemmingway-1-oQ4e-mtp from the oMLX model browser, then load it as an LLM. After the model is loaded, the following request uses the local OpenAI-compatible endpoint:

curl -s http://127.0.0.1:31423/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Hemmingway-1-oQ4e-mtp",
    "messages": [
      {
        "role": "user",
        "content": "Write one vivid sentence about rain on a city window."
      }
    ],
    "max_tokens": 80,
    "temperature": 0.7,
    "enable_thinking": false,
    "stream": false
  }'

Set enable_thinking to false when you want direct prose without visible planning. Runtime defaults and the registered model identifier can vary with the local oMLX installation.

Verification

The finished artifact passed a local oMLX smoke test on 2026-09-20. oMLX loaded the model, returned a complete chat response with finish_reason: stop, and unloaded it without error.

The artifact contains four safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards. This smoke test confirms that the files load and generate through oMLX; it does not establish quality parity with the BF16 source model.

Limitations

Quantization can change word choice, coherence, and instruction following. A controlled BF16 comparison has not been published for this build.

The sensitivity pass used oMLX's oqe_code_multilingual calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately.

The original model's documented limitations and acceptable-use guidance also apply to this quantized release.

License

The source model is released under the Apache 2.0 license. This quantized derivative uses the same license; refer to the source repository for the upstream model card and attribution.

Feedback

Send compatibility reports through this repository's Community tab and include your oMLX version and Apple hardware.

Downloads last month
426
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sixstringzen/Hemmingway-1-oQ4e-mtp

Base model

Qwen/Qwen3.8-27B
Quantized
(18)
this model

Collection including sixstringzen/Hemmingway-1-oQ4e-mtp