Syzygy Research

Mach-2 Additive Medium GGUF

GGUF build of Mach-2 Additive Medium for llama.cpp: Qwen3.8-Flash-Next at 1.70 bits per weight, 26.7 GB for the 125.7B text parameters, plus the model's n-gram embedding table at bf16 (102.4 GB).

This GGUF requires the Mach-1 fork of llama.cpp: SyzygyResearch/llama.cpp-mach1. Mainline llama.cpp will not load it: the model uses custom compressed tensor payloads and decode ops that only exist in the fork.

Benchmarks

Average score vs GPU memory

This GGUF holds the same weights as Mach-2 Additive Medium. Retention is Mach-2 Additive Medium's score divided by that of Qwen3.8-Flash-Next BF16, with both run on the same harness, settings and tasks.

Benchmark Tasks Mach-2 Additive Medium BF16 Retention
Humanity's Last Exam (text) 2,158 33.87 39.85 85.0%
GPQA Diamond 198 x 5 90.51 92.22 98.1%
AIME 2026 30 x 8 96.25 95.83 100.4%
Terminal-Bench 2.1 89 78.65 82.02 95.9%
AutomationBench 1.0.6 600 71.67 70.29 102.0%
NL2Repo 98 56.12 60.06 93.4%
DeepSWE v1.1 113 48.67 49.56 98.2%
τ³-Banking 97 x 5 45.10 46.41 97.2%
  • Humanity's Last Exam: up to 163,840 output tokens, graded by Artificial Analysis's judge.
  • GPQA Diamond: mean accuracy over 5 samples per question. Answers that reached 81,920 tokens were re-asked at 163,840.
  • AIME 2026: mean accuracy over 8 samples per problem, up to 81,920 output tokens.
  • Terminal-Bench 2.1: Terminus 2, one attempt per task, 4 hours per task.
  • AutomationBench: up to 50 steps per task, scored by Artificial Analysis's rule.
  • NL2Repo: OpenHands without internet access; the 102 tasks BF16 finished, less 4 whose BF16 test run hit the 1-hour scoring limit. Test runs cut at that limit are left out rather than scored 0. Mach-2 Additive Medium's score averages two runs.
  • DeepSWE: Claude Code as the agent, all 113 tasks, up to 10.5 hours per task. A task counts as solved only if the agent solved it within 700 steps.
  • τ³-Banking: tau2-bench 1.0.1, 5 trials per task, GPT-5.4 mini as the simulated user and judge, as Artificial Analysis runs it. Pass rate over the 459 conversations both models completed; 26 that outgrew the context window or lost the simulated user to an API error are left out.

Empty answers and answers still unfinished at the token limit score 0. Agents and limits differ from those behind Qwen's model card, so compare these columns with each other rather than with scores published elsewhere.

Files

File Size Notes
Mach-2-Additive-Medium.mach1.gguf 130.1 GB text model: 26.7 GB of weights + 102.4 GB n-gram table

Memory

With -ngl 99 the 26.7 GB of weights go to the GPU, so a GPU with 32 GB or more holds them with room for the KV cache. The n-gram table stays in host memory and is read a few rows per token. Run with --mlock so it stays in RAM: the machine needs about 105 GB of free RAM and ulimit -l unlimited. Without it, table pages the OS drops are re-read from disk during decoding. The model was checked on one NVIDIA A100 80 GB.

Quick start

Build the fork with CUDA:

git clone https://github.com/SyzygyResearch/llama.cpp-mach1
cd llama.cpp-mach1
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

Run directly from Hugging Face:

./build/bin/llama-cli -hf SyzygyResearch/Mach-2-Additive-Medium-GGUF -ngl 99 -fa on -c 32768 --mlock \
  --temp 1.0 --top-p 0.95 --top-k 20

or serve a local download with an OpenAI-compatible API:

./build/bin/llama-server -m Mach-2-Additive-Medium.mach1.gguf -ngl 99 -fa on -c 32768 --mlock --jinja \
  --temp 1.0 --top-p 0.95 --top-k 20

Sample at temperature 1.0, top_p 0.95, top_k 20. Without these settings the model can loop. The chat template is stored in the GGUF.

Downloads last month
2
GGUF
Model size
66B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SyzygyResearch/Mach-2-Additive-Medium-GGUF

Quantized
(1)
this model