Q2_0 g64 ternary is ~5Γ— slower than Q4_K_M on Samsung S22 CPU - missing optimized kernel?

#6
by chrisoutwright - opened

TL;DR: On a Snapdragon 8 Gen 1 phone (Termux, CPU-only, 4 threads), the Ternary-Bonsai-8B Q2_0 g64 GGUF (2.1 GB) generates at ~3.3 t/s, while LFM2.5-8B-A1B-GGUF model family at Q4_K_M (5.15 GB) does 15–21 t/s - measured the same morning on the same phone.

The 2-bit ternary file is 2.5Γ— smaller but ~5Γ— slower, which should be impossible if generation were just memory-bandwidth-bound. Output quality on my probes is fine.

My hypothesis: there is no optimized CPU path (repack layout / vecdot kernel) for the Q2_0 g64 block format, so it runs the generic fallback dequant. Is that expected, and is a kernel planned?

Device & setup

Device Samsung Galaxy S22 Ultra, Snapdragon 8 Gen 1
CPU 8Γ— Cortex-A (1Γ— 3.0 GHz + 3Γ— 2.5 GHz + 4Γ— 1.8 GHz), 4 threads used for inference
RAM ~11.2 GB usable (12 GB part), 4 GB zram swap
OS / runtime Android 14 (One UI 6), Termux
Engine llama.cpp, custom Android build (Vulkan + CPU backends), commit 790cf51
Ternary flags -t 4 -c 16384 -b 512 -ub 256 --jinja --reasoning-budget 300 --device none --no-repack

Both models were benchmarked the same morning, same phone, same thread count. Prompts: a 158-token and a ~550-token prompt, a 37-token math question with reasoning-budget 300, and a haiku request.

Results

Model (file size) Prompt t/s Gen t/s RSS (anon + file) Swap used by model
LFM2.5-8B-A1B Q4_K_M (5.15 GB) 33–42 15–21 ~0.3 + 5.1 GB 0
Ternary-Bonsai-8B Q2_0 g64 (2.1 GB) 4.0 3.2–3.4 2.6 + 2.1 GB 0

Notes on fairness:

  • The Q4_K_M run used the stock (repacked) layout; the ternary run used --device none --no-repack. So the ternary got the less optimized path β€” the gap is, if anything, understated.
  • The phone was warmer during the ternary run (later in the morning, after more benchmarks), which shaves a few t/s off absolute numbers but doesn't come close to explaining a 5Γ— gap.

Why I think it's a kernel issue, not a quant issue

  1. Speed doesn't scale with size. If token generation were memory-bandwidth-bound (as it is for every standard GGUF on CPU), the 2.1 GB ternary model should be the faster one β€” it streams ~2.5Γ— less weight per token than Q4_K_M. Instead it's ~5Γ— slower. That inverts the expected ordering, which points at per-token compute cost, i.e. the dequant/matmul path.
  2. Prompt processing is ~8–10Γ— slow too (4.0 t/s vs 33–42 t/s). Both pp and tg going through the same weight-matrix path being uniformly slow fits "generic fallback kernel" rather than, say, a KV-cache or attention issue.
  3. Large anonymous allocation at load: the ternary model shows 2.6 GB RssAnon on top of its 2.1 GB file mapping (Q4_K_M shows ~0.3 GB anon). Something is being expanded/dequantized into an anonymous buffer at load time β€” consistent with a format that isn't being served in its native layout.
  4. --no-repack was applied to the ternary run. If Q2_0 g64 simply has no repack support and no vecdot/matmul specialization, it lands on the slowest generic code path, while Q4_K_M keeps its optimized block kernels.

Questions

  1. Is there an optimized CPU kernel (or repack layout) for the Q2_0 g64 ternary block type? If not, is that the expected explanation for the ~5Γ— gap?
  2. Is the 2.6 GB anonymous buffer at load expected (e.g. runtime expansion to an intermediate type)? It roughly doubles the effective footprint.
  3. Any guidance on what a good target is β€” e.g. should Q2_0 g64 on 4 CPU threads be within ~10–20% of Q4_K_M speed once a proper kernel exists (it should be faster, since it moves less data per token)?

Reproduce

llama-server -m Ternary-Bonsai-8B-Q2_0_g64.gguf \
  -t 4 -c 16384 -b 512 -ub 256 --jinja \
  --reasoning-budget 300 --device none --no-repack \
  --port 8090 --host 127.0.0.1

# then any chat completion, e.g.:
curl -s localhost:8090/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"max_tokens":60,"messages":[{"role":"user","content":"What is the capital of France? Answer in one short sentence."}]}'
# check timings.predicted_per_second

Happy to run further A/B variants (repack on/off, other thread counts, -fa, KV cache types, both models under identical flags) if it helps isolate the bottleneck.

Prism ML org

Hi, thanks for the note, yeah your observation makes sense. The Q2_0 is not as optimized for all hardware. Most likely some operations are going on slower fallback kernels.

That being said the other model you are comparing with is a mixture of expert with only 1B activate weights during token generation and that means a lot less operations per token needed, so that could be another reason for speed difference. Our 8B model is dense model so all the 8B weights are active.

My hypothesis: there is no optimized CPU path (repack layout / vecdot kernel) for the Q2_0 g64 block format, so it runs the generic fallback dequant. Is that expected, and is a kernel planned?

Understand when you're working on CPU you'll have a penalty not only from limited number of processors/cores, but also a penalty based on alignment.

For speed purposes CPU's prefer to be aligned by their register size, which is going to be 32bit or 64bit. So something that may take say 10 cycles to grab the operation for a 4byte aligned block, may be 20 cycles longer for a single byte when the address isn't divisible by 4. Add shifts, AND, multiply, then xor and shifting back to put it back into the same sized space before re-writing the byte block, what may be say 50 cycles total for read multiply/add write for 1 bytes could be 10x longer depending on how it is written, nevermind if the instructions can't be run in parallel/out of order (based on unrelated registers having work that doesn't need to wait), which is how a number of modern CPU's work which is an indirect penalty as well. And if the transformer relies on FPU processing which in theory only takes a few cycles, doing FPU instruction calls you have to load the numbers in via instructions, tell it the operation to do, then do an FWAIT instruction before you can get it back, then convert it down vs just using built-in integer instructions.

I don't have the concrete numbers at the moment to give a real example of how many cycles for a single read/process/write but if you're going to run on CPU and you don't have super limited RAM, then probably use the Q8_0 models if possible. Though with Q4 it's shift 4 or AND 0xf to grab the two different halves so that's quite a bit faster.

Most programs compiled will align data in the register/size bounds for speed purposes, even if you lose a lot of space to padding, in which case you have to override structures to follow specific size/types if it's for a legacy process/format.

Sign up or log in to comment