Confucius4-R2T2-Q4_K_M-GGUF / QUANTIZATION.md
Nairod785's picture
Final measured CUDA/CPU speed (drift-free build)
b1ea192 verified
|
Raw
History Blame Contribute Delete
9.84 kB

Confucius4-R2T2 — Quantization Technical Report & Research Log

Companion technical report for Confucius4-R2T2-Q4_K_M-GGUF. Records empirical quantization measurements, ablation findings, and validation benchmarks on R2T2. For the general per-tensor quantization arms methodology and principles, see QUANTIZATION_ARMS.md.


Result: the floor is 1.187 GB — 52% below the 2.478 GB Q8_0 reference.

Metric Reference Baseline Shipped Release (Arm M)
File r2t2-q8_0.gguf r2t2-q4_k_m.gguf
Size 2.478 GB 1.187 GB
Composition Full Q8_0 / F16 embed Tower @ Q4_K · Gate/Up @ Q4_K · Attention @ Q4_K · down_proj @ Q6_K · embed_tokens @ Q2_K

Historical note: The intermediate experimental test arms evaluated during the research sweep were scratch builds used for the ablation study. The final selected model—**r2t2-q4_k_m.gguf** (Arm M)—is the published artifact provided in this repository.


1. Provenance

Component Value
Source netease-youdao/Confucius4-R2T2
Revision snapshot 185ce639118ad1362d049ca0d8ed04b6ec5cd6c9
Weights A single 4.076 GB model.safetensors
Side Files Embedded tokenizer, processor config, generation config
Tensor Census Tower: 147 · Attention: 112 · Gate/Up: 56 · Down: 28

2. The Arm Inventory

embed is the tied thinker.model.embed_tokens.weight, which feeds both the input lookup and the output logits. The converter pins it to F16 by default; the table below shows where that pin was deliberately overridden.

Arm File Attention Gate/Up down_proj Embed Tower Size (GB) Verdict
q8 r2t2-q8_0 Q8_0 Q8_0 Q8_0 F16 BF16/Q8_0 2.478 Reference
— tower-q4-lm-q8 Q8_0 Q8_0 Q8_0 F16 BF16/Q4_K 2.321 Tower floor probe
q6_k r2t2-q6_k Q6_K Q6_K Q6_K F16 BF16/Q6_K 2.060 Passes
q5_k r2t2-q5_k Q5_K Q5_K Q5_K F16 BF16/Q5_K 1.832 Fails Russian; slower
— lm-q4-tower-q8 Q4_K Q4_K Q4_K F16 BF16/Q8_0 1.773 Fails
q4_k r2t2-q4_k Q4_K Q4_K Q4_K F16 BF16/Q4_K 1.616 Fails
C r2t2-C Q6_K Q6_K Q6_K Q4_K BF16/Q4_K 1.532 Passes
D r2t2-D Q6_K Q4_K Q6_K Q4_K BF16/Q4_K 1.351 Passes
De3 r2t2-De3 Q6_K Q4_K Q6_K Q3_K BF16/Q4_K 1.309 Passes
Dt3 r2t2-Dt3 Q6_K Q4_K Q6_K Q4_K BF16/Q3_K 1.309 Empty output
D5 r2t2-D5 Q6_K Q4_K Q5_K Q4_K BF16/Q4_K 1.304 Fails German
Da4 r2t2-Da4 Q4_K Q4_K Q6_K Q4_K BF16/Q4_K 1.260 Passes
E r2t2-E Q6_K Q4_K Q4_K Q4_K BF16/Q4_K 1.260 Fails German
J r2t2-J Q4_K Q4_K Q6_K Q3_K BF16/Q4_K 1.219 Marginal — fails French
K r2t2-K Q6_K Q4_K Q4_K Q3_K BF16/Q4_K 1.219 Fails German
M r2t2-q4_k_m Q4_K Q4_K Q6_K Q2_K BF16/Q4_K 1.187 Floor — passes all
L r2t2-L Q4_K Q4_K Q5_K Q3_K BF16/Q4_K 1.172 Fails German
G r2t2-G Q4_K Q4_K Q4_K Q4_K BF16/Q4_K 1.169 Fails German
N r2t2-N Q4_K Q4_K Q5_K Q2_K BF16/Q4_K 1.140 Fails German
G3 r2t2-G3 Q4_K Q4_K Q4_K Q3_K BF16/Q4_K 1.128 Fails German
P r2t2-P Q4_K Q3_K Q6_K Q2_K BF16/Q4_K 1.093 Empty output

3. Why Arm M Works

The Error Budget is Cumulative, Not Per-Tensor

What a transformer block can afford depends on what the rest of the network has already spent:

  • Da4 (attention @ Q4_K, down_proj @ Q6_K) passes.
  • E (attention @ Q6_K, down_proj @ Q4_K) also passes early probes.
  • G (both @ Q4_K) fails.

Each individual Q4_K move fits on its own; both together exceed the budget. The LM decoder affords exactly one of {attention, down_proj} at Q4_K. Arm M spends that budget on attention while holding down_proj at Q6_K.

down_proj is the Hard Floor

down_proj @ Q6_K is the single non-negotiable constraint. Every arm that put it at Q5_K or Q4_K—D5, E, L, N, K, G, G3—emitted English for German input. There is no arm in the table with down_proj below Q6_K that passes.

Architectural mechanism: down_proj writes straight into the residual stream, so its error propagates through every remaining layer and compounds. In contrast, gate_proj/up_proj error passes through SwiGLU, which bounds it—allowing gate/up to tolerate Q4_K.

The Embedding is the Cheapest Lever

embed_tokens is a lookup: its error enters once and leaves once without compounding. It contains 311 M parameters, making it the largest single size lever per unit of quality. Arm M quantizes it to Q2_K—reducing the floor from 1.26 GB (Da4) to 1.187 GB.

Low-Bit Matmuls are a Cliff

  • Dt3 (tower @ Q3_K) → Empty output
  • P (gate/up @ Q3_K) → Empty output

Both produced nothing, not degraded text. There is no gradual size/quality trade below Q4_K on a matmul—the representation stops working abruptly.


4. The Marginal Arm: J

Arm Attention down_proj Embed Outcome
M Q4_K Q6_K Q2_K Passes everything
J Q4_K Q6_K Q3_K Fails French clip
Da4 Q4_K Q6_K Q4_K Passes everything
De3 Q6_K Q6_K Q3_K Passes everything

J and M differ in nothing but the embedding quant, and the higher-precision one (Q3_K) failed. At attention @ Q4_K the arm is at the edge of the error budget; embed @ Q3_K tips it over, while embed @ Q2_K happens not to. De3 confirms Q3_K embed is fully stable once attention is raised to Q6_K.


5. Speed Benchmarks

Measured on CPU (pooled ratio against Q8_0, positive = faster):

Arm vs Q8_0 Speed Note
tower-q4-lm-q8 ±0% Isolates runtime to LM matmuls; tower quant is compute-free
r2t2-q6_k +5% Passes
r2t2-C +12% Passes
r2t2-D +15% Passes
r2t2-D5 +18% Fails German
r2t2-E +22% Fails German
r2t2-q4_k_m (M) +20%+ Floor — passes all
r2t2-q5_k −20% (SLOWER) Strictly dominated; unoptimized kernels

Key findings:

  1. Q8_0 is not the fast path: Every arm with the decoder LM quantized beats Q8_0.
  2. Q5_K is strictly dominated: It is ~20% slower than Q8_0 and switches to English on Russian. Q6_K and Q4_K have well-optimized vectorized kernels; Q5_K does not.

5b. Measured speed (transcribe.cpp, CUDA and CPU)

Hardware: RTX 4070 Laptop GPU (8 GB) and i9-13900H. Each figure is 3 runs per configuration: the first is discarded and runs 2 and 3 are averaged. Decode includes prefill; tokens/s = generated tokens / decode time.

Backend Clip Total Realtime Encode Decode Decode tokens/s
CUDA German, 29.3 s (49 tokens) 421 ms 69.6x 50 ms 357 ms 137
CUDA English (JFK), 11.0 s (29 tokens) 232 ms 47.4x 26 ms 200 ms 145
CPU German, 29.3 s 9.32 s 3.1x 2.47 s 6.83 s 7.2
CPU English (JFK), 11.0 s 4.13 s 2.7x 0.91 s 3.22 s 9.0

Steady-state CUDA decode is 6.0 ms per token (166 tokens/s). Prefill of the ~400-token audio prompt takes ~52 ms of the decode figure.

Where the CUDA speed comes from:

  • Arm M vs Q8_0: decode is ~1.6x faster, because decode is weight-bandwidth bound.
  • CUDA graphs (GGML_CUDA_GRAPHS=ON, now the transcribe.cpp default for CUDA builds): +15 % decode. The per-token graph is small and launch-bound.
  • Attention reads only the filled KV rows: 7.08 -> 6.60 ms per token. Before, the whole padded cache window was scanned.
  • Matrix-vector row blocking at K = 2048: the Q2_K output-logits matmul is -38 %, which removes most of the embedding cost.

Accuracy: all of the above is numerically identical to the unoptimized build. On 100 FLEURS test clips (20 each of fr, de, en, es, it), transcripts are 100/100 byte-identical and WER is unchanged.

Two further CUDA options were measured and left off, because they change rounding and flipped 14/100 transcripts (WER within noise, but not identical):

  • TRANSCRIBE_QKV_PACK=1: fused Q/K/V matmul.
  • TRANSCRIBE_ENCODER_FLASH=1

Together they are worth ~5 % decode.

Optional speculative drafting: with a parakeet transcript as a draft prior, decode is 1.84x faster at K = 5 and byte-identical. Set it with TRANSCRIBE_SPEC_PRIOR_TEXT; it is experimental.


6. Reproduction Recipe

To build r2t2-q4_k_m.gguf from the base Safetensors weights:

# Census: tower 147, attn 112, gate/up 56, down 28
# Base: q6_k
# Overrides:
#   tower: q4_k
#   attention (q, k, v, o_proj): q4_k
#   gate/up (gate_proj, up_proj): q4_k
#   embed_tokens: q2_k
#   down_proj: left at base q6_k (non-negotiable floor)

audiocpp_gguf \
  --input /path/to/Confucius4-R2T2/model.safetensors \
  --root /path/to/Confucius4-R2T2 \
  --family confucius4_r2t2 \
  --model-spec model_specs/confucius4_r2t2.json \
  --output r2t2-q4_k_m.gguf \
  --type q6_k \
  --keep-type thinker.model.embed_tokens.weight=q2_k \
  "${ATTN_OVERRIDES[@]}" \
  "${GATEUP_OVERRIDES[@]}" \
  "${TOWER_OVERRIDES[@]}" \
  --overwrite

Run from bash, not cmd.exe, to prevent command-line length truncation.