Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
danielhanchen 
posted an update 3 days ago
Post
823
We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. 🤯

We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...

1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.

GGUF: unsloth/Kimi-K3-GGUF
GitHub repo: https://github.com/unslothai/unsloth

The headline is 1-bit vs Opus 5. The number underneath it is that K3 already shipped 4-bit.

moonshotai/Kimi-K3's config.json is mxfp4-pack-quantized, num_bits 4, group_size 32. 2.78T params in 1,560.94 GB of safetensors. That is 4.492 bits per weight before anyone quantizes anything.

Your file tree, paged with the cursor (the tree API caps at 50 entries and undercounts silently):

UD-Q8_K_XL 1,561.16 GB, 4.493 bpw
UD-Q4_K_XL 1,508.67 GB, 4.342 bpw
UD-Q2_K_XL 861.28 GB, 2.479 bpw
UD-IQ2_XXS 711.07 GB, 2.046 bpw
UD-IQ1_M 648.87 GB, 1.867 bpw
UD-IQ1_S 594.00 GB, 1.709 bpw

Q8_K_XL is 222 MB larger than Moonshot's own upload. Q4_K_XL saves 3.3%. Your card calls Q8 lossless and it is, but for a reason worth saying out loud: there is no fp16 here to be lossless against. On this model the top half of the ladder is a re-encode, not a compression.

The 4x B200 line falls out of the same table. 768 GB of HBM. IQ1_S leaves 174 GB for KV and activations, IQ1_M leaves 119, IQ2_XXS leaves 57, Q2_K_XL does not fit at all. So "1-bit" reads like a quality choice and is really the rung with headroom on one node.

Which makes the open question sharper than the aquarium. The "Q4 is safe, under 2 bits it falls apart" curve was fit on bf16 sources. Here IQ1_S is a second lossy pass on an already-lossy 4-bit grid, and only partly: the ignore list keeps self_attn, shared_experts, the dense MLP, lm_head and the vision tower out of mxfp4, so 57.2B params are true bf16 and only the 2.72T routed experts get quantized twice. The dtype counts add up exactly, and 4-bit packed plus one uint8 scale per 32 predicts 1,560.86 GB against 1,560.94 actual.

Has anyone measured whether K3's quality curve bends earlier than a bf16-native MoE of the same shape? One prompt and one sample cannot see it, and it is the number that decides whether IQ1_S is a bargain or a trap.