avlp12 commited on
Commit
8994ede
·
verified ·
1 Parent(s): 37ec3a0

Link live Q6 two-box sibling

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -25,7 +25,7 @@ base_model_relation: quantized
25
 
26
  # Inkling-975B-Alis-MLX-Dynamic-3.7bpw
27
 
28
- > The **quality / golden-spot** tier of the Inkling · Alis MLX Dynamic family — siblings: the [\~2.7 bpw size-optimal build](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-2.7bpw) and a \~6.5 bpw two-box Q6 performance build (uploading as its certification completes). Full family: [Inkling 975B · Alis MLX Dynamic collection](https://huggingface.co/collections/avlp12/inkling-975b-alis-mlx-dynamic-6a5f2d32de21740f9fdfc390).
29
 
30
  **Apple Silicon (MLX) mixed-precision quantization of [thinkingmachines/Inkling](https://huggingface.co/thinkingmachines/Inkling)** — a 975B-class multimodal Mixture-of-Experts model (66 hybrid decoder layers, **256 routed experts (top-6) + 2 shared** per MoE layer, hidden 6144, sliding-window attention + short-convolution hybrid, vision + audio front-ends, 201K vocab).
31
 
@@ -79,7 +79,7 @@ The result is conservative by construction: **you get provably-not-worse-than-ba
79
  |---|---|---|---|
80
  | **this** — quality / golden spot | **3.71** | 409 GiB | single 512 GB Mac, best quality in one box |
81
  | [capacity / size-optimal](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-2.7bpw) | 2.72 | 299 GiB | single Mac with generous headroom / smaller boxes |
82
- | Q6 teacher (two-box) | \~6.5 | \~790 GiB | maximum fidelity, 2 × 512 GB pipeline serving |
83
 
84
  ---
85
 
 
25
 
26
  # Inkling-975B-Alis-MLX-Dynamic-3.7bpw
27
 
28
+ > The **quality / golden-spot** tier of the Inkling · Alis MLX Dynamic family — siblings: the [\~2.7 bpw size-optimal build](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-2.7bpw) and the [\~6.6 bpw two-box Q6 performance build](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-6.6bpw). Full family: [Inkling 975B · Alis MLX Dynamic collection](https://huggingface.co/collections/avlp12/inkling-975b-alis-mlx-dynamic-6a5f2d32de21740f9fdfc390).
29
 
30
  **Apple Silicon (MLX) mixed-precision quantization of [thinkingmachines/Inkling](https://huggingface.co/thinkingmachines/Inkling)** — a 975B-class multimodal Mixture-of-Experts model (66 hybrid decoder layers, **256 routed experts (top-6) + 2 shared** per MoE layer, hidden 6144, sliding-window attention + short-convolution hybrid, vision + audio front-ends, 201K vocab).
31
 
 
79
  |---|---|---|---|
80
  | **this** — quality / golden spot | **3.71** | 409 GiB | single 512 GB Mac, best quality in one box |
81
  | [capacity / size-optimal](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-2.7bpw) | 2.72 | 299 GiB | single Mac with generous headroom / smaller boxes |
82
+ | [Q6 teacher (two-box)](https://huggingface.co/avlp12/Inkling-975B-Alis-MLX-Dynamic-6.6bpw) | 6.60 | 728 GiB | maximum fidelity, 2 × 512 GB pipeline serving |
83
 
84
  ---
85