--- language: en tags: - quantized - mlx - conversational - image-text-to-text - audio-text-to-text - moe base_model: - thinkingmachines/Inkling-Small base_model_relation: quantized library_name: mlx pipeline_tag: image-text-to-text --- # Inkling-Small See Inkling-Small in action: [demonstration videos](https://youtube.com/xcreate) #### Tested with an M3 Ultra 512 GiB using [Inferencer app v2.2.2](https://inferencer.com) - Multimodal inference: ~25 tokens/s @ ~278 GiB
Q9 achieves near lossless accuracy in our coding test 
| Quantization (bpw) | Perplexity | Token Accuracy | Missed Divergence | Size |
|---|---|---|---|---|
| Q4.5 | 1.11169 | 94.93% | 24.01% | 165.99 GB |
| Q6.5 | 1.06942 | 98.63% | 10.37% | 239.73 GB |
| Q9 | 1.06713 | 99.22% | 10.12% | 331.92 GB |
| Base | Untested | 100% | 0.000% | 597.59 GB |