Tessera-4B-Preview
Collection
9 items • Updated
How to use sahilchachra/Tessera-4B-Preview-optiq-5bpw-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Tessera-4B-Preview-optiq-5bpw-mlx sahilchachra/Tessera-4B-Preview-optiq-5bpw-mlx
MLX quantization of sahilchachra/Tessera-4B-Preview for Apple Silicon.
Variant: OptiQ mixed-precision (target 5.0 bpw)
Disk size: 3094 MB
Quantized by: sahilchachra
Unlike uniform 4-bit quantization (which forces every layer onto the same bit grid and often collapses reasoning at low bit widths), this model was quantized with mlx-optiq using per-layer KL-sensitivity analysis:
lm_head, and a few middle blocks) get more bits; layers that tolerate aggressive quantization get fewer.optiq_mixed_precision (mlx-optiq)uniform_4bit248 quantizable components total. OptiQ allocated bits non-uniformly based on KL sensitivity:
| Bits | Components | Share |
|---|---|---|
| 8-bit | 50 | 20.2% |
| 6-bit | 113 | 45.6% |
| 4-bit | 85 | 34.3% |
| Total | 248 | 100.0% |
Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.
| This model | Naive 8-bit | FP16 baseline | |
|---|---|---|---|
| Decode tok/s (avg, long traces) | 76.51 | 57.3 | 31.96 |
| Peak memory (GB) | 3.776 | 5.017 | 8.82 |
| Disk size (MB) | 3094 | 4282 | 8907 |
| Benchmark | This model | Naive 8-bit | FP16 baseline | n |
|---|---|---|---|---|
| MATH-500 (math reasoning) | 83.3% (answered 29/30) | 86.7% (answered 30/30) | 86.7% (answered 28/30) | 30 |
| IFEval (instruction following) | 33.3% | 30.0% | 30.0% | 30 |
| GSM8K (math, accuracy) | 96.7% | 96.7% | 93.3% | 30 |
| HumanEval (code, pass@1) | 90.0% | 90.0% | 86.7% | 30 |
| MMLU (knowledge, accuracy) | 63.3% | 63.3% | 56.7% | 30 |
| Level | This model | Naive 8-bit | FP16 baseline |
|---|---|---|---|
| level 1 | 83.3% | 83.3% | 83.3% |
| level 2 | 100.0% | 100.0% | 100.0% |
| level 3 | 66.7% | 66.7% | 66.7% |
| level 4 | 83.3% | 83.3% | 83.3% |
| level 5 | 83.3% | 100.0% | 100.0% |
| Context length | Decode tok/s |
|---|---|
| ~128 tokens | 77.9 |
| ~256 tokens | 77.7 |
| ~512 tokens | 77.8 |
| ~1024 tokens | 77.6 |
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("sahilchachra/tessera-4b-optiq-5bpw-mlx")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)
| Model | Variant |
|---|---|
| sahilchachra/tessera-4b-4bit-mlx | Affine int4 |
| sahilchachra/tessera-4b-8bit-mlx | Affine int8 |
| sahilchachra/tessera-4b-mxfp4-mlx | Block float MX FP4 |
| sahilchachra/tessera-4b-mxfp8-mlx | Block float MX FP8 |
| sahilchachra/tessera-4b-optiq-5bpw-mlx | OptiQ mixed-precision (target 5.0 bpw) ← this model |
See sahilchachra/Tessera-4B-Preview for full model details and intended use.
4-bit