mabaeyens commited on
Commit
9e7c491
·
verified ·
1 Parent(s): d2823d3

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +30 -7
README.md CHANGED
@@ -19,32 +19,55 @@ extra_gated_description: If you want to learn more about how we process your per
19
  tags:
20
  - mistral-common
21
  - mlx
 
 
 
 
 
 
 
 
22
  pipeline_tag: image-text-to-text
23
  base_model: mistralai/Ministral-3-14B-Base-2512
24
  ---
25
 
26
  # mlx-community/Ministral-3-14B-Base-2512-4bit
27
 
28
- The largest model in the Ministral 3 family, **Ministral 3 14B Base 2512** is a vision-language model: a text backbone paired
29
  with a vision encoder, supporting image understanding alongside text. This is
30
  the **base pre-trained** checkpoint — not instruction- or chat-tuned. For
31
  chat/instruction-following use cases, use the
32
  [Instruct variant](https://huggingface.co/mlx-community/Ministral-3-14B-Instruct-2512-4bit)
33
  instead; this base checkpoint is intended for custom post-training/fine-tuning.
34
 
 
 
 
 
 
 
 
 
35
  This is an MLX conversion of [`mistralai/Ministral-3-14B-Base-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Base-2512),
36
  converted with [mlx-vlm](https://github.com/Blaizzy/mlx-vlm). Refer to the
37
  [original model card](https://huggingface.co/mistralai/Ministral-3-14B-Base-2512) for the full
38
  description, capabilities, and license terms.
39
 
40
- ## Quantization notes
 
 
 
 
 
 
 
 
 
41
 
42
- - Language model layers: 4-bit, group_size=64, affine mode
43
- - Vision tower and multimodal projector are kept at **full precision** only
44
- the language backbone is quantized, per mlx-vlm's standard policy of not
45
- quantizing multimodal modules
46
  - Blended average: **5.416 bits per weight** across all parameters
47
- - Output size on disk: **9.47GB**
48
 
49
  ## Ministral 3 family
50
 
 
19
  tags:
20
  - mistral-common
21
  - mlx
22
+ - ministral
23
+ - ministral-3
24
+ - vision-language
25
+ - multimodal
26
+ - quantized
27
+ - edge
28
+ - 4-bit
29
+ - base-model
30
  pipeline_tag: image-text-to-text
31
  base_model: mistralai/Ministral-3-14B-Base-2512
32
  ---
33
 
34
  # mlx-community/Ministral-3-14B-Base-2512-4bit
35
 
36
+ This is, **Ministral 3 14B Base 2512** is a vision-language model: a text backbone paired
37
  with a vision encoder, supporting image understanding alongside text. This is
38
  the **base pre-trained** checkpoint — not instruction- or chat-tuned. For
39
  chat/instruction-following use cases, use the
40
  [Instruct variant](https://huggingface.co/mlx-community/Ministral-3-14B-Instruct-2512-4bit)
41
  instead; this base checkpoint is intended for custom post-training/fine-tuning.
42
 
43
+ > **Community note.** Structural check confirms the vision tower and
44
+ > multimodal projector were carried over intact (not dropped, which is a real
45
+ > failure mode for text-only conversion tools on vision-language models).
46
+ > Functional check confirms both text-only and image+text generation produce
47
+ > coherent output. Converted and verified by a single maintainer running
48
+ > local MLX tooling -- not independently reviewed by anyone else; please open
49
+ > a discussion if you hit anything unexpected.
50
+
51
  This is an MLX conversion of [`mistralai/Ministral-3-14B-Base-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Base-2512),
52
  converted with [mlx-vlm](https://github.com/Blaizzy/mlx-vlm). Refer to the
53
  [original model card](https://huggingface.co/mistralai/Ministral-3-14B-Base-2512) for the full
54
  description, capabilities, and license terms.
55
 
56
+ ## Heads up
57
+
58
+ - **Base model, not instruct-tuned** — expect raw completion behavior, not
59
+ chat-following. Don't expect it to follow instructions well.
60
+ - **Vision retained at full precision** — only the language backbone is
61
+ quantized; the vision tower and multimodal projector are untouched bf16,
62
+ per mlx-vlm's standard policy of not quantizing multimodal modules.
63
+ - **Output size on disk: 9.47GB**
64
+
65
+ ## Provenance
66
 
67
+ - Source: [`mistralai/Ministral-3-14B-Base-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Base-2512) (BF16)
68
+ - Language model layers: **4-bit** affine quantization, group_size=64
69
+ - Vision tower + multimodal projector: kept at full precision (not quantized)
 
70
  - Blended average: **5.416 bits per weight** across all parameters
 
71
 
72
  ## Ministral 3 family
73