ashxhart commited on
Commit
f321918
·
verified ·
1 Parent(s): c444e5d

Add Mac memory chooser and reproducible demo prompt

Browse files
Files changed (1) hide show
  1. README.md +50 -0
README.md CHANGED
@@ -19,12 +19,16 @@ tags:
19
  - 4bit
20
  ---
21
 
 
22
  # Solar-Open2-250B-MLX-4bit
23
 
 
24
  Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B), converted for Apple Silicon / MLX workflows.
25
 
 
26
  ## Details
27
 
 
28
  - Source model: `upstage/Solar-Open2-250B`
29
  - Quantization: 4-bit affine, group size 64
30
  - Local size: 131G
@@ -32,10 +36,13 @@ Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Ope
32
  - Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
33
  - Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
34
 
 
35
  ## Important runtime notes
36
 
 
37
  Solar Open2 is not yet a stock `mlx-lm` architecture in many installs. This repo includes `solar_open2.py`; launch with `--trust-remote-code` when serving or loading from Hugging Face.
38
 
 
39
  ```bash
40
  mlx_lm.server \
41
  --model Vontra/Solar-Open2-250B-MLX-4bit \
@@ -47,30 +54,73 @@ mlx_lm.server \
47
  --max-tokens 32768
48
  ```
49
 
 
50
  You may see a `transformers` warning that mentions loading `model_type=solar_open2` into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
51
 
 
52
  The tokenizer template uses Solar/Whale-style tool markers such as `<|tool_call:start|>` and `<|tool_arg:start|>`. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured `tool_calls`. Plain text generation does not need this parser.
53
 
 
54
  ## Use with MLX
55
 
 
56
  This repo includes a small `solar_open2.py` MLX loader because upstream `mlx-lm` does not yet ship native Solar Open 2 support.
57
 
 
58
  ```bash
59
  pip install -U mlx-lm
60
  ```
61
 
 
62
  ```python
63
  from mlx_lm import load, generate
64
 
 
65
  model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
66
  prompt = "Write a short Python function that validates an IPv4 CIDR string."
67
  print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
68
  ```
69
 
 
70
  ## Notes
71
 
 
72
  This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
73
 
 
74
  ## License
75
 
 
76
  The source model is released under the Upstage Solar License. A copy is included in `LICENSE`. Please review the upstream model card and license before use or redistribution.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19
  - 4bit
20
  ---
21
 
22
+
23
  # Solar-Open2-250B-MLX-4bit
24
 
25
+
26
  Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B), converted for Apple Silicon / MLX workflows.
27
 
28
+
29
  ## Details
30
 
31
+
32
  - Source model: `upstage/Solar-Open2-250B`
33
  - Quantization: 4-bit affine, group size 64
34
  - Local size: 131G
 
36
  - Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
37
  - Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
38
 
39
+
40
  ## Important runtime notes
41
 
42
+
43
  Solar Open2 is not yet a stock `mlx-lm` architecture in many installs. This repo includes `solar_open2.py`; launch with `--trust-remote-code` when serving or loading from Hugging Face.
44
 
45
+
46
  ```bash
47
  mlx_lm.server \
48
  --model Vontra/Solar-Open2-250B-MLX-4bit \
 
54
  --max-tokens 32768
55
  ```
56
 
57
+
58
  You may see a `transformers` warning that mentions loading `model_type=solar_open2` into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
59
 
60
+
61
  The tokenizer template uses Solar/Whale-style tool markers such as `<|tool_call:start|>` and `<|tool_arg:start|>`. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured `tool_calls`. Plain text generation does not need this parser.
62
 
63
+
64
  ## Use with MLX
65
 
66
+
67
  This repo includes a small `solar_open2.py` MLX loader because upstream `mlx-lm` does not yet ship native Solar Open 2 support.
68
 
69
+
70
  ```bash
71
  pip install -U mlx-lm
72
  ```
73
 
74
+
75
  ```python
76
  from mlx_lm import load, generate
77
 
78
+
79
  model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
80
  prompt = "Write a short Python function that validates an IPv4 CIDR string."
81
  print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
82
  ```
83
 
84
+
85
  ## Notes
86
 
87
+
88
  This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
89
 
90
+
91
  ## License
92
 
93
+
94
  The source model is released under the Upstage Solar License. A copy is included in `LICENSE`. Please review the upstream model card and license before use or redistribution.
95
+
96
+
97
+
98
+ <!-- vontra-chooser-start -->
99
+ ## Choose for your Mac
100
+
101
+ [64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b)
102
+
103
+ No measured memory tier is assigned here. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
104
+
105
+ ### Runtime and evidence
106
+
107
+ The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results.
108
+
109
+ ### Quick start and demo prompt
110
+
111
+ ```bash
112
+ hf download Vontra/Solar-Open2-250B-MLX-4bit --local-dir ./models/Solar-Open2-250B-MLX-4bit
113
+ ```
114
+
115
+ Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
116
+
117
+ Try this in a new chat with a 128-token output limit:
118
+
119
+ ```text
120
+ Explain why the sky looks blue in three short sentences.
121
+ ```
122
+
123
+ This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
124
+
125
+ [Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra)
126
+ <!-- vontra-chooser-end -->