Text Generation
MLX
Safetensors
English
Korean
Japanese
solar_open2
solar
solar-open2
Mixture of Experts
quantized
4bit
conversational
4-bit precision
Instructions to use TensorFold/Solar-Open2-250B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/Solar-Open2-250B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("TensorFold/Solar-Open2-250B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TensorFold/Solar-Open2-250B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Solar-Open2-250B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TensorFold/Solar-Open2-250B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use TensorFold/Solar-Open2-250B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "TensorFold/Solar-Open2-250B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "TensorFold/Solar-Open2-250B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TensorFold/Solar-Open2-250B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use TensorFold/Solar-Open2-250B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Solar-Open2-250B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TensorFold/Solar-Open2-250B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TensorFold/Solar-Open2-250B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Solar-Open2-250B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TensorFold/Solar-Open2-250B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Add Mac memory chooser and reproducible demo prompt
Browse files
README.md
CHANGED
|
@@ -19,12 +19,16 @@ tags:
|
|
| 19 |
- 4bit
|
| 20 |
---
|
| 21 |
|
|
|
|
| 22 |
# Solar-Open2-250B-MLX-4bit
|
| 23 |
|
|
|
|
| 24 |
Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B), converted for Apple Silicon / MLX workflows.
|
| 25 |
|
|
|
|
| 26 |
## Details
|
| 27 |
|
|
|
|
| 28 |
- Source model: `upstage/Solar-Open2-250B`
|
| 29 |
- Quantization: 4-bit affine, group size 64
|
| 30 |
- Local size: 131G
|
|
@@ -32,10 +36,13 @@ Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Ope
|
|
| 32 |
- Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
|
| 33 |
- Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
|
| 34 |
|
|
|
|
| 35 |
## Important runtime notes
|
| 36 |
|
|
|
|
| 37 |
Solar Open2 is not yet a stock `mlx-lm` architecture in many installs. This repo includes `solar_open2.py`; launch with `--trust-remote-code` when serving or loading from Hugging Face.
|
| 38 |
|
|
|
|
| 39 |
```bash
|
| 40 |
mlx_lm.server \
|
| 41 |
--model Vontra/Solar-Open2-250B-MLX-4bit \
|
|
@@ -47,30 +54,73 @@ mlx_lm.server \
|
|
| 47 |
--max-tokens 32768
|
| 48 |
```
|
| 49 |
|
|
|
|
| 50 |
You may see a `transformers` warning that mentions loading `model_type=solar_open2` into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
|
| 51 |
|
|
|
|
| 52 |
The tokenizer template uses Solar/Whale-style tool markers such as `<|tool_call:start|>` and `<|tool_arg:start|>`. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured `tool_calls`. Plain text generation does not need this parser.
|
| 53 |
|
|
|
|
| 54 |
## Use with MLX
|
| 55 |
|
|
|
|
| 56 |
This repo includes a small `solar_open2.py` MLX loader because upstream `mlx-lm` does not yet ship native Solar Open 2 support.
|
| 57 |
|
|
|
|
| 58 |
```bash
|
| 59 |
pip install -U mlx-lm
|
| 60 |
```
|
| 61 |
|
|
|
|
| 62 |
```python
|
| 63 |
from mlx_lm import load, generate
|
| 64 |
|
|
|
|
| 65 |
model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
|
| 66 |
prompt = "Write a short Python function that validates an IPv4 CIDR string."
|
| 67 |
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
|
| 68 |
```
|
| 69 |
|
|
|
|
| 70 |
## Notes
|
| 71 |
|
|
|
|
| 72 |
This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
|
| 73 |
|
|
|
|
| 74 |
## License
|
| 75 |
|
|
|
|
| 76 |
The source model is released under the Upstage Solar License. A copy is included in `LICENSE`. Please review the upstream model card and license before use or redistribution.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
- 4bit
|
| 20 |
---
|
| 21 |
|
| 22 |
+
|
| 23 |
# Solar-Open2-250B-MLX-4bit
|
| 24 |
|
| 25 |
+
|
| 26 |
Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B), converted for Apple Silicon / MLX workflows.
|
| 27 |
|
| 28 |
+
|
| 29 |
## Details
|
| 30 |
|
| 31 |
+
|
| 32 |
- Source model: `upstage/Solar-Open2-250B`
|
| 33 |
- Quantization: 4-bit affine, group size 64
|
| 34 |
- Local size: 131G
|
|
|
|
| 36 |
- Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
|
| 37 |
- Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
|
| 38 |
|
| 39 |
+
|
| 40 |
## Important runtime notes
|
| 41 |
|
| 42 |
+
|
| 43 |
Solar Open2 is not yet a stock `mlx-lm` architecture in many installs. This repo includes `solar_open2.py`; launch with `--trust-remote-code` when serving or loading from Hugging Face.
|
| 44 |
|
| 45 |
+
|
| 46 |
```bash
|
| 47 |
mlx_lm.server \
|
| 48 |
--model Vontra/Solar-Open2-250B-MLX-4bit \
|
|
|
|
| 54 |
--max-tokens 32768
|
| 55 |
```
|
| 56 |
|
| 57 |
+
|
| 58 |
You may see a `transformers` warning that mentions loading `model_type=solar_open2` into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
|
| 59 |
|
| 60 |
+
|
| 61 |
The tokenizer template uses Solar/Whale-style tool markers such as `<|tool_call:start|>` and `<|tool_arg:start|>`. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured `tool_calls`. Plain text generation does not need this parser.
|
| 62 |
|
| 63 |
+
|
| 64 |
## Use with MLX
|
| 65 |
|
| 66 |
+
|
| 67 |
This repo includes a small `solar_open2.py` MLX loader because upstream `mlx-lm` does not yet ship native Solar Open 2 support.
|
| 68 |
|
| 69 |
+
|
| 70 |
```bash
|
| 71 |
pip install -U mlx-lm
|
| 72 |
```
|
| 73 |
|
| 74 |
+
|
| 75 |
```python
|
| 76 |
from mlx_lm import load, generate
|
| 77 |
|
| 78 |
+
|
| 79 |
model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
|
| 80 |
prompt = "Write a short Python function that validates an IPv4 CIDR string."
|
| 81 |
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
|
| 82 |
```
|
| 83 |
|
| 84 |
+
|
| 85 |
## Notes
|
| 86 |
|
| 87 |
+
|
| 88 |
This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
|
| 89 |
|
| 90 |
+
|
| 91 |
## License
|
| 92 |
|
| 93 |
+
|
| 94 |
The source model is released under the Upstage Solar License. A copy is included in `LICENSE`. Please review the upstream model card and license before use or redistribution.
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
<!-- vontra-chooser-start -->
|
| 99 |
+
## Choose for your Mac
|
| 100 |
+
|
| 101 |
+
[64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b)
|
| 102 |
+
|
| 103 |
+
No measured memory tier is assigned here. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
|
| 104 |
+
|
| 105 |
+
### Runtime and evidence
|
| 106 |
+
|
| 107 |
+
The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results.
|
| 108 |
+
|
| 109 |
+
### Quick start and demo prompt
|
| 110 |
+
|
| 111 |
+
```bash
|
| 112 |
+
hf download Vontra/Solar-Open2-250B-MLX-4bit --local-dir ./models/Solar-Open2-250B-MLX-4bit
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
|
| 116 |
+
|
| 117 |
+
Try this in a new chat with a 128-token output limit:
|
| 118 |
+
|
| 119 |
+
```text
|
| 120 |
+
Explain why the sky looks blue in three short sentences.
|
| 121 |
+
```
|
| 122 |
+
|
| 123 |
+
This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
|
| 124 |
+
|
| 125 |
+
[Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra)
|
| 126 |
+
<!-- vontra-chooser-end -->
|