--- base_model: upstage/Solar-Open2-250B base_model_relation: quantized library_name: vllm pipeline_tag: text-generation language: - en - ko - ja tags: - quantization - nvfp4 - vllm - moe - nota --- # Solar Open2 250B — Nota NVFP4 [Nota AI](https://www.nota.ai/) presents a 4-bit quantized release of [Upstage](https://www.upstage.ai/)'s [Solar Open2 250B](https://huggingface.co/upstage/Solar-Open2-250B), produced with Nota AI's proprietary quantization technology specialized for Mixture-of-Experts (MoE) large language models. ## Highlights - **NVFP4 (4-bit float, W4A4)** — `group_size=16`, packed in the `llm-compressor` (compressed-tensors) format for direct serving in [vLLM](https://github.com/vllm-project/vllm). Both weights and activations are quantized to 4-bit floating point. - > **Requires NVIDIA Blackwell.** NVFP4 relies on the FP4 tensor cores introduced in > the Blackwell architecture (e.g. B200 / GB200), so inference must run on a > Blackwell-class GPU. Earlier architectures (Hopper, Ada, Ampere) do not support > NVFP4 execution. - **Nota AI's proprietary MoE quantization framework.** This release is built upon a suite of techniques developed by Nota AI to preserve model quality under aggressive low-bit quantization of MoE architectures: - A MoE-specialized calibration-dataset construction method, which achieved **1st place across all tracks** at the NVIDIA Nemotron Hackathon. - [**DREAM-MoE**](https://openreview.net/pdf?id=Wyhqwjl51A) and [**SRA-MoE**](https://openreview.net/pdf?id=H0NoX02erJ), two quantization algorithms proposed by Nota AI (published at the *ICML 2026 Workshop on AdaptFM*), which preserve MoE routing decisions and align expert-routing behavior throughout the quantization process. ## Performance ### Weight footprint | Precision | Weight footprint | | ------------------ | :--------------: | | BF16 | 500.6 GB | | W4A4 (NVFP4-Nota) | 153.3 GB | ### Benchmarks | Benchmark | BF16 | W4A4 (NVFP4-Nota) | | ----------------------- | :--------------: | :---------------: | | Tau2-Bench | 75.20 | 75.08 | | HLE | 27.88 | 27.66 | | GPQA Diamond | 86.26 | 85.45 | | IFBench | 80.00 | 81.02 | | LiveCodeBench (v5–v6) | 87.03 | 88.55 | | MMLU-Pro | 86.19 | 86.15 | | AIME 2026 (EN) | 95.67 | 96.67 | | IFEval (EN) | 94.09 | 92.61 | | HMMT | 92.05 | 90.15 | | KMMLU-Pro | 78.38 | 77.93 | | HAE-RAE Bench v1.1 | 73.84 | 72.84 | | AIME (KO) | 97.67 | 97.00 | | KBL | 75.51 | 75.40 | | KBank-MMLU | 80.80 | 80.68 | | KorMedMCQA | 92.99 | 93.05 | | **Avg.** | **81.57** | **81.35** | ## Usage This model is packed in the NVFP4 (compressed-tensors) format and can be served directly with vLLM on a **Blackwell-class GPU**: ```bash vllm serve nota-ai/Solar-Open2-250B-Nota-NVFP4 \ --tensor-parallel-size 8 \ --trust-remote-code ``` See the original model card for prompt format and further details: [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B). ## Citation ```bibtex @inproceedings{park2026dreammoe, title = {{DREAM-MoE}: Downstream Routing Error-Aware Margin-Preserving Quantization for Mixture-of-Experts Large Language Models}, author = {Park, Hancheol and Lee, Geonho and Kim, Tae-Ho}, booktitle = {ICML 2026 Workshop on Adaptive Foundation Models (AdaptFM)}, year = {2026}, url = {https://openreview.net/forum?id=Wyhqwjl51A}, } @inproceedings{lee2026sramoe, title = {{SRA-MoE}: Output-Aware Selective Router Alignment for MoE Quantization}, author = {Lee, Geonho and Park, Hancheol and Lee, Seunghyun and Choi, Jungwook and Kim, Tae-Ho}, booktitle = {ICML 2026 Workshop on Adaptive Foundation Models (AdaptFM)}, year = {2026}, url = {https://openreview.net/forum?id=H0NoX02erJ}, } ```