hancheolp's picture
Upload Solar Open2 Nota NVFP4
7a4d07f verified
|
Raw
History Blame Contribute Delete
5.49 kB
metadata
base_model: upstage/Solar-Open2-250B
base_model_relation: quantized
library_name: vllm
pipeline_tag: text-generation
language:
  - en
  - ko
  - ja
tags:
  - quantization
  - nvfp4
  - vllm
  - moe
  - nota

Solar Open2 250B — Nota NVFP4

Nota AI presents a 4-bit quantized release of Upstage's Solar Open2 250B, produced with Nota AI's proprietary quantization technology specialized for Mixture-of-Experts (MoE) large language models.

Highlights

  • NVFP4 (4-bit float, W4A4)group_size=16, packed in the llm-compressor (compressed-tensors) format for direct serving in vLLM. Both weights and activations are quantized to 4-bit floating point.
    • Requires NVIDIA Blackwell. NVFP4 relies on the FP4 tensor cores introduced in the Blackwell architecture (e.g. B200 / GB200), so inference must run on a Blackwell-class GPU. Earlier architectures (Hopper, Ada, Ampere) do not support NVFP4 execution.

  • Nota AI's proprietary MoE quantization framework. This release is built upon a suite of techniques developed by Nota AI to preserve model quality under aggressive low-bit quantization of MoE architectures:
    • A MoE-specialized calibration-dataset construction method, which achieved 1st place across all tracks at the NVIDIA Nemotron Hackathon.
    • DREAM-MoE and SRA-MoE, two quantization algorithms proposed by Nota AI (published at the ICML 2026 Workshop on AdaptFM), which preserve MoE routing decisions and align expert-routing behavior throughout the quantization process.

License

This repository contains both model weights and code, which are licensed under different terms:

  1. MODEL WEIGHTS (*.safetensors) The Upstage Solar License is based on the permissive Apache License 2.0 with specific mandatory conditions to support the Solar ecosystem.

    Key requirements for Derivative AI Models (create / train / fine-tune / distill / improve using Solar Open 2):

    • Naming: prefix your model name with "Solar" (e.g., Solar-MyModel-v1).
    • Attribution: prominently display "Built with Solar" in related public-facing materials.
    • Notice: include a copy of the Upstage Solar License with your derivative model.
  2. CODE (*.py, *.json, *.jinja files) Licensed under Apache License 2.0 See: https://www.apache.org/licenses/LICENSE-2.0

Performance

Weight footprint

Precision Weight footprint
BF16 500.6 GB
Nota NVFP4 153.3 GB

Benchmarks

Benchmark BF16 Nota NVFP4
Tau2-Bench 75.20 75.08
HLE 27.88 27.66
GPQA Diamond 86.26 85.45
IFBench 80.00 81.02
LiveCodeBench (v5–v6) 87.03 88.55
MMLU-Pro 86.19 86.15
AIME 2026 (EN) 95.67 96.67
IFEval (EN) 94.09 92.61
HMMT 92.05 90.15
KMMLU-Pro 78.38 77.93
HAE-RAE Bench v1.1 73.84 72.84
AIME (KO) 97.67 97.00
KBL 75.51 75.40
KBank-MMLU 80.80 80.68
KorMedMCQA 92.99 93.05
Avg. 81.57 81.35

Usage

This model is packed in the NVFP4 (compressed-tensors) format and can be served directly with vLLM on a Blackwell-class GPU:

vllm serve nota-ai/Solar-Open2-250B-Nota-NVFP4 \
  --tensor-parallel-size 4 \
  --trust-remote-code
  • Set --tensor-parallel-size according to the number of GPUs available in your serving environment.

  • For reasoning and tool-calling, use the reasoning parser and tool parser provided in the original upstage/Solar-Open2-250B model card.

  • See the original model card for the prompt format, parser configuration, and further details.

Citation

@inproceedings{park2026dreammoe,
  title     = {{DREAM-MoE}: Downstream Routing Error-Aware Margin-Preserving Quantization for Mixture-of-Experts Large Language Models},
  author    = {Park, Hancheol and Lee, Geonho and Kim, Tae-Ho},
  booktitle = {ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)},
  year      = {2026},
  url       = {https://openreview.net/forum?id=Wyhqwjl51A},
}

@inproceedings{lee2026sramoe,
  title     = {{SRA-MoE}: Output-Aware Selective Router Alignment for MoE Quantization},
  author    = {Lee, Geonho and Park, Hancheol and Lee, Seunghyun and Choi, Jungwook and Kim, Tae-Ho},
  booktitle = {ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)},
  year      = {2026},
  url       = {https://openreview.net/forum?id=H0NoX02erJ},
}