--- license: apache-2.0 base_model: BAAI/AREX-2 base_model_relation: quantized library_name: vllm pipeline_tag: image-text-to-text tags: - qwen3_5 - arex-2 - compressed-tensors - w4a16 - int4 - quantized - vllm - agent - reasoning - tool-use - vision - long-context ---
AREX

AREX-2 27B · INT4 W4A16

Symmetric group-128 weight-only quantization for vLLM.

Base model · Paper · Project · vLLM

Format W4A16 Weights INT4 Activations BF16 License Apache 2.0

> [!NOTE] > This is a numerical quantization of [BAAI/AREX-2](https://huggingface.co/BAAI/AREX-2), not a fine-tune. It was exported from official revision `d4e3502f92d9e889031c2d04e387ab2eb3520268`. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT4 checkpoint. ## Model summary AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the [upstream model card](https://huggingface.co/BAAI/AREX-2) for architecture, capabilities, evaluation protocols, and usage guidance. ## Dual RTX 3090 deployment **Validation scope:** the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested. | Runtime property | Tested profile | |---|---| | GPUs | 2×RTX 3090 (Ampere `sm_86`) | | Tensor parallelism | 2 | | Activation dtype | BF16 | | KV cache | FP8 E4M3 | | CPU KV offload | 28 GiB, using the deployment's vLLM backport | | Maximum model length | 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens | | Maximum active sequences | 1 | | Optional speculation | External DFlash2 W4A16 drafter; not included in this repository | ### Memory and capacity | Metric | Result | |---|---:| | Indexed tensor payload | 18.60 GB / 17.33 GiB | | vLLM GPU memory after short probes | 16,784 MiB used and 7,394 MiB free per GPU | | Long-context capacity | 128K prompt requests benchmarked; the configured 262,144-token limit was not tested | GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit. ### Measured serving performance The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile. | Prompt tokens | New tokens | Median TTFT (s) | Prefill (tok/s) | Decode median (tok/s; min–max) | |---:|---:|---:|---:|---:| | 1,027 | 512 | 0.69 | 1479 | 225.6 (191.3–260.0) | | 1,028 | 1,024 | 0.71 | 1450 | 149.0 (88.0–209.9) | | 8,194 | 512 | 5.73 | 1430 | 109.8 (75.9–143.7) | | 8,196 | 1,024 | 5.79 | 1416 | 266.2 (253.9–278.4) | | 32,002 | 512 | 25.17 | 1273 | 66.4 (65.2–67.7) | | 32,002 | 1,024 | 28.71 | 1115 | 125.0 (63.3–186.8) | | 64,002 | 512 | 74.38 | 861 | 78.5 (56.7–100.3) | | 64,002 | 1,024 | 77.99 | 821 | 89.8 (57.6–122.0) | | 128,002 | 512 | 180.24 | 710 | 42.6 (39.3–45.9) | | 128,003 | 1,024 | 183.37 | 698 | 84.0 (81.6–86.4) | ## Quantization fidelity This export uses data-free, symmetric round-to-nearest (RTN) quantization. The aggregate relative L2 reconstruction error over the quantized weights is **0.12160**. This is a weight-space diagnostic only; it is not perplexity, logit/KLD agreement, or a task-quality score. No calibration dataset, behavioral evaluation, or quality comparison against the BF16 source has been run. ## Checkpoint profile | Property | Value | |---|---| | Base checkpoint | `BAAI/AREX-2`, revision `d4e3502f92d9e889031c2d04e387ab2eb3520268` | | Quantization | Symmetric INT4 W4A16, group size 128, RTN; no calibration | | Quantized modules | 400 two-dimensional Linear weights | | Runtime format | `compressed-tensors` / `pack-quantized` | | Kernel dispatch | `CompressedTensorsWNA16` → `MarlinLinearKernel` in the tested runtime | | Scale storage | FP16 | | Preserved source precision | Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors | | Runtime | vLLM with `compressed-tensors` support; this is not a GGUF checkpoint | The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately. ## Quantization design | Component | Treatment | |---|---| | 400 eligible two-dimensional Linear weights | Symmetric INT4, group size 128, round-to-nearest | | Vision tower, embeddings, output head, recurrent GDN gates | Retained at source precision | | Other non-Linear and excluded tensors | Retained at source precision | The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model. ## Why W4A16 W4A16 stores the selected Linear weights in 4-bit integers with group-wise scales while keeping activations at 16-bit precision. The smaller tensor payload does not guarantee higher inference speed or equivalent quality. Benchmark on the intended hardware and workload before choosing this variant for production. ## Operational notes | Question | Guidance | |---|---| | Which runtime is validated? | The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation. | | Can llama.cpp load this checkpoint? | No. It uses the vLLM `compressed-tensors` format, not GGUF. | | Is vision validated? | The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured. | | Is 262K serving validated? | No. The source model's context setting is retained, but long-context capacity and quality were not tested. | | Is DFlash2 included? | No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets. | ## Files and provenance The repository contains the sharded SafeTensors checkpoint, `model.safetensors.index.json`, model/tokenizer configuration, and the upstream Apache-2.0 `LICENSE`. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index. ## Acknowledgements and license This repository repackages numerical weights derived from **BAAI/AREX-2**. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.