---
license: apache-2.0
base_model: BAAI/AREX-2
base_model_relation: quantized
library_name: vllm
pipeline_tag: image-text-to-text
tags:
- qwen3_5
- arex-2
- compressed-tensors
- w4a16
- int4
- quantized
- vllm
- agent
- reasoning
- tool-use
- vision
- long-context
---
> [!NOTE]
> This is a numerical quantization of [BAAI/AREX-2](https://huggingface.co/BAAI/AREX-2), not a fine-tune. It was exported from official revision `d4e3502f92d9e889031c2d04e387ab2eb3520268`. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT4 checkpoint.
## Model summary
AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the [upstream model card](https://huggingface.co/BAAI/AREX-2) for architecture, capabilities, evaluation protocols, and usage guidance.
## Dual RTX 3090 deployment
**Validation scope:** the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested.
| Runtime property | Tested profile |
|---|---|
| GPUs | 2×RTX 3090 (Ampere `sm_86`) |
| Tensor parallelism | 2 |
| Activation dtype | BF16 |
| KV cache | FP8 E4M3 |
| CPU KV offload | 28 GiB, using the deployment's vLLM backport |
| Maximum model length | 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens |
| Maximum active sequences | 1 |
| Optional speculation | External DFlash2 W4A16 drafter; not included in this repository |
### Memory and capacity
| Metric | Result |
|---|---:|
| Indexed tensor payload | 18.60 GB / 17.33 GiB |
| vLLM GPU memory after short probes | 16,784 MiB used and 7,394 MiB free per GPU |
| Long-context capacity | 128K prompt requests benchmarked; the configured 262,144-token limit was not tested |
GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit.
### Measured serving performance
The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile.
| Prompt tokens | New tokens | Median TTFT (s) | Prefill (tok/s) | Decode median (tok/s; min–max) |
|---:|---:|---:|---:|---:|
| 1,027 | 512 | 0.69 | 1479 | 225.6 (191.3–260.0) |
| 1,028 | 1,024 | 0.71 | 1450 | 149.0 (88.0–209.9) |
| 8,194 | 512 | 5.73 | 1430 | 109.8 (75.9–143.7) |
| 8,196 | 1,024 | 5.79 | 1416 | 266.2 (253.9–278.4) |
| 32,002 | 512 | 25.17 | 1273 | 66.4 (65.2–67.7) |
| 32,002 | 1,024 | 28.71 | 1115 | 125.0 (63.3–186.8) |
| 64,002 | 512 | 74.38 | 861 | 78.5 (56.7–100.3) |
| 64,002 | 1,024 | 77.99 | 821 | 89.8 (57.6–122.0) |
| 128,002 | 512 | 180.24 | 710 | 42.6 (39.3–45.9) |
| 128,003 | 1,024 | 183.37 | 698 | 84.0 (81.6–86.4) |
## Quantization fidelity
This export uses data-free, symmetric round-to-nearest (RTN) quantization. The aggregate relative L2 reconstruction error over the quantized weights is **0.12160**. This is a weight-space diagnostic only; it is not perplexity, logit/KLD agreement, or a task-quality score. No calibration dataset, behavioral evaluation, or quality comparison against the BF16 source has been run.
## Checkpoint profile
| Property | Value |
|---|---|
| Base checkpoint | `BAAI/AREX-2`, revision `d4e3502f92d9e889031c2d04e387ab2eb3520268` |
| Quantization | Symmetric INT4 W4A16, group size 128, RTN; no calibration |
| Quantized modules | 400 two-dimensional Linear weights |
| Runtime format | `compressed-tensors` / `pack-quantized` |
| Kernel dispatch | `CompressedTensorsWNA16` → `MarlinLinearKernel` in the tested runtime |
| Scale storage | FP16 |
| Preserved source precision | Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors |
| Runtime | vLLM with `compressed-tensors` support; this is not a GGUF checkpoint |
The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately.
## Quantization design
| Component | Treatment |
|---|---|
| 400 eligible two-dimensional Linear weights | Symmetric INT4, group size 128, round-to-nearest |
| Vision tower, embeddings, output head, recurrent GDN gates | Retained at source precision |
| Other non-Linear and excluded tensors | Retained at source precision |
The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model.
## Why W4A16
W4A16 stores the selected Linear weights in 4-bit integers with group-wise scales while keeping activations at 16-bit precision. The smaller tensor payload does not guarantee higher inference speed or equivalent quality. Benchmark on the intended hardware and workload before choosing this variant for production.
## Operational notes
| Question | Guidance |
|---|---|
| Which runtime is validated? | The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation. |
| Can llama.cpp load this checkpoint? | No. It uses the vLLM `compressed-tensors` format, not GGUF. |
| Is vision validated? | The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured. |
| Is 262K serving validated? | No. The source model's context setting is retained, but long-context capacity and quality were not tested. |
| Is DFlash2 included? | No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets. |
## Files and provenance
The repository contains the sharded SafeTensors checkpoint, `model.safetensors.index.json`, model/tokenizer configuration, and the upstream Apache-2.0 `LICENSE`. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index.
## Acknowledgements and license
This repository repackages numerical weights derived from **BAAI/AREX-2**. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.