Qwen3.8-Flash-Next W4 G128 for Ascend 300I

This is an experimental text-serving checkpoint for Ascend 310P. It requires the pinned OpenSensor vLLM fork, the OpenSensor vLLM Ascend development runtime, and compiled custom operators. It is not a drop-in model for stock vLLM Ascend.

This is a complete four-bit checkpoint derived from Qwen/Qwen3.8-Flash-Next at revision de4b8e4d43b917e7706784d8bb445c9af86a3540. It targets two Atlas 300I Duo cards, exposed as four Ascend 310P3 devices.

The checkpoint stores routed-expert weights as packed signed INT4 with asymmetric groups of 128. Runtime policy determines the activation path:

Model policy Selected expert computation
W4 with activation quantization permitted Native W4A8: per-group INT8 activation quantization and packing, then INT4 x INT4 to INT32
W4 requiring FP16 activations Existing W4A16 backend
Companion W8A8 Dynamic Existing W8A8 backend, unchanged

Selection uses checkpoint quantization metadata, hardware support, and supported shapes before graph capture. Native W4A8 is explicit; the runtime does not silently change activation precision.

Quantization and checkpoint format

Ascend ModelSlim IR converted 73,728 routed-expert projections from the original BF16 checkpoint using asymmetric per-group min/max RTN. Each group contains 128 weights. Two signed four-bit values are packed into each byte, with the low nibble first; scales are FP16 and zero points are INT8. The checkpoint advertises qwen4exp_w4a16_group_v1 in text_config.ascend_expert_quantization.

Router, shared experts, attention, Gated DeltaNet, embeddings, the PLE table, vision components, and MTP weights remain floating-point where present. Floating BF16 source tensors were exported as FP16 for the 310P serving path; FP32 source tensors remain FP32.

The complete export contains 1,610 Safetensors shards and 222,746 tensors. Tensor payload is 181,637,185,528 bytes (169.16 GiB); the model directory with metadata is 181,713,896,545 bytes. The approximately 102.40 GB PLE table is included in that total. See provenance for source, conversion, byte accounting, and validation details.

Native W4A8 result

The retained native backend keeps expert weights packed as INT4 and fuses activation quantization and INT4 packing on device. The c1 decode shape uses a compile-time streamed-weight schedule; c2-c4 instantiate the resident-weight schedule. This split avoids the aggregate-throughput regression seen when the streamed pipeline was compiled into every shape.

Measurements used the real checkpoint, two Atlas 300I Duo cards, TP4 with expert parallelism, MTP2, FULL_DECODE_ONLY graphs, a 262,144-token serving limit, four scheduler slots, 512 batched tokens, and fixed worker CPU affinity. Warmup runs are excluded from the medians.

Workload Resident baseline Retained native W4A8 Change
c1 decode, median 28.1735 tok/s 29.8610 tok/s +5.99%
c1 decode, peak - 30.3191 tok/s -
c4 aggregate decode, median 59.6759 tok/s 59.5233 tok/s -0.26%

Candidate c1 runs were 29.9512, 30.3191, 29.8610, 29.7040, and 29.6911 tok/s. Candidate c4 aggregate runs were 62.1914, 59.5233, and 58.2283 tok/s. The c4 difference is within observed run-to-run variance.

The fixed 228-question zero-shot quality gate scored 206/228 with zero invalid answers. This is a fixed sampled gate, not an official full-dataset MMLU score. Across the validation service run, interval counters reported 7,412 accepted speculative tokens from 9,078 drafted tokens (81.65%). Five real-NPU changing-input graph tests passed, including c1 gate/down, 120-row c4 fallback for both projection widths, and fused down with changing inputs, routes, and weights.

Machine-readable results are in native-w4a8-results-20260928.json. Implementation and full test notes are in opensensor/vllm-ascend@d4d2c42f1 and its WEIGHT-PIPELINE-C1-RESULTS.md.

Grouped MTP update

Revision opensensor/vllm-ascend@17954180d adds an explicit mtp_expert_execution=w8a8_grouped mode. It quantizes the MTP checkpoint's routed experts once at load time, stores the local expert banks in CANN FRACTAL_NZ format, and runs both draft expert projections through the 310P grouped W8A8 path. The established w8a16_routed draft path remains the default, so this does not silently change MTP activation precision.

On the same four-device service configuration, five c1 runs measured 29.7374 tok/s median, 30.8504 tok/s mean, and 34.3382 tok/s peak. Three c4 trials measured 61.3700, 60.8789, and 59.4403 aggregate tok/s, for a 60.8789 tok/s median. This is a 2.28% gain over the retained native-W4A8 c4 median while c1 changed by -0.41%, within run-to-run variation. All four c4 requests ran concurrently with none queued.

The candidate passed model loading, memory profiling, graph warmup, a coherent generation smoke test, and the measured runs. Focused host coverage passed 71 tests with one skip. The earlier 206/228 quality gate was not rerun for this isolated draft-path change. Raw results and test notes are in qwen-next-gains-20260930.

September 30 PLE frontend update

Revision opensensor/vllm-ascend@d4d2c42f1 removes the decode-time NPU-to-host synchronization from PLE n-gram hashing, reuses a persistent parallel host gather pool, and adds the fused 310P PLE gate/short-convolution epilogue. Hashing reads the runner's existing CPU token mirrors; no additional token transfer is introduced.

On real checkpoint tensors, the exact one-token PLE frontend fell from 1.423 ms to 0.470 ms. Its components measured 0.024 ms for host hashing, 0.312 ms for hot host gather, 0.032 ms for pageable host-to-device copy, and 0.356 ms for the 2560-to-12800 FP16 projection. The transfer is already small enough that pinning it is not justified by these measurements.

The same revision adds an opt-in dynamic-W8A8 PLE projection. It must be selected explicitly with ascend_expert_quantization.ple_projection_execution=w8a8_dynamic; FP16 remains the default, so the runtime never silently changes PLE activation precision. Including activation quantization, the projection measured 0.201 ms versus 0.356 ms for FP16. Relative L2 error was 0.58% and cosine similarity was 0.999983 at the measured one-token shape.

Paired whole-model c1 trials used five 512-token completions. Dynamic W8A8 measured 27.16 tok/s median versus 26.99 tok/s for the exact projection. MTP acceptance changed between outputs, so target-step time is the clearer kernel measure: 85.18 ms versus 89.48 ms, a 4.81% reduction. The W8A8 trials ranged from 26.24 to 32.57 tok/s. Six c4 runs measured 57.96 tok/s aggregate median; the last three warm runs measured 58.57 tok/s median and peaked at 59.39 tok/s. A changing-input graph replay returned the correct 17 x 19 = 323 result.

Machine-readable stage and whole-model evidence is in the commit's qwen-ple-frontend-20260930 artifact directory. These short-prompt measurements do not validate full-window accuracy or throughput.

The earlier W4A16 runtime study measured 14.71 tok/s at short context and 14.60 tok/s near 23.4K context. Those historical results remain in benchmark-summary-20260927.json; they are not directly comparable to the later native-W4A8 service because the runtime and serving configuration changed.

October 2 GPQA Diamond end-to-end result

The native W4A8 service completed all 198 GPQA Diamond questions on two Atlas 300I Duo cards (four Ascend 310P3 chips, TP4/EP4). The matched reference was one RTX 6000 Pro running the same base model as an UD-IQ4_XS GGUF in llama.cpp. Both used the same serialized AISBench zero-shot chain-of-thought prompts (dataset version b1ed2c), temperature 0, seed 1024, top-k 20, top-p 1, and an 8,192-token output ceiling. The servers' chat templates and quantization formats differ, so this is a serving-stack comparison rather than an isolated hardware or quantization test.

Measure Ascend W4 RTX IQ4_XS
AISBench-style Answer: X score 140/198 (70.71%) 140/198 (70.71%)
Responses with a final answer 117/198 124/198
Correct answers in final response text 114/198 (57.58%) 120/198 (60.61%)
Hit the 8,192-token output limit 81/198 74/198
Total output tokens 949,014 957,402

The benchmark extractor scans the combined response for the last Answer: X. It therefore credited 26 Ascend and 20 RTX cases where a correct letter appeared in unfinished reasoning but generation stopped before any final answer. The final-response row above is a separate completion diagnostic, not the configured benchmark score. When both services did finish an answer for a question, all 112 letters agreed (109 correct, three wrong). Chemistry caused most output-limit hits: 62/93 on Ascend and 57/93 on RTX.

The final 106 Ascend cases used MTP2 with FULL_DECODE_ONLY C3/C4 graphs captured at [9, 12] and four concurrent requests. On those same 106 cases, median client-observed output rate was 12.50 tokens/s on Ascend versus 61.72 tokens/s on RTX with three slots. Ascend delivered 48.93 aggregate output tokens/s over that 2 h 52 m 46 s phase. The RTX full run delivered 183.11 aggregate output tokens/s over 1 h 27 m 9 s; those wall rates cover different case windows. The Ascend phase completed without an eager decode fallback, zero-acceptance interval, or server error. Earlier Ascend cases were resumed across server changes, so no single uninterrupted whole-run Ascend wall rate is claimed.

The measured development runtime also synchronizes pending recurrent-state NPU writes before a prefix-Mamba state slot is evicted or reused. The GPQA client saves each completed case and resumes missing IDs; two prior timeout rows were successfully replayed. The healthy final phase does not isolate which reliability change prevented the earlier degraded server behavior.

This run does not evaluate the companion W8A8 checkpoint, full-context accuracy, multimodal input, or MTP-versus-target-only quality parity.

October 5 prefill and profiling update

The native-W4 development service now has an explicitly selected 2,560-token expert chunk, built-in FP16 SwiGLU, and the CANN route finalizer. The native pack and matmul operators support 25,600 top-10 routed rows, and the projection schedule retains the smaller row tile through this batch size. The source chunk default remains 1,536 tokens; selecting 2,560 requires a matching rebuilt OPP package. The CANN finalizer remains opt-in because its rounding differs from the torch finalizer. Model weights and checkpoint format are unchanged.

Three matched 23,410-token cold prompts on TP4/EP4, MTP2 and decode graphs [3, 6] produced the following preliminary service results:

Configuration Mean cold TTFT Effective prompt rate
1,536-token expert chunks, built-in SwiGLU, CANN finalizer 64.292 s 364.1 tok/s
2,560-token expert chunks, revised projection schedule, same activation/finalizer 62.032 s 377.4 tok/s

This is a 3.52% reduction in mean TTFT across three cases. Every request reported zero cached tokens. Two 32-token outputs matched exactly; one opening phrase changed. The earlier run used a different service instance and day, so run order and device conditions were not fully controlled. A one-layer 23,410-token partial improved from 457.74 to 420.97 ms with exact output parity; it excludes attention, shared experts, TP reduction and scheduling.

An October 5 resident-server replay of the saved case-1 prompt measured 62.392 s TTFT versus its previous 61.840 s, with zero cached tokens and identical output. A short-context 512-token decode observation measured 30.91 tok/s; replaying that request measured 31.87 tok/s, but later output wording differed. These are individual observations, not a demonstrated speedup in decode or a guarantee at long context. Separate conversational requests around 22K–26K context measured approximately 26 tok/s decode.

The service reported 1,068,936 cache tokens and 4.08x planner concurrency at 262,144 tokens per request, and captured both decode graphs. This is startup capacity accounting, not a measured four-request full-window workload. Sustained thermal behavior, concurrent-request latency and broader quality remain unqualified for the latest prefill configuration. The October 2 GPQA result above was obtained with the earlier serving configuration and was not rerun for this prefill update.

Four-rank named profiling measured approximately 95.7% device occupancy. Native W4 projections, QSA K/V gathers and large hyperconnection casts remain major targets. Long host copy calls overlapped active device execution almost entirely, so their host duration cannot be counted as freely removable latency. The published trace PNGs show both the task timeline and measured occupancy.

Two further candidates remain experimental and are not serving defaults:

  • A separately named sparse QSA gather avoids fallback cache reads for masked groups. It passed 32 NPU correctness cases and produced exact full-attention outputs. Selecting it only for sparse tiles reduced first-chunk synthetic parallel-gather attention from 192.566 to 188.835 ms (1.94%); dense tiles retain the existing operator. Real multi-request/prefix dispatch and model TTFT remain unqualified.
  • A resident Python candidate reuses the FP16 normalized hyperconnection projection operand, removing a repeated cast while retaining FP32 mixing math. Its host tests passed, but NPU performance and serving recapture are pending. It is inactive in the measured service.

The implementation and evidence are published at opensensor/vllm-ascend@ffa980c1d0ac9f31cd1dbbd06cd4bd3286571032. See the batching results and trace figures, runtime runbook, and resident experiment status.

Serving example

The October 5 measured prefill profile uses these software revisions:

Build and install the 310P custom operators from that vLLM Ascend revision, activate the matching CANN environment, and place that plugin ahead of any stock installation. The measured host also used the project CPU-affinity helper; core IDs depend on the host topology. The fixed cache value below is specific to the qualified four-device compact-state layout; use it only with the matching runtime and operator package described in the runbook.

export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
export VLLM_ASCEND_KV_CACHE_FRACTION=0.80
export VLLM_USE_BREAKABLE_CUDAGRAPH=1
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=3000
export TASK_QUEUE_ENABLE=1
export OMP_NUM_THREADS=1

vllm serve matteiuspi/Qwen3.8-Flash-Next-W4A16-G128-300i \
  --served-model-name qwen38-w4-batch2560-cann-finalize-candidate \
  --dtype float16 \
  --tensor-parallel-size 4 \
  --no-async-scheduling \
  --disable-custom-all-reduce \
  --max-model-len 262144 \
  --max-num-batched-tokens 2560 \
  --max-num-seqs 4 \
  --gpu-memory-utilization 0.965 \
  --kv-cache-memory 88673894400 \
  --language-model-only \
  --enable-expert-parallel \
  --enable-ep-weight-filter \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --mamba-cache-mode align \
  --enable-chunked-prefill \
  --enable-prefix-caching \
  --enable-prompt-tokens-details \
  --speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
  --compilation-config '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[3,6]}' \
  --hf-overrides '{"text_config":{"ascend_expert_quantization":{"backend":"cube_310_int4_a8","activation_quantization":"int8_per_group","grouped_activation":"cann_builtin_fp16","grouped_finalize":"cann_v2","grouped_prefill_chunk_tokens":2560,"bits":4,"format":"qwen4exp_w4a16_group_v1","group_size":128,"lm_head_execution":"w8a8_dynamic","ple_projection_execution":"w8a8_dynamic","offset_dtype":"int8","packing":"signed_int4_low_nibble_first_in_axis","scale_dtype":"float16","shared_expert_execution":"tp_sharded","symmetric":false}}}' \
  --limit-mm-per-prompt '{"image":0,"video":0}'

Do not add --quantization ascend: model-specific metadata selects the W4 loader, and the W4 loader rejects that generic override. Removing the native backend override keeps the checkpoint's FP16-activation W4A16 path.

Validation scope and limitations

  • The public source revision and custom operator implementation are pinned, but users must build the 310P operators locally; stock vLLM Ascend does not contain this backend.
  • The 262,144-token value is a configured serving limit. Full-window accuracy and throughput have not been validated.
  • Four scheduler slots were active during aggregate testing; four simultaneous full 262K windows were not filled or measured.
  • The 228-question gate is useful regression evidence, not a broad substitute for GPQA, coding, or official full-dataset MMLU evaluation.
  • Multimodal input and flashcomm1 are not validated for this configuration.
  • Cold prefill remains workload-sensitive.

The model weights are distributed under the Qwen Community License 1.0, inherited from the base model. Review its terms before use. The vLLM Ascend code is separate from these weights and retains its own license.

Please cite the original Qwen model and technical report when using this checkpoint.

Downloads last month
767
Safetensors
Model size
121B params
Tensor type
F16
·
I8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for matteiuspi/Qwen3.8-Flash-Next-W4A16-G128-300i

Quantized
(365)
this model

Collection including matteiuspi/Qwen3.8-Flash-Next-W4A16-G128-300i