Qwen3-4B-Instruct-2507 β€” fraQtl KV-Cache Compression Sidecars

Model Overview

Description

This is the artifact that serves nine concurrent β‰ˆ128K-context users on one A100 β€” 4.5Γ— the users fp16 holds without preemption and 1.8Γ— FP8's (9 vs 5), at 2Γ— fp16's total decode throughput and 1.14Γ— FP8's peak (134.1 vs 66.6 vs 117.6 tok/s aggregate) β€” with nine distinct passkeys planted at placements spanning 0–100% of the context and 9/9 recovered exactly. The receipt for that sentence, and every other number on this page, is in this repository.

The sidecars are calibrated per-layer, per-KV-head eigenbasis artifacts consumed by the fraqtl runtime (a prebuilt binary vLLM attention backend). Base model weights are not modified or redistributed.

128K concurrency ladder

Batch-1 receipts, three contexts

License/Terms of Use

Sidecars and receipts: Apache 2.0. Base model weights download from Qwen/Qwen3-4B-Instruct-2507 under its own license. The fraqtl runtime wheel is proprietary, free to install and run for verification and evaluation (license).

Deployment Geography

Global.

Use Case

Teams serving LLMs with vLLM who need more concurrent long-context users per GPU, longer contexts within a fixed memory budget, or lower KV-cache cost per token β€” chatbots, RAG systems, and agentic workloads with large contexts.

Release Date

Hugging Face 07/2026 via this repository. Companion model: Mistral-7B-Instruct-v0.3 kit.

References

Model Architecture

Artifact type: Calibrated KV-cache eigenbasis sidecars (K + V) consumed by a compressed-page attention runtime for vLLM.

Compression architecture: fraQtl inserts a compression membrane between the model and vLLM's paged KV cache. On write, K and V are stored in a compressed page format: a calibrated protected subspace kept at high precision (rank 16 of 128 for K, rank 32 of 128 for V, per layer per KV head β€” the sidecars in this repository) plus low-bit tails on the remaining dimensions β€” V compressed more aggressively than K, whose error softmax amplifies. Logical rank stays 128; nothing is truncated. On read, the attention kernel consumes compressed pages directly at tensor-core speed β€” no decompress-then-attend step, which is why capacity gains do not cost decode bandwidth. The exact tail format is part of the unpublished method.

How the protected subspaces are calibrated is not published. The sidecars are the calibrated artifacts β€” sufficient to run and verify every number on this page.

Input / Output

  • Input: text (any workload the base model supports)
  • Output: text
  • Context length: receipted at 8K, 32K, and 128K. Qwen3's native 262K window is used as-is (no RoPE overrides). No 256K claims (see Model Limitations).

Software Integration

Supported runtime engine: vLLM 0.20.2 (torch 2.11.0 β€” the receipt-validated stack)

Supported hardware microarchitecture: NVIDIA Ampere (SM80 / A100). Hopper/Blackwell not yet supported.

Operating system: Linux

Integration uses vLLM's standard out-of-tree plugin mechanism (vllm.general_plugins entry point): pip install the wheel and the backend registers in every engine worker process.

Calibration Dataset

  • Dataset: wikitext-2-raw-v1 (test split)
  • Data collection method: Automated
  • Properties: K sidecar calibrated on 16 sequences Γ— 1024 tokens; per-layer, per-KV-head. No task data, no benchmark data, no retrieval-test content.

Usage

Reproduce every batch-1 number on this page with one command (β‰ˆ$15 of rented A100 time):

pip install modal && modal setup
modal secret create huggingface HF_TOKEN=hf_...
modal run fraqtl_repro_receipts.py --model qwen3-4b-instruct-2507 --contexts 8k,32k,128k

Prints the three-arm table with per-arm needle-in-a-haystack results. Docker fallback: --local. Script: fraqtl_repro_receipts.py.

The 9-user receipt (2026-08-14)

Nine users chart

Arm Users @β‰ˆ128K Aggregate decode Per user
fp16 2 (holds 2 without preemption) 66.6 tok/s ~33
FP8 5 (holds 5 without preemption) 117.6 tok/s (its peak, at 4 users) ~29 at peak
fraQtl D2 9 134.1 tok/s 14.9

Same GPU (A100-80GB), same model, vLLM custom backend, CUDA graphs on, prefix caching off. KV pool: 1,255,376 tokens. Recipe as executed: K 16 protected dims, V 32 protected dims, low-bit tails, logical rank 128 (KV pool fingerprint: 1,255,376 tokens).

Measurement detail, stated precisely: the retrieval check ran 131,040-token prompts per user (+32 generated); the timing pass ran 125,795-token prompts per user with 1,024 decode tokens each (9,216 total) β€” both β‰ˆ128K workloads. Baseline arms were measured at 130,048-token prompts (a 0.76% difference, disclosed). Needle grid: nine distinct passkeys, placements spanning 0–100% of context (depths 0.0 / 0.1 / 0.25 / 0.25 / 0.4 / 0.5 / 0.75 / 0.9 / 1.0) β€” 9/9 found. Zero kernel fallbacks, zero CUDA errors, byte-exact roundtrip witness on the storage format.

Run IDs: needles ap-pBrqbRbG2UJqb5HvQDbMg4 Β· timing ap-U44I0iXVh2vrYX9ALCqIeB Β· byte-exactness ap-fMsMp4N2eYHHmfwRXqOCGG. Receipt JSON: receipts/batch9_128k_2026-08-14/.

Stated plainly: per-user decode at this capacity is 14.9 tok/s (fp16 gives 33 β€” for two users; at 7 users fp16's per-user collapses to ~9.3 under preemption). All three arms passed 100% of retrieval checks in the baseline depth grids (189/189) β€” the baselines compared against are not broken ones. Separate D2-family quality evidence (related recipe lineage, not this exact recipe): "On the 2026-06-11 hardened paired packet, D2 located 63/63 NIAH needles (fp16-parity); observed failures are rare exact-ID transcription noise (5%) and synthetic 3-hop state-tracking flips (~17%, symmetric), both root-caused; real multi-hop QA (LongBench) shows no F1 gap." The direct quality evidence for this exact recipe is the 9/9 grid above.

Evaluation

Three arms, same command, same GPU (A100-80GB), retrieval gate per arm: fraQtl D2 vs fp16 (stock vLLM) vs fp8-KV (kv_cache_dtype=fp8, the strongest available baseline).

Batch-1 decode and KV capacity

ctx arm decode tok/s NIAH KV pool (tokens)
8K fraQtl D2 115.9 PASS 1,012,896
8K fp16 125.74 PASS 424,496
8K fp8 128.08 PASS 834,912
32K fraQtl D2 101.52 PASS 984,432
32K fp16 104.18 PASS 412,992
32K fp8 105.81 PASS 811,920
128K fraQtl D2 69.5 PASS 845,136
128K fp16 52.59 PASS 361,776
128K fp8 72.7 PASS 709,488

128K cells: chunked prefill disabled, max_num_batched_tokens=131072, 1024 decode tokens (32-token subtractive timing is a measurement artifact at 128K). Receipt run IDs: 8K ap-E4uir4SUtT5YLmKSJq2diz, 32K ap-cQiTpbsP6wBYPUkZeXLjnm, 128K ap-QuvidvohDU0DVYEKUomCKt / ap-uCQmv4B5FI985G1lfyxitA / ap-Dgvl6PXVpRBc9JLx1b1U7M.

128K concurrency (aggregate decode tok/s, 1024 tokens per user, per-user needle verified)

users fraQtl D2 fp8 fp16
1 69.5 72.7 54.0
2 100.8 97.8 66.6
3 116.5 110.2 62.1
4 130.1 117.6 67.3
5 118.7 105.0 61.8
6 140.9 99.6 67.2
7 121.4 105.0 65.0

Preemption-free capacity: D2 6 users, fp8 5, fp16 2. At each arm's own clean maximum: D2 140.9 tok/s @ 6 Β· fp8 105.0 @ 5 Β· fp16 66.6 @ 2 β€” 2.1Γ— fp16 and 1.3Γ— fp8 aggregate throughput. D2 unchunked prefill holds β‰ˆ3,950 tok/s at every rung (β‰₯ fp16's 3,820, > fp8's 2,720). Run IDs: ap-uCQmv4B5FI985G1lfyxitA, ap-Dgvl6PXVpRBc9JLx1b1U7M, ap-BLmwYqynVkeUk5S6i8s9zG.

32K high-batch (8–24 users, aggregate tok/s, all needles green)

users fraQtl D2 fp8 fp16
8 335.9 349.6 265.6
12 363.0 259.5 254.9
16 408.5 333.9 233.0
20 409.1 379.8 249.1
24 429.7 432.3 246.0

At 24 users @32K: 1.75Γ— fp16, parity with fp8 β€” the vs-fp8 advantage lives at 128K, where fp8's pool runs out first. Run ID: ap-sISJ6RY7TOm3WQASL8YCJX.

Retrieval verification

Needle-in-a-haystack passkey grids, 7 depths Γ— 3 keys per context, exact-match gated: 21/21 per arm at 8K, 32K, and 128K β€” 189/189 total. Run IDs: ap-l7cHhG8pwbKqndTqKRxBfU, ap-YzAkwaYtW7FrUuFf5vAfhC, ap-zjyaiI4b7Rl4VujTuozW7O, ap-utQpsypNBJQrwrXHzRqhTN.

Speculative decoding compatibility (2026-08-24)

Speculative decoding compatibility: receipted. ngram drafting over the compressed cache at 9 users Γ— 131K went from 0.76Γ— (root-caused: kernel served one query token per cache pass) to 1.04Γ— after the fix (bit-exact paired-token kernel, run ap-7Mfjwuudo5qNpDG86QHfaO). Claim scope: no longer a loss at c=9 β€” the rung where drafting is expected to stop paying; middle-rung receipts (c=4) in progress. Mean acceptance length ~4.0 with plain ngram lookup at 131K.

Independent reproduction

All nine batch-1 cells were re-run (2026-07-06) through the exact public path in this repository β€” prebuilt wheel + these sidecars, no internal source: all cells NIAH PASS; D2 KV pools byte-identical at all three contexts (1,012,896 / 984,432 / 845,136); decode within run-to-run band. Receipt JSONs: receipts/repro_2026_07_06/.

Evaluation methodology notes

  1. Retrieval = NIAH passkey grids, exact-match gated. We never say "lossless".
  2. fp8-KV is the strongest available baseline and appears in every table.
  3. Batch and context are stated on every number; batch-1 short-context decode is weight-bound β€” parity is the ceiling for every KV method there.
  4. 2026-07-03: a batch-routing bug (mixed prefill+decode steps) dropped needles at batch β‰₯ 2; root-caused, fixed, full ladder re-receipted green. Pre-fix multi-user results were never published.

Model Limitations

  • Hardware: SM80 (A100) only; the runtime binary will not run on H100 or consumer GPUs.
  • 128K prefill: D2 receipts run with chunked prefill disabled (known 6.7Γ— D2 prefill regression under chunking; fix in progress). Stock arms unaffected.
  • 256K: no claims β€” D2 has a known precision gap at 256K depth 0.25 (fp16/fp8 controls pass); capacity receipts exist, quality gate open.
  • Compression is lossy by construction; the evidence standard is retrieval verification, not bit-identity. Behavior on tasks other than retrieval is not separately benchmarked here.
  • The base model's own limitations and biases are inherited unchanged (weights are not modified).

Files

File What
sidecar_real_u_qwen3-4b-instruct-2507.bin V eigenbasis sidecar (per-layer, per-kv-head U matrices)
sidecar_real_u_qwen3-4b-instruct-2507.meta.json V sidecar provenance (dims, calibration corpus/size, sha256)
qwen3-4b-instruct-2507-k16-int3.fraqtl-k-eigenbasis.bin K eigenbasis sidecar (RAQK format, rank 16)
qwen3-4b-instruct-2507-k16-int3.fraqtl-k-eigenbasis.receipt.json K sidecar build receipt
receipts/repro_2026_07_06/ Independent reproduction receipts
fraqtl_repro_receipts.py One-command reproduction script

More Models

Same recipe, receipt format, and measurement standards per model: Mistral-7B-Instruct-v0.3 (live). Next: Llama-3.1-8B, DeepSeek-R1-Distill-Llama-8B, Mistral-Nemo. Request a model by opening a discussion β€” calibration is one factory run.

llama.cpp membrane runtime β€” same sidecars, second engine (2026-08-25)

The sidecars in this repository also power the fraQtl membrane inside a llama.cpp fork β€” compressed KV read natively at decode.

llama.cpp receipts (membrane runtime, 1Γ—A100-80GB, Q4_K_M weights)

Decode tok/s, single sequence (identical prompts; baseline stated per row):

Context membrane q8_0 baseline ratio baseline build
8K 120.8 127.5 0.95Γ— same binary
32K 101.4 90.0 1.13Γ— upstream master (commit pinned in receipt)
128K 68.8 38.4 1.79Γ— upstream master (commit pinned in receipt)

Upstream fp16 KV at 128K: 74.1 tok/s β€” membrane reaches 92% of fp16 decode with ~2.45Γ— less KV memory.

Concurrent capacity ("users" = N parallel sequences via -np N, each with an independent KV stream and a full 32,832-token context allocation; exact command line embedded in every receipt JSON):

32K tokens/user membrane-exclusive q8_0
Max users before OOM 36 28
Aggregate tok/s at max 335.5 164.2

Retrieval verification: needle-retrieval at 7 depths (0.05–0.95) Γ— distinct high-entropy keys at 8K and 32K, plus 128K at depths 0.1/0.5/0.9 β€” 24/24 retrieved; q8_0 baseline 7/7 (parity at every depth). Multi-user isolation verified (distinct keys per stream, users finishing at different times).

How to run (llama.cpp)

Runtime ships as a prebuilt binary β€” same posture as the vLLM wheel: you get the built runtime and the calibrated sidecars; kernel source is not distributed. Download: fraQtl/fraqtl-membrane-llamacpp-runtime (Linux x86_64, CUDA, SM80+SM90, manifest-sha'd). Flags: FRAQTL_MEMBRANE=1 FRAQTL_MEMBRANE_EXCLUSIVE=1 + --fraqtl-kv --fraqtl-eigenbasis <v-sidecar> --fraqtl-kv-protect 32 --fraqtl-k-eigenbasis <k-sidecar> --fraqtl-sink-tokens 0 --fraqtl-residual-window 0]

llama.cpp fine print

  • Weights are quantized Q4_K_M (bartowski GGUF); fraQtl compresses the KV cache, not the weights.
  • Short-context (≀8K) decode is β‰ˆ95% of q8_0 β€” the advantage grows with context and with concurrent load.
  • Prefill: 128K ingest 54.5s vs q8's 31.6s (1.7Γ—) β€” improving; numbers in receipts.
  • All receipts (JSON, with exact cmd + env + upstream commit) reproducible via the fraQtl gate harness.

More from fraQtl

The on-device lane β€” calibration-aware Hi-Fi GGUFs with KLD receipts: Gemma-4-E2B (runs an offline phone agent) Β· Llama-3.1-8B + Llama-3.2-1B draft pair Β· Qwen3.6-35B MoE. Same standard per artifact: pinned provenance, measured numbers, losses disclosed. Org page: huggingface.co/fraQtl.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for fraQtl/qwen3-4b-instruct-2507-kv-sidecars

Adapter
(5723)
this model

Collection including fraQtl/qwen3-4b-instruct-2507-kv-sidecars