- Qwen3-4B-Instruct-2507 β fraQtl KV-Cache Compression Sidecars
- Model Overview
- Model Architecture
- Input / Output
- Software Integration
- Calibration Dataset
- Usage
- The 9-user receipt (2026-08-14)
- Evaluation
- Batch-1 decode and KV capacity
- 128K concurrency (aggregate decode tok/s, 1024 tokens per user, per-user needle verified)
- 32K high-batch (8β24 users, aggregate tok/s, all needles green)
- Retrieval verification
- Speculative decoding compatibility (2026-08-24)
- Independent reproduction
- Evaluation methodology notes
- Model Limitations
- Files
- More Models
- llama.cpp membrane runtime β same sidecars, second engine (2026-08-25)
- llama.cpp receipts (membrane runtime, 1ΓA100-80GB, Q4_K_M weights)
- How to run (llama.cpp)
- llama.cpp fine print
- More from fraQtl
Qwen3-4B-Instruct-2507 β fraQtl KV-Cache Compression Sidecars
Model Overview
Description
This is the artifact that serves nine concurrent β128K-context users on one A100 β 4.5Γ the users fp16 holds without preemption and 1.8Γ FP8's (9 vs 5), at 2Γ fp16's total decode throughput and 1.14Γ FP8's peak (134.1 vs 66.6 vs 117.6 tok/s aggregate) β with nine distinct passkeys planted at placements spanning 0β100% of the context and 9/9 recovered exactly. The receipt for that sentence, and every other number on this page, is in this repository.
The sidecars are calibrated per-layer, per-KV-head eigenbasis artifacts consumed by the fraqtl runtime (a prebuilt binary vLLM attention backend). Base model weights are not modified or redistributed.
License/Terms of Use
Sidecars and receipts: Apache 2.0. Base model weights download from Qwen/Qwen3-4B-Instruct-2507 under its own license. The fraqtl runtime wheel is proprietary, free to install and run for verification and evaluation (license).
Deployment Geography
Global.
Use Case
Teams serving LLMs with vLLM who need more concurrent long-context users per GPU, longer contexts within a fixed memory budget, or lower KV-cache cost per token β chatbots, RAG systems, and agentic workloads with large contexts.
Release Date
Hugging Face 07/2026 via this repository. Companion model: Mistral-7B-Instruct-v0.3 kit.
References
- Runtime wheel + build provenance: fraQtl/fraqtl-sm80-runtime
- Independent reproduction receipts:
receipts/repro_2026_07_06/
Model Architecture
Artifact type: Calibrated KV-cache eigenbasis sidecars (K + V) consumed by a compressed-page attention runtime for vLLM.
Compression architecture: fraQtl inserts a compression membrane between the model and vLLM's paged KV cache. On write, K and V are stored in a compressed page format: a calibrated protected subspace kept at high precision (rank 16 of 128 for K, rank 32 of 128 for V, per layer per KV head β the sidecars in this repository) plus low-bit tails on the remaining dimensions β V compressed more aggressively than K, whose error softmax amplifies. Logical rank stays 128; nothing is truncated. On read, the attention kernel consumes compressed pages directly at tensor-core speed β no decompress-then-attend step, which is why capacity gains do not cost decode bandwidth. The exact tail format is part of the unpublished method.
How the protected subspaces are calibrated is not published. The sidecars are the calibrated artifacts β sufficient to run and verify every number on this page.
Input / Output
- Input: text (any workload the base model supports)
- Output: text
- Context length: receipted at 8K, 32K, and 128K. Qwen3's native 262K window is used as-is (no RoPE overrides). No 256K claims (see Model Limitations).
Software Integration
Supported runtime engine: vLLM 0.20.2 (torch 2.11.0 β the receipt-validated stack)
Supported hardware microarchitecture: NVIDIA Ampere (SM80 / A100). Hopper/Blackwell not yet supported.
Operating system: Linux
Integration uses vLLM's standard out-of-tree plugin mechanism (vllm.general_plugins entry point): pip install the wheel and the backend registers in every engine worker process.
Calibration Dataset
- Dataset: wikitext-2-raw-v1 (test split)
- Data collection method: Automated
- Properties: K sidecar calibrated on 16 sequences Γ 1024 tokens; per-layer, per-KV-head. No task data, no benchmark data, no retrieval-test content.
Usage
Reproduce every batch-1 number on this page with one command (β$15 of rented A100 time):
pip install modal && modal setup
modal secret create huggingface HF_TOKEN=hf_...
modal run fraqtl_repro_receipts.py --model qwen3-4b-instruct-2507 --contexts 8k,32k,128k
Prints the three-arm table with per-arm needle-in-a-haystack results. Docker fallback: --local. Script: fraqtl_repro_receipts.py.
The 9-user receipt (2026-08-14)
| Arm | Users @β128K | Aggregate decode | Per user |
|---|---|---|---|
| fp16 | 2 (holds 2 without preemption) | 66.6 tok/s | ~33 |
| FP8 | 5 (holds 5 without preemption) | 117.6 tok/s (its peak, at 4 users) | ~29 at peak |
| fraQtl D2 | 9 | 134.1 tok/s | 14.9 |
Same GPU (A100-80GB), same model, vLLM custom backend, CUDA graphs on, prefix caching off. KV pool: 1,255,376 tokens. Recipe as executed: K 16 protected dims, V 32 protected dims, low-bit tails, logical rank 128 (KV pool fingerprint: 1,255,376 tokens).
Measurement detail, stated precisely: the retrieval check ran 131,040-token prompts per user (+32 generated); the timing pass ran 125,795-token prompts per user with 1,024 decode tokens each (9,216 total) β both β128K workloads. Baseline arms were measured at 130,048-token prompts (a 0.76% difference, disclosed). Needle grid: nine distinct passkeys, placements spanning 0β100% of context (depths 0.0 / 0.1 / 0.25 / 0.25 / 0.4 / 0.5 / 0.75 / 0.9 / 1.0) β 9/9 found. Zero kernel fallbacks, zero CUDA errors, byte-exact roundtrip witness on the storage format.
Run IDs: needles ap-pBrqbRbG2UJqb5HvQDbMg4 Β· timing ap-U44I0iXVh2vrYX9ALCqIeB Β· byte-exactness ap-fMsMp4N2eYHHmfwRXqOCGG. Receipt JSON: receipts/batch9_128k_2026-08-14/.
Stated plainly: per-user decode at this capacity is 14.9 tok/s (fp16 gives 33 β for two users; at 7 users fp16's per-user collapses to ~9.3 under preemption). All three arms passed 100% of retrieval checks in the baseline depth grids (189/189) β the baselines compared against are not broken ones. Separate D2-family quality evidence (related recipe lineage, not this exact recipe): "On the 2026-06-11 hardened paired packet, D2 located 63/63 NIAH needles (fp16-parity); observed failures are rare exact-ID transcription noise (5%) and synthetic 3-hop state-tracking flips (~17%, symmetric), both root-caused; real multi-hop QA (LongBench) shows no F1 gap." The direct quality evidence for this exact recipe is the 9/9 grid above.
Evaluation
Three arms, same command, same GPU (A100-80GB), retrieval gate per arm: fraQtl D2 vs fp16 (stock vLLM) vs fp8-KV (kv_cache_dtype=fp8, the strongest available baseline).
Batch-1 decode and KV capacity
| ctx | arm | decode tok/s | NIAH | KV pool (tokens) |
|---|---|---|---|---|
| 8K | fraQtl D2 | 115.9 | PASS | 1,012,896 |
| 8K | fp16 | 125.74 | PASS | 424,496 |
| 8K | fp8 | 128.08 | PASS | 834,912 |
| 32K | fraQtl D2 | 101.52 | PASS | 984,432 |
| 32K | fp16 | 104.18 | PASS | 412,992 |
| 32K | fp8 | 105.81 | PASS | 811,920 |
| 128K | fraQtl D2 | 69.5 | PASS | 845,136 |
| 128K | fp16 | 52.59 | PASS | 361,776 |
| 128K | fp8 | 72.7 | PASS | 709,488 |
128K cells: chunked prefill disabled, max_num_batched_tokens=131072, 1024 decode tokens (32-token subtractive timing is a measurement artifact at 128K). Receipt run IDs: 8K ap-E4uir4SUtT5YLmKSJq2diz, 32K ap-cQiTpbsP6wBYPUkZeXLjnm, 128K ap-QuvidvohDU0DVYEKUomCKt / ap-uCQmv4B5FI985G1lfyxitA / ap-Dgvl6PXVpRBc9JLx1b1U7M.
128K concurrency (aggregate decode tok/s, 1024 tokens per user, per-user needle verified)
| users | fraQtl D2 | fp8 | fp16 |
|---|---|---|---|
| 1 | 69.5 | 72.7 | 54.0 |
| 2 | 100.8 | 97.8 | 66.6 |
| 3 | 116.5 | 110.2 | 62.1 |
| 4 | 130.1 | 117.6 | 67.3 |
| 5 | 118.7 | 105.0 | 61.8 |
| 6 | 140.9 | 99.6 | 67.2 |
| 7 | 121.4 | 105.0 | 65.0 |
Preemption-free capacity: D2 6 users, fp8 5, fp16 2. At each arm's own clean maximum: D2 140.9 tok/s @ 6 Β· fp8 105.0 @ 5 Β· fp16 66.6 @ 2 β 2.1Γ fp16 and 1.3Γ fp8 aggregate throughput. D2 unchunked prefill holds β3,950 tok/s at every rung (β₯ fp16's 3,820, > fp8's 2,720). Run IDs: ap-uCQmv4B5FI985G1lfyxitA, ap-Dgvl6PXVpRBc9JLx1b1U7M, ap-BLmwYqynVkeUk5S6i8s9zG.
32K high-batch (8β24 users, aggregate tok/s, all needles green)
| users | fraQtl D2 | fp8 | fp16 |
|---|---|---|---|
| 8 | 335.9 | 349.6 | 265.6 |
| 12 | 363.0 | 259.5 | 254.9 |
| 16 | 408.5 | 333.9 | 233.0 |
| 20 | 409.1 | 379.8 | 249.1 |
| 24 | 429.7 | 432.3 | 246.0 |
At 24 users @32K: 1.75Γ fp16, parity with fp8 β the vs-fp8 advantage lives at 128K, where fp8's pool runs out first. Run ID: ap-sISJ6RY7TOm3WQASL8YCJX.
Retrieval verification
Needle-in-a-haystack passkey grids, 7 depths Γ 3 keys per context, exact-match gated: 21/21 per arm at 8K, 32K, and 128K β 189/189 total. Run IDs: ap-l7cHhG8pwbKqndTqKRxBfU, ap-YzAkwaYtW7FrUuFf5vAfhC, ap-zjyaiI4b7Rl4VujTuozW7O, ap-utQpsypNBJQrwrXHzRqhTN.
Speculative decoding compatibility (2026-08-24)
Speculative decoding compatibility: receipted. ngram drafting over the compressed cache at 9 users Γ 131K went from 0.76Γ (root-caused: kernel served one query token per cache pass) to 1.04Γ after the fix (bit-exact paired-token kernel, run ap-7Mfjwuudo5qNpDG86QHfaO). Claim scope: no longer a loss at c=9 β the rung where drafting is expected to stop paying; middle-rung receipts (c=4) in progress. Mean acceptance length ~4.0 with plain ngram lookup at 131K.
Independent reproduction
All nine batch-1 cells were re-run (2026-07-06) through the exact public path in this repository β prebuilt wheel + these sidecars, no internal source: all cells NIAH PASS; D2 KV pools byte-identical at all three contexts (1,012,896 / 984,432 / 845,136); decode within run-to-run band. Receipt JSONs: receipts/repro_2026_07_06/.
Evaluation methodology notes
- Retrieval = NIAH passkey grids, exact-match gated. We never say "lossless".
- fp8-KV is the strongest available baseline and appears in every table.
- Batch and context are stated on every number; batch-1 short-context decode is weight-bound β parity is the ceiling for every KV method there.
- 2026-07-03: a batch-routing bug (mixed prefill+decode steps) dropped needles at batch β₯ 2; root-caused, fixed, full ladder re-receipted green. Pre-fix multi-user results were never published.
Model Limitations
- Hardware: SM80 (A100) only; the runtime binary will not run on H100 or consumer GPUs.
- 128K prefill: D2 receipts run with chunked prefill disabled (known 6.7Γ D2 prefill regression under chunking; fix in progress). Stock arms unaffected.
- 256K: no claims β D2 has a known precision gap at 256K depth 0.25 (fp16/fp8 controls pass); capacity receipts exist, quality gate open.
- Compression is lossy by construction; the evidence standard is retrieval verification, not bit-identity. Behavior on tasks other than retrieval is not separately benchmarked here.
- The base model's own limitations and biases are inherited unchanged (weights are not modified).
Files
| File | What |
|---|---|
sidecar_real_u_qwen3-4b-instruct-2507.bin |
V eigenbasis sidecar (per-layer, per-kv-head U matrices) |
sidecar_real_u_qwen3-4b-instruct-2507.meta.json |
V sidecar provenance (dims, calibration corpus/size, sha256) |
qwen3-4b-instruct-2507-k16-int3.fraqtl-k-eigenbasis.bin |
K eigenbasis sidecar (RAQK format, rank 16) |
qwen3-4b-instruct-2507-k16-int3.fraqtl-k-eigenbasis.receipt.json |
K sidecar build receipt |
receipts/repro_2026_07_06/ |
Independent reproduction receipts |
fraqtl_repro_receipts.py |
One-command reproduction script |
More Models
Same recipe, receipt format, and measurement standards per model: Mistral-7B-Instruct-v0.3 (live). Next: Llama-3.1-8B, DeepSeek-R1-Distill-Llama-8B, Mistral-Nemo. Request a model by opening a discussion β calibration is one factory run.
llama.cpp membrane runtime β same sidecars, second engine (2026-08-25)
The sidecars in this repository also power the fraQtl membrane inside a llama.cpp fork β compressed KV read natively at decode.
llama.cpp receipts (membrane runtime, 1ΓA100-80GB, Q4_K_M weights)
Decode tok/s, single sequence (identical prompts; baseline stated per row):
| Context | membrane | q8_0 baseline | ratio | baseline build |
|---|---|---|---|---|
| 8K | 120.8 | 127.5 | 0.95Γ | same binary |
| 32K | 101.4 | 90.0 | 1.13Γ | upstream master (commit pinned in receipt) |
| 128K | 68.8 | 38.4 | 1.79Γ | upstream master (commit pinned in receipt) |
Upstream fp16 KV at 128K: 74.1 tok/s β membrane reaches 92% of fp16 decode with ~2.45Γ less KV memory.
Concurrent capacity ("users" = N parallel sequences via -np N, each
with an independent KV stream and a full 32,832-token context allocation;
exact command line embedded in every receipt JSON):
| 32K tokens/user | membrane-exclusive | q8_0 |
|---|---|---|
| Max users before OOM | 36 | 28 |
| Aggregate tok/s at max | 335.5 | 164.2 |
Retrieval verification: needle-retrieval at 7 depths (0.05β0.95) Γ distinct high-entropy keys at 8K and 32K, plus 128K at depths 0.1/0.5/0.9 β 24/24 retrieved; q8_0 baseline 7/7 (parity at every depth). Multi-user isolation verified (distinct keys per stream, users finishing at different times).
How to run (llama.cpp)
Runtime ships as a prebuilt binary β same posture as the vLLM wheel: you
get the built runtime and the calibrated sidecars; kernel source is not
distributed. Download:
fraQtl/fraqtl-membrane-llamacpp-runtime
(Linux x86_64, CUDA, SM80+SM90, manifest-sha'd). Flags: FRAQTL_MEMBRANE=1 FRAQTL_MEMBRANE_EXCLUSIVE=1 + --fraqtl-kv --fraqtl-eigenbasis <v-sidecar> --fraqtl-kv-protect 32 --fraqtl-k-eigenbasis <k-sidecar> --fraqtl-sink-tokens 0 --fraqtl-residual-window 0]
llama.cpp fine print
- Weights are quantized Q4_K_M (bartowski GGUF); fraQtl compresses the KV cache, not the weights.
- Short-context (β€8K) decode is β95% of q8_0 β the advantage grows with context and with concurrent load.
- Prefill: 128K ingest 54.5s vs q8's 31.6s (1.7Γ) β improving; numbers in receipts.
- All receipts (JSON, with exact cmd + env + upstream commit) reproducible via the fraQtl gate harness.
More from fraQtl
The on-device lane β calibration-aware Hi-Fi GGUFs with KLD receipts: Gemma-4-E2B (runs an offline phone agent) Β· Llama-3.1-8B + Llama-3.2-1B draft pair Β· Qwen3.6-35B MoE. Same standard per artifact: pinned provenance, measured numbers, losses disclosed. Org page: huggingface.co/fraQtl.
- Downloads last month
- 14
Model tree for fraQtl/qwen3-4b-instruct-2507-kv-sidecars
Base model
Qwen/Qwen3-4B-Instruct-2507

