Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,96 +1,92 @@
|
|
| 1 |
-
---
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
##
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
#
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
#
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
``
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
```bibtex
|
| 83 |
-
@article{wang2026flashmemory,
|
| 84 |
-
title = {FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention},
|
| 85 |
-
author = {Yan Wang and Qifan Zhang and Jiachen Yu and Tian Liang and Dongyang Ma and
|
| 86 |
-
Xiang Hu and Zibo Lin and Chunyang Li and Zhichao Wang and Jia Li and
|
| 87 |
-
Yujiu Yang and Haitao Mi and Dong Yu},
|
| 88 |
-
year = {2026},
|
| 89 |
-
journal = {arXiv preprint arXiv:2606.09079},
|
| 90 |
-
url = {https://
|
| 91 |
-
}
|
| 92 |
-
```
|
| 93 |
-
|
| 94 |
-
## License
|
| 95 |
-
|
| 96 |
-
MIT
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
library_name: sglang
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- deepseek-v4
|
| 7 |
+
- long-context
|
| 8 |
+
- kv-cache
|
| 9 |
+
- sparse-attention
|
| 10 |
+
- retriever
|
| 11 |
+
- sglang
|
| 12 |
+
base_model: deepseek-ai/DeepSeek-V4
|
| 13 |
+
arxiv: 2606.09079
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# FlashMemory-DeepSeek-V4
|
| 17 |
+
|
| 18 |
+
**FlashMemory** is a trained **Memory-Indexer** retriever for **DeepSeek-V4 Compressed-Sparse-Attention (CSA)** KV-cache. Instead of scoring the full history every decode step, it predicts which ~10β15% of chunks the next 64 tokens will attend to β only those stay on GPU, the rest are offloaded to CPU.
|
| 19 |
+
|
| 20 |
+
- π **[Paper (arXiv:2606.09079)](https://arxiv.org/abs/2606.09079)**
|
| 21 |
+
- π€ **[Model weights (this repo)](https://huggingface.co/libertywing/FlashMemory-Deepseek-V4)**
|
| 22 |
+
- π» **[Code on GitHub](https://github.com/libertywing/FlashMemory-Deepseek-V4)**
|
| 23 |
+
|
| 24 |
+
## Highlights
|
| 25 |
+
|
| 26 |
+
FM-DS-V4 **matches or beats the DS-V4-Flash full-attention baseline** on long-context accuracy while dramatically cutting serving cost:
|
| 27 |
+
|
| 28 |
+
- At **1M context**, GPU KV cache shrinks by **~90%** (3.73 β 0.37 GB).
|
| 29 |
+
- Per-decode-token compute drops from **118.9 β 35.4 GFLOP (0.30Γ)**.
|
| 30 |
+
- **2.7Γ aggregate throughput**, and max concurrency rises from **11 β 40 (3.6Γ)**.
|
| 31 |
+
|
| 32 |
+
| Context | KMAX | Ours conc / throughput | Baseline conc / throughput | Gain conc | Gain throughput |
|
| 33 |
+
|---------|------|------------------------|----------------------------|-----------|-----------------|
|
| 34 |
+
| 256K | 96 | 76 / ~2759 tok/s | 47 / ~1584 tok/s | 1.6Γ | 1.7Γ |
|
| 35 |
+
| 512K | 192 | 60 / ~2008 tok/s | 25 / ~1028 tok/s | 2.4Γ | 2.0Γ |
|
| 36 |
+
| 1M | 384 | 30 / ~1266 tok/s | 11 / ~455 tok/s | 2.7οΏ½οΏ½ | 2.8Γ |
|
| 37 |
+
|
| 38 |
+
On LongBench-v2 / LongMemEval / RULER it matches or exceeds the full-attention baseline while keeping only ~10β15% of CSA KV on GPU. Ablations (**Recency-10%**, **Random-10%**) with the same KV budget confirm the gains come from **learned relevance**, not a positional/budget artifact.
|
| 39 |
+
|
| 40 |
+
## How it works (in brief)
|
| 41 |
+
|
| 42 |
+
Two indexers cooperate:
|
| 43 |
+
|
| 44 |
+
- **Level 1 β Memory Indexer (every 64 steps):** a trained retriever scores the history's compressed keys and selects query-critical chunks β a `resident_set` recalled from CPU to GPU.
|
| 45 |
+
- **Level 2 β Lightning Indexer (every step):** the native top-512 runs confined to the `resident_set`, at full speed inside cuda-graph.
|
| 46 |
+
|
| 47 |
+
The GPU/CPU split is the key: only recalled chunks + compressed indexer keys live on GPU; the bulk KV cache sits on CPU and is pulled on demand.
|
| 48 |
+
|
| 49 |
+
> For the full architecture, retriever math, hyperparameters and training details, see the **[paper](https://arxiv.org/abs/2606.09079)** and the **[GitHub code](https://github.com/libertywing/FlashMemory-Deepseek-V4)**.
|
| 50 |
+
|
| 51 |
+
## Inference (PD-disaggregated)
|
| 52 |
+
|
| 53 |
+
Three servers β **P** (prefill), **D** (decode, all offload/recall), **router**. Clients hit the router at `http://<router>:31503/v1/chat/completions`.
|
| 54 |
+
|
| 55 |
+
```bash
|
| 56 |
+
pip install -e sglang/python # also: pip install sgl_kernel==0.3.21
|
| 57 |
+
# download checkpoints/ from this HF repo
|
| 58 |
+
|
| 59 |
+
# D (decode) β GPU 0β7, 512K example:
|
| 60 |
+
TGT_CONC=60 TGT_CTX=524288 CTX_LEN=1100000 bash launch_decode.sh
|
| 61 |
+
# P (prefill) β GPU 0β7:
|
| 62 |
+
CTX_LEN=1100000 SWA_RATIO=0.1 HOST=<P_IP> bash launch_prefill.sh
|
| 63 |
+
# router:
|
| 64 |
+
PREFILL_IP=<P_IP> DECODE_IP=<D_IP> bash launch_router.sh
|
| 65 |
+
```
|
| 66 |
+
|
| 67 |
+
All optimizations are env-gated (gate-off = exact DS-V4 baseline). Key switches:
|
| 68 |
+
`SGLANG_DECODE_SWAP_P`, `SGLANG_PATHP_INDEX_K_OFFLOAD`, `SGLANG_PATHP_SCORE_RESIDENT`,
|
| 69 |
+
`SGLANG_PATHP_PAGE_RECALL`, `SGLANG_PATHP_CUDAGRAPH`, `SGLANG_PATHP_ASYNC_RECALL`,
|
| 70 |
+
`SGLANG_PATHP_FUSED_REMAP`.
|
| 71 |
+
|
| 72 |
+
## Checkpoint
|
| 73 |
+
|
| 74 |
+
Download `checkpoints/` into the repo root (default `top3_R930_joint.pt`, CSA layers 10/12/20).
|
| 75 |
+
|
| 76 |
+
## License
|
| 77 |
+
|
| 78 |
+
MIT
|
| 79 |
+
|
| 80 |
+
## Citation
|
| 81 |
+
|
| 82 |
+
```bibtex
|
| 83 |
+
@article{wang2026flashmemory,
|
| 84 |
+
title = {FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention},
|
| 85 |
+
author = {Yan Wang and Qifan Zhang and Jiachen Yu and Tian Liang and Dongyang Ma and
|
| 86 |
+
Xiang Hu and Zibo Lin and Chunyang Li and Zhichao Wang and Jia Li and
|
| 87 |
+
Yujiu Yang and Haitao Mi and Dong Yu},
|
| 88 |
+
year = {2026},
|
| 89 |
+
journal = {arXiv preprint arXiv:2606.09079},
|
| 90 |
+
url = {https://arxiv.org/abs/2606.09079},
|
| 91 |
+
}
|
| 92 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|