libertywing commited on
Commit
5b5caa5
Β·
verified Β·
1 Parent(s): 70431ba

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +92 -96
README.md CHANGED
@@ -1,96 +1,92 @@
1
- ---
2
- datasets:
3
- - ruler
4
- - longmemeval
5
- - longbench-v2
6
- - mrcr
7
- language: en
8
- license: mit
9
- pipeline_tag: feature-extraction
10
- tags:
11
- - deepseek-v4
12
- - retrieval
13
- - kv-cache
14
- - sparse-attention
15
- - long-context
16
- - flashmemory
17
- ---
18
-
19
- # FlashMemory DS-V4 Retriever
20
-
21
- A lightweight retriever that sparsifies **DeepSeek-V4 Compressed-Sparse-Attention (CSA) KV-cache**. Given a decode-token hidden state, it predicts which compressed-K chunks the next ~64 tokens will attend to β€” keeping only those on GPU, offloading the rest.
22
-
23
- This model is the Neural Memory Indexer presented in [FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention](https://huggingface.co/papers/2606.09079).
24
-
25
- Detailed code and inference scripts can be found on GitHub: [libertywing/FlashMemory-Deepseek-V4](https://github.com/libertywing/FlashMemory-Deepseek-V4).
26
-
27
- ## Performance
28
-
29
- In downstream evaluation, it matches or beats the full-attention baseline on reasoning-heavy long-context tasks (**RULER, LongMemEval, LongBench V2**) while reducing KV-cache usage by **~85–90%**. Precise needle-retrieval tasks require an additional threshold-fallback mechanism (not in this release).
30
-
31
- ## Quick start
32
-
33
- To use this model, you will need the `retriever.py` file from the [official repository](https://github.com/libertywing/FlashMemory-Deepseek-V4).
34
-
35
- ```bash
36
- pip install torch safetensors
37
- python demo.py --ckpt weights/flashmemory_ds_v4.safetensors
38
- ```
39
-
40
- ## Usage
41
-
42
- ```python
43
- from retriever import FlashMemoryRetriever
44
-
45
- model = FlashMemoryRetriever.from_checkpoint(
46
- "weights/flashmemory_ds_v4.safetensors", device="cuda"
47
- )
48
-
49
- # hidden: [B, 4096] decode-token hidden state
50
- # comp_k: [B, N, 132] uint8 compressed CSA keys
51
- # positions: [B] int64 token positions
52
-
53
- # Per-layer sigmoid scores: {"l10": [B,N], "l12": [B,N], "l20": [B,N]}
54
- per_layer = model(hidden, comp_k, positions)
55
-
56
- # Cross-layer ensemble (mode="max" or "mean")
57
- scores = model.ensemble(hidden, comp_k, positions, mode="max") # [B, N]
58
-
59
- # Boolean keep mask
60
- keep = model.select_topk(hidden, comp_k, positions, top_k=512) # top-K
61
- keep = model.select_topk(hidden, comp_k, positions, threshold=0.5) # threshold
62
- ```
63
-
64
- **`compressed_k` format:** each chunk = 128 bytes `float8_e4m3` values + 4 bytes `float32` scale. See `make_mock_compressed_k()` in `demo.py`.
65
-
66
- ## Architecture
67
-
68
- 3-layer joint model (`l10`, `l12`, `l20`), 128 heads, 2048 LoRA rank. Per-layer sigmoid scores are ensembled (`max` or `mean`) per chunk.
69
-
70
- ```
71
- hidden [B,4096] β†’ q-proj β†’ RoPE(YaRN) β†’ Hadamard β†’ q [B,128,128]
72
- β†’ weights_proj β†’ fused_w [B,128]
73
- compressed_k β†’ FP8 dequant β†’ k [B,N,128]
74
-
75
- score = sigmoid( Ξ£( relu(k @ qα΅€) Β· fused_w ) ) ∈ [0,1]
76
- ```
77
-
78
- ## Citation
79
-
80
- If you use FlashMemory in your research, please cite:
81
-
82
- ```bibtex
83
- @article{wang2026flashmemory,
84
- title = {FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention},
85
- author = {Yan Wang and Qifan Zhang and Jiachen Yu and Tian Liang and Dongyang Ma and
86
- Xiang Hu and Zibo Lin and Chunyang Li and Zhichao Wang and Jia Li and
87
- Yujiu Yang and Haitao Mi and Dong Yu},
88
- year = {2026},
89
- journal = {arXiv preprint arXiv:2606.09079},
90
- url = {https://huggingface.co/papers/2606.09079},
91
- }
92
- ```
93
-
94
- ## License
95
-
96
- MIT
 
1
+ ---
2
+ license: mit
3
+ library_name: sglang
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - deepseek-v4
7
+ - long-context
8
+ - kv-cache
9
+ - sparse-attention
10
+ - retriever
11
+ - sglang
12
+ base_model: deepseek-ai/DeepSeek-V4
13
+ arxiv: 2606.09079
14
+ ---
15
+
16
+ # FlashMemory-DeepSeek-V4
17
+
18
+ **FlashMemory** is a trained **Memory-Indexer** retriever for **DeepSeek-V4 Compressed-Sparse-Attention (CSA)** KV-cache. Instead of scoring the full history every decode step, it predicts which ~10–15% of chunks the next 64 tokens will attend to β€” only those stay on GPU, the rest are offloaded to CPU.
19
+
20
+ - πŸ“„ **[Paper (arXiv:2606.09079)](https://arxiv.org/abs/2606.09079)**
21
+ - πŸ€— **[Model weights (this repo)](https://huggingface.co/libertywing/FlashMemory-Deepseek-V4)**
22
+ - πŸ’» **[Code on GitHub](https://github.com/libertywing/FlashMemory-Deepseek-V4)**
23
+
24
+ ## Highlights
25
+
26
+ FM-DS-V4 **matches or beats the DS-V4-Flash full-attention baseline** on long-context accuracy while dramatically cutting serving cost:
27
+
28
+ - At **1M context**, GPU KV cache shrinks by **~90%** (3.73 β†’ 0.37 GB).
29
+ - Per-decode-token compute drops from **118.9 β†’ 35.4 GFLOP (0.30Γ—)**.
30
+ - **2.7Γ— aggregate throughput**, and max concurrency rises from **11 β†’ 40 (3.6Γ—)**.
31
+
32
+ | Context | KMAX | Ours conc / throughput | Baseline conc / throughput | Gain conc | Gain throughput |
33
+ |---------|------|------------------------|----------------------------|-----------|-----------------|
34
+ | 256K | 96 | 76 / ~2759 tok/s | 47 / ~1584 tok/s | 1.6Γ— | 1.7Γ— |
35
+ | 512K | 192 | 60 / ~2008 tok/s | 25 / ~1028 tok/s | 2.4Γ— | 2.0Γ— |
36
+ | 1M | 384 | 30 / ~1266 tok/s | 11 / ~455 tok/s | 2.7οΏ½οΏ½ | 2.8Γ— |
37
+
38
+ On LongBench-v2 / LongMemEval / RULER it matches or exceeds the full-attention baseline while keeping only ~10–15% of CSA KV on GPU. Ablations (**Recency-10%**, **Random-10%**) with the same KV budget confirm the gains come from **learned relevance**, not a positional/budget artifact.
39
+
40
+ ## How it works (in brief)
41
+
42
+ Two indexers cooperate:
43
+
44
+ - **Level 1 β€” Memory Indexer (every 64 steps):** a trained retriever scores the history's compressed keys and selects query-critical chunks β†’ a `resident_set` recalled from CPU to GPU.
45
+ - **Level 2 β€” Lightning Indexer (every step):** the native top-512 runs confined to the `resident_set`, at full speed inside cuda-graph.
46
+
47
+ The GPU/CPU split is the key: only recalled chunks + compressed indexer keys live on GPU; the bulk KV cache sits on CPU and is pulled on demand.
48
+
49
+ > For the full architecture, retriever math, hyperparameters and training details, see the **[paper](https://arxiv.org/abs/2606.09079)** and the **[GitHub code](https://github.com/libertywing/FlashMemory-Deepseek-V4)**.
50
+
51
+ ## Inference (PD-disaggregated)
52
+
53
+ Three servers β€” **P** (prefill), **D** (decode, all offload/recall), **router**. Clients hit the router at `http://<router>:31503/v1/chat/completions`.
54
+
55
+ ```bash
56
+ pip install -e sglang/python # also: pip install sgl_kernel==0.3.21
57
+ # download checkpoints/ from this HF repo
58
+
59
+ # D (decode) β€” GPU 0–7, 512K example:
60
+ TGT_CONC=60 TGT_CTX=524288 CTX_LEN=1100000 bash launch_decode.sh
61
+ # P (prefill) β€” GPU 0–7:
62
+ CTX_LEN=1100000 SWA_RATIO=0.1 HOST=<P_IP> bash launch_prefill.sh
63
+ # router:
64
+ PREFILL_IP=<P_IP> DECODE_IP=<D_IP> bash launch_router.sh
65
+ ```
66
+
67
+ All optimizations are env-gated (gate-off = exact DS-V4 baseline). Key switches:
68
+ `SGLANG_DECODE_SWAP_P`, `SGLANG_PATHP_INDEX_K_OFFLOAD`, `SGLANG_PATHP_SCORE_RESIDENT`,
69
+ `SGLANG_PATHP_PAGE_RECALL`, `SGLANG_PATHP_CUDAGRAPH`, `SGLANG_PATHP_ASYNC_RECALL`,
70
+ `SGLANG_PATHP_FUSED_REMAP`.
71
+
72
+ ## Checkpoint
73
+
74
+ Download `checkpoints/` into the repo root (default `top3_R930_joint.pt`, CSA layers 10/12/20).
75
+
76
+ ## License
77
+
78
+ MIT
79
+
80
+ ## Citation
81
+
82
+ ```bibtex
83
+ @article{wang2026flashmemory,
84
+ title = {FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention},
85
+ author = {Yan Wang and Qifan Zhang and Jiachen Yu and Tian Liang and Dongyang Ma and
86
+ Xiang Hu and Zibo Lin and Chunyang Li and Zhichao Wang and Jia Li and
87
+ Yujiu Yang and Haitao Mi and Dong Yu},
88
+ year = {2026},
89
+ journal = {arXiv preprint arXiv:2606.09079},
90
+ url = {https://arxiv.org/abs/2606.09079},
91
+ }
92
+ ```