File size: 16,582 Bytes
8c958bc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
# Agent runbook — set up `LordNeel/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8` end-to-end

**For**: an autonomous code agent (Claude Code, Codex, Cursor, Aider, etc.) running on a fresh-ish Linux box with the right hardware. Stop and ask the human only when explicitly told.

**Goal**: a vLLM OpenAI-compatible server on port 8000 serving the MTP-enabled model at the validated 524k profile, with smoke tests passing.

---

## Phase 0 — verify hardware (HARD STOP if not satisfied)

Run `nvidia-smi --query-gpu=name,memory.total --format=csv,noheader` and `nvidia-smi topo -m`.

Validated SKUs (any of these works without code changes):
-**RTX PRO 6000 Blackwell Max-Q Workstation Edition** (96 GB ea, sm_120, no NVLink) ← reference platform
- 2× RTX PRO 6000 Blackwell Server (96 GB ea, sm_120, with NVLink) ← also expected to work
- 2× DGX Spark / GB10 (sm_121)
- 8× H200 SXM (sm_90)

**Stop and ask the human if any of the following is true**:
- Fewer than 2 GPUs with ≥ 96 GB each (TP=2 with 524k KV cache won't fit)
- More than 2 visible GPUs and you'd need to pick which 2 (set `CUDA_VISIBLE_DEVICES` deliberately, don't guess)
- GPU is on a topology where `nvidia-smi topo -m` reports `SYS` (cross-NUMA) — this can cause more comm overhead; not a blocker but flag it

Detect Max-Q-specific quirks: if the GPU name contains "Max-Q" you MUST pass `--disable-custom-all-reduce` later. The serve script already does this.

---

## Phase 1 — system prereqs (no sudo)

Working dir: pick `$HOME/dsv4-local` (the reference layout). All paths in this doc assume that.

```bash
mkdir -p $HOME/dsv4-local
cd $HOME/dsv4-local
```

Required:
- Python 3.12 (system or via miniforge — agent's choice; reference uses miniforge3)
- Driver: NVIDIA 580.x or newer (check `nvidia-smi`)
- gcc/g++ 13.x (system) — used as host compiler for nvcc
- Git, curl, jq, ca-certificates (apt should already have these)

Forbidden without explicit human approval:
- Modifying `/etc`, kernel modules, drivers, or system CUDA
- Installing system packages with sudo (use conda for CUDA toolkit instead)
- Touching `/proc/sys/*` or `nvidia-smi -pl`/`-lgc` (those need sudo)

---

## Phase 2 — local CUDA toolkit via conda (no sudo)

```bash
# 2a. Miniforge3 (only if conda not already available)
test -x $HOME/dsv4-local/miniforge3/bin/conda || (
  curl -fsSLO https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
  bash Miniforge3-Linux-x86_64.sh -b -p $HOME/dsv4-local/miniforge3
)
$HOME/dsv4-local/miniforge3/bin/conda init bash >/dev/null

# 2b. CUDA toolkit 12.9 in a self-contained env
$HOME/dsv4-local/miniforge3/bin/conda create -y -p $HOME/dsv4-local/conda-cuda \
  -c nvidia -c conda-forge cuda-toolkit=12.9 ccache

# 2c. Verify
$HOME/dsv4-local/conda-cuda/bin/nvcc --version | tail -3
```

Common gotcha: conda's default put CUDA headers under `targets/x86_64-linux/include/` instead of `include/`. The vLLM build will fail with "Could NOT find CUDA". Fix by symlinking once:

```bash
CUDA_HOME=$HOME/dsv4-local/conda-cuda
for f in $(ls $CUDA_HOME/targets/x86_64-linux/include 2>/dev/null); do
  test -e $CUDA_HOME/include/$f || ln -s ../targets/x86_64-linux/include/$f $CUDA_HOME/include/$f
done
```

---

## Phase 3 — Python venv + base deps

```bash
mkdir -p $HOME/dsv4-local/venvs
python3.12 -m venv $HOME/dsv4-local/venvs/vllm-dsv4
source $HOME/dsv4-local/venvs/vllm-dsv4/bin/activate
python -m pip install --upgrade pip wheel setuptools
pip install \
  "torch==2.11.0+cu128" --index-url https://download.pytorch.org/whl/cu128
pip install \
  "transformers==5.8.0" \
  "compressed-tensors==0.15.0.1" \
  "safetensors>=0.7" \
  "huggingface_hub>=1.14" \
  hf_transfer \
  "openai" \
  "rich" \
  "datasets"
```

(If conda gcc-14 is on PATH, force system gcc/g++ for the upcoming vLLM build:
`export CC=/usr/bin/gcc CXX=/usr/bin/g++ CUDAHOSTCXX=/usr/bin/g++ NVCC_PREPEND_FLAGS="-ccbin /usr/bin/g++"`.)

---

## Phase 4 — patched vLLM fork

```bash
mkdir -p $HOME/dsv4-local/src
cd $HOME/dsv4-local/src
git clone https://github.com/jasl/vllm.git vllm
cd vllm
# Validated pin. This is itself a cherry-pick of neuralmagic/vllm@f910a73 ("support
# ct quantization" — Kyle Sayers); pinning here means you do NOT need the extra
# cherry-pick step that older versions of this runbook documented.
git checkout b158e5001f097b193b4fa8c4fc7bb32bb4e32e9e

# pasta-paul's packed_modules_mapping patch (adds the class-level dict on
# DeepseekV4ForCausalLM that the cherry-pick references but doesn't define).
curl -fsSL https://raw.githubusercontent.com/pasta-paul/dsv4-flash-w4a16-fp8/main/scripts/patch_v4_packed_mapping.py \
  -o /tmp/patch_packed.py
python /tmp/patch_packed.py vllm/model_executor/models/deepseek_v4.py
```

### Acti's MTP-loader patches (~30 lines, three edits in `vllm/model_executor/models/deepseek_v4_mtp.py`)

Apply each of these or you'll hit the six load errors documented in the model card.

#### Patch 4a — pass `prefix=` to e_proj / h_proj

In `class DeepSeekV4MultiTokenPredictorLayer.__init__`, change both ReplicatedLinear constructions:

```python
self.e_proj = ReplicatedLinear(
    config.hidden_size, config.hidden_size,
    bias=False, return_bias=False,
    quant_config=quant_config,
    prefix=f"{prefix}.e_proj",          # <-- ADD THIS
)
self.h_proj = ReplicatedLinear(
    config.hidden_size, config.hidden_size,
    bias=False, return_bias=False,
    quant_config=quant_config,
    prefix=f"{prefix}.h_proj",          # <-- ADD THIS
)
```

#### Patch 4b — `packed_modules_mapping` on `DeepSeekV4MTP`

Add a class attribute right above `__init__`:

```python
class DeepSeekV4MTP(nn.Module):
    packed_modules_mapping = {                          # <-- ADD THIS BLOCK
        "fused_wqa_wkv": ["wq_a", "wkv"],
        "fused_wkv_wgate": ["wkv", "wgate"],
        "gate_up_proj": ["w1", "w3"],
    }

    def __init__(self, *, vllm_config, prefix=""):
        ...
```

#### Patch 4c — `.weight_scale` (no `_inv`) in MTP loader

In `DeepSeekV4MTP.load_weights` find the suffix pick logic and change:

```python
suffix = (
    expert_scale_suffix
    if _EXPERT_SCALE_RE.search(name)
    else ".weight_scale"                # <-- was ".weight_scale_inv"
)
```

If anything goes sideways, sanity-check against the local diff after applying — `git diff vllm/model_executor/models/deepseek_v4_mtp.py` should show three hunks: (a) `prefix=f"{prefix}.e_proj"` and `prefix=f"{prefix}.h_proj"` added in the `DeepSeekV4MultiTokenPredictorLayer.__init__` constructor calls, (b) the `packed_modules_mapping` dict added as a class attribute on `DeepSeekV4MTP` immediately above its `__init__`, and (c) the `else ".weight_scale_inv"` line in `DeepSeekV4MTP.load_weights` changed to `else ".weight_scale"`.

### Build vLLM

```bash
cd $HOME/dsv4-local/src/vllm
export CUDA_HOME=$HOME/dsv4-local/conda-cuda
export PATH=/usr/bin:$CUDA_HOME/bin:$PATH
export LD_LIBRARY_PATH=$CUDA_HOME/lib:$LD_LIBRARY_PATH
export CPATH=$CUDA_HOME/include:$CPATH
export TORCH_CUDA_ARCH_LIST=12.0a
export CUDA_ARCH_LIST=120a
export CC=/usr/bin/gcc CXX=/usr/bin/g++ CUDAHOSTCXX=/usr/bin/g++
export NVCC_PREPEND_FLAGS="-ccbin /usr/bin/g++"
pip install -e .            # ~20-30 min on a fast box
python -c "import vllm; print(vllm.__version__)"
python -c "from vllm.model_executor.models.deepseek_v4 import DeepseekV4ForCausalLM; print('OK')"
```

If the build fails: the most common causes are (a) host gcc too new for nvcc 12.9, fixed by exporting CC/CXX above; (b) Triton can't find `pyconfig.h`, fix by `export CPATH=$CUDA_HOME/include:/usr/include/x86_64-linux-gnu/python3.12:$CPATH`. Stop and ask the human if those don't fix it — don't pivot to vanilla vLLM.

---

## Phase 5 — model download (143 GB)

```bash
mkdir -p $HOME/dsv4-local/models
HF_HUB_ENABLE_HF_TRANSFER=1 hf download \
  LordNeel/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8 \
  --local-dir $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8
```

This is the v2 GPTQ release. Expect ~30-60 min on a 100 MB/s connection. The 4 base shards (~50 GB each) will hash-dedupe against pasta-paul's repo if you've already downloaded that.

Verification:

```bash
ls $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/*.safetensors | wc -l    # → 5
du -sh $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8                       # → ~146G
python -c "
from safetensors import safe_open
import json
idx = json.load(open('$HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/model.safetensors.index.json'))
keys = list(idx['weight_map'])
print('total tensors:', len(keys))
print('mtp.0.* tensors:', sum(1 for k in keys if k.startswith('mtp.')))
print('mtp.0.ffn.experts.0.w1.weight_packed present:',
      'mtp.0.ffn.experts.0.w1.weight_packed' in keys)
"
```

Expected: ~102k tensors, 2338 mtp.* tensors, w1.weight_packed True.

---

## Phase 6 — write serve script + env

```bash
mkdir -p $HOME/dsv4-local/scripts
cat > $HOME/dsv4-local/scripts/serve_524k.sh <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
ROOT=$HOME/dsv4-local
source "$ROOT/venvs/vllm-dsv4/bin/activate"

export CUDA_VISIBLE_DEVICES=0,1
export CUDA_HOME=$ROOT/conda-cuda
export PATH=/usr/bin:$CUDA_HOME/bin:$PATH
export TRITON_PTXAS_PATH=$CUDA_HOME/bin/ptxas
export LD_LIBRARY_PATH=$CUDA_HOME/lib:${LD_LIBRARY_PATH:-}
export CPATH=$CUDA_HOME/include:${CPATH:-}
export TORCH_CUDA_ARCH_LIST=12.0a
export CUDA_ARCH_LIST=120a

export VLLM_USE_FLASHINFER_SAMPLER=0
export VLLM_ENABLE_DEEPSEEK_V4_SPARSE_MLA_WARMUP=0
export VLLM_ENGINE_READY_TIMEOUT_S=3600
export PYTHONUNBUFFERED=1

# NCCL: validated for Max-Q PCIe topology
export NCCL_DEBUG=WARN
export NCCL_P2P_DISABLE=1
export NCCL_SHM_DISABLE=0
export NCCL_IB_DISABLE=1
# small-msg latency tuning (drops TTFT 154ms -> 91ms on Max-Q at zero decode cost)
export NCCL_PROTO=LL
export NCCL_ALGO=Ring
export NCCL_MIN_NCHANNELS=8
export NCCL_NTHREADS=512

MODEL=$ROOT/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8

exec vllm serve "$MODEL" \
  --served-model-name deepseek-v4-flash deepseek-v4-flash-mtp DSV4-W4A16-FP8 deepseek-ai/DeepSeek-V4-Flash \
  --tensor-parallel-size 2 \
  --kv-cache-dtype fp8 \
  --block-size 256 \
  --max-model-len 524288 \
  --max-num-seqs 2 --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.93 \
  --tokenizer-mode deepseek_v4 \
  --tool-call-parser deepseek_v4 --enable-auto-tool-choice \
  --reasoning-parser deepseek_v4 \
  --trust-remote-code \
  --disable-custom-all-reduce \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
  --host 0.0.0.0 --port 8000
EOF
chmod +x $HOME/dsv4-local/scripts/serve_524k.sh
```

**Hardware-specific edits to make**:
- If running on RTX PRO 6000 **Server** (with NVLink): you can drop `--disable-custom-all-reduce` AND you can drop the `NCCL_P2P_DISABLE`/`NCCL_PROTO=LL` block. CustomAllreduce works on NVLink and is a bit faster.
- If running on H200: same as Server (NVLink available); also you'll likely run TP=4 or higher — adjust `--tensor-parallel-size` and `CUDA_VISIBLE_DEVICES` accordingly. **Do NOT use TP≥4 with W4A16 quants** — there's an upstream MoE scale-sharding bug ([vllm-project/vllm#41511](https://github.com/vllm-project/vllm/issues/41511)).
- If running on DGX Spark / GB10: same as Max-Q (no NVLink). Keep all flags.

---

## Phase 7 — start the server, wait for ready

```bash
setsid bash $HOME/dsv4-local/scripts/serve_524k.sh > $HOME/dsv4-local/serve.log 2>&1 &
echo "started pid=$!"

# Wait for readiness (≤5 min cold)
until grep -q "Application startup complete" $HOME/dsv4-local/serve.log 2>/dev/null; do
  if grep -qE "(Traceback|RuntimeError|Engine core init.*failed|Worker proc.*died)" $HOME/dsv4-local/serve.log; then
    echo "FAILED — see $HOME/dsv4-local/serve.log"
    tail -50 $HOME/dsv4-local/serve.log
    exit 1
  fi
  sleep 2
done
curl -fsS http://127.0.0.1:8000/health -o /dev/null && echo "health OK"
```

Expected timing on a 2× Max-Q: weight load ~30 s, KV cache + warmup ~30 s, graph capture ~5 s, total ~75 s wall.

### Failure modes to recognize and remediate

| Symptom in log | Cause | Fix |
|---|---|---|
| `NotImplementedError: DeepSeekV4 requires FP8 attention quantization` | Loaded a non-FP8-attn variant (e.g. Intel/AutoRound) | Use the LordNeel repo, NOT a different DSV4 quant |
| `Unable to find matching target for ''` | Patch 4a (e_proj/h_proj prefix) missing | Apply Patch 4a |
| `KeyError: attn.fused_wqa_wkv.weight_scale_inv` (mtp_block) | Patch 4b (`packed_modules_mapping` on `DeepSeekV4MTP`) missing | Apply Patch 4b |
| `ValueError: Unable to find matching target for model.layers.43.e_proj` | Out-of-date config.json (ignore list missing layer 43 prefixes) | Re-pull the model — that bug is fixed in v2 |
| Hangs at `gpu_worker.py:619 CUDA graph pool memory` then loops `shm_broadcast.py:681` | CustomAllreduce deadlock on Max-Q | Add `--disable-custom-all-reduce` |
| `KeyError: experts.w13_weight` | Loaded the v1 RTN ckpt by mistake | Pull v2 (current main branch on Hub) |

---

## Phase 8 — smoke tests

```bash
# Non-streaming chat
curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" -H "Authorization: Bearer EMPTY" \
  --data '{
    "model":"deepseek-v4-flash","stream":false,"max_tokens":64,"temperature":0.0,
    "messages":[{"role":"user","content":"Say OK."}]
  }' | python -c "import json,sys; d=json.load(sys.stdin); print(d['choices'][0]['message'].get('content') or '(empty content — check reasoning_content)')"

# Streaming
curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" -H "Authorization: Bearer EMPTY" \
  --data '{
    "model":"deepseek-v4-flash","stream":true,"max_tokens":32,"temperature":0.0,
    "messages":[{"role":"user","content":"hi"}]
  }' | head -20
```

If `content` is empty even on simple prompts, that's the deepseek_v4 reasoning-parser putting output into `reasoning_content`. Read both fields:

```python
msg = response["choices"][0]["message"]
text = msg.get("content") or msg.get("reasoning_content") or ""
```

---

## Phase 9 — hand-off

Once Phase 8 passes, the agent should:

1. Print the listen address: `http://<host>:8000/v1`
2. Print the served model names: `deepseek-v4-flash`, `deepseek-v4-flash-mtp`, `DSV4-W4A16-FP8`, `deepseek-ai/DeepSeek-V4-Flash`
3. Note that the API key is the literal string `EMPTY` (vLLM's default for unauthenticated)
4. Print expected single-stream decode TPS for the configured profile (524k → ~85; 128k → ~111)
5. Tell the human about reasoning_parser and `--reasoning-parser deepseek_v4` (their client may need to read both `content` and `reasoning_content`)

Do NOT proceed to fine-tuning, hot-swapping, or modifying the running server without explicit human approval.

---

## Phase 10 — optional: GPU clock unlock (sudo, +0-1% TPS)

If on Max-Q and the human wants the last 1%:

```bash
# REQUIRES SUDO. Stop and ask the human first.
sudo nvidia-smi -pm 1
sudo nvidia-smi -pl 325 -i 0; sudo nvidia-smi -pl 325 -i 1
sudo nvidia-smi -lgc 2700 -i 0; sudo nvidia-smi -lgc 2700 -i 1
# revert: sudo nvidia-smi -rgc -i 0,1; sudo nvidia-smi -pl 300 -i 0,1
```

Honest expectation: Max-Q firmware caps the achievable boost at ~2325 MHz under sustained 99% SM load regardless of the lock value. The +1% gain is real but small. Skip this unless the human asks for it.

---

## Quick reference: where things live in the reference layout

```
$HOME/dsv4-local/
├── conda-cuda/                      # CUDA 12.9 toolkit (no system change)
├── miniforge3/                      # conda installer
├── venvs/vllm-dsv4/                 # the only Python venv you need
├── src/vllm/                        # patched jasl/vllm fork
├── models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/
└── scripts/serve_524k.sh
```

## Questions an agent should ask the human (only when triggered)

| Trigger | Ask |
|---|---|
| < 2× 96 GB GPUs found | "Hardware doesn't fit a 524k TP=2 profile. Want to try smaller profiles or stop?" |
| Build fails with cuda/nvcc errors after the documented fixes | "Build broken. Want me to pivot to vanilla vLLM (won't load this model) or stop?" |
| Phase 7 hangs > 5 min with no log progress | "Server isn't responding; want me to dump py-spy stacks (needs sudo) or kill and retry with `--enforce-eager`?" |
| Phase 10 (sudo clock unlock) | always ask before running |
| User asks about TP > 2 | "Upstream W4A16 MoE scale-sharding bug — issue #41511. TP=4 will OOM/error. Want to stop or proceed anyway?" |