Agent runbook — set up LordNeel/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8 end-to-end
For: an autonomous code agent (Claude Code, Codex, Cursor, Aider, etc.) running on a fresh-ish Linux box with the right hardware. Stop and ask the human only when explicitly told.
Goal: a vLLM OpenAI-compatible server on port 8000 serving the MTP-enabled model at the validated 524k profile, with smoke tests passing.
Phase 0 — verify hardware (HARD STOP if not satisfied)
Run nvidia-smi --query-gpu=name,memory.total --format=csv,noheader and nvidia-smi topo -m.
Validated SKUs (any of these works without code changes):
- 2× RTX PRO 6000 Blackwell Max-Q Workstation Edition (96 GB ea, sm_120, no NVLink) ← reference platform
- 2× RTX PRO 6000 Blackwell Server (96 GB ea, sm_120, with NVLink) ← also expected to work
- 2× DGX Spark / GB10 (sm_121)
- 8× H200 SXM (sm_90)
Stop and ask the human if any of the following is true:
- Fewer than 2 GPUs with ≥ 96 GB each (TP=2 with 524k KV cache won't fit)
- More than 2 visible GPUs and you'd need to pick which 2 (set
CUDA_VISIBLE_DEVICESdeliberately, don't guess) - GPU is on a topology where
nvidia-smi topo -mreportsSYS(cross-NUMA) — this can cause more comm overhead; not a blocker but flag it
Detect Max-Q-specific quirks: if the GPU name contains "Max-Q" you MUST pass --disable-custom-all-reduce later. The serve script already does this.
Phase 1 — system prereqs (no sudo)
Working dir: pick $HOME/dsv4-local (the reference layout). All paths in this doc assume that.
mkdir -p $HOME/dsv4-local
cd $HOME/dsv4-local
Required:
- Python 3.12 (system or via miniforge — agent's choice; reference uses miniforge3)
- Driver: NVIDIA 580.x or newer (check
nvidia-smi) - gcc/g++ 13.x (system) — used as host compiler for nvcc
- Git, curl, jq, ca-certificates (apt should already have these)
Forbidden without explicit human approval:
- Modifying
/etc, kernel modules, drivers, or system CUDA - Installing system packages with sudo (use conda for CUDA toolkit instead)
- Touching
/proc/sys/*ornvidia-smi -pl/-lgc(those need sudo)
Phase 2 — local CUDA toolkit via conda (no sudo)
# 2a. Miniforge3 (only if conda not already available)
test -x $HOME/dsv4-local/miniforge3/bin/conda || (
curl -fsSLO https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
bash Miniforge3-Linux-x86_64.sh -b -p $HOME/dsv4-local/miniforge3
)
$HOME/dsv4-local/miniforge3/bin/conda init bash >/dev/null
# 2b. CUDA toolkit 12.9 in a self-contained env
$HOME/dsv4-local/miniforge3/bin/conda create -y -p $HOME/dsv4-local/conda-cuda \
-c nvidia -c conda-forge cuda-toolkit=12.9 ccache
# 2c. Verify
$HOME/dsv4-local/conda-cuda/bin/nvcc --version | tail -3
Common gotcha: conda's default put CUDA headers under targets/x86_64-linux/include/ instead of include/. The vLLM build will fail with "Could NOT find CUDA". Fix by symlinking once:
CUDA_HOME=$HOME/dsv4-local/conda-cuda
for f in $(ls $CUDA_HOME/targets/x86_64-linux/include 2>/dev/null); do
test -e $CUDA_HOME/include/$f || ln -s ../targets/x86_64-linux/include/$f $CUDA_HOME/include/$f
done
Phase 3 — Python venv + base deps
mkdir -p $HOME/dsv4-local/venvs
python3.12 -m venv $HOME/dsv4-local/venvs/vllm-dsv4
source $HOME/dsv4-local/venvs/vllm-dsv4/bin/activate
python -m pip install --upgrade pip wheel setuptools
pip install \
"torch==2.11.0+cu128" --index-url https://download.pytorch.org/whl/cu128
pip install \
"transformers==5.8.0" \
"compressed-tensors==0.15.0.1" \
"safetensors>=0.7" \
"huggingface_hub>=1.14" \
hf_transfer \
"openai" \
"rich" \
"datasets"
(If conda gcc-14 is on PATH, force system gcc/g++ for the upcoming vLLM build:
export CC=/usr/bin/gcc CXX=/usr/bin/g++ CUDAHOSTCXX=/usr/bin/g++ NVCC_PREPEND_FLAGS="-ccbin /usr/bin/g++".)
Phase 4 — patched vLLM fork
mkdir -p $HOME/dsv4-local/src
cd $HOME/dsv4-local/src
git clone https://github.com/jasl/vllm.git vllm
cd vllm
# Validated pin. This is itself a cherry-pick of neuralmagic/vllm@f910a73 ("support
# ct quantization" — Kyle Sayers); pinning here means you do NOT need the extra
# cherry-pick step that older versions of this runbook documented.
git checkout b158e5001f097b193b4fa8c4fc7bb32bb4e32e9e
# pasta-paul's packed_modules_mapping patch (adds the class-level dict on
# DeepseekV4ForCausalLM that the cherry-pick references but doesn't define).
curl -fsSL https://raw.githubusercontent.com/pasta-paul/dsv4-flash-w4a16-fp8/main/scripts/patch_v4_packed_mapping.py \
-o /tmp/patch_packed.py
python /tmp/patch_packed.py vllm/model_executor/models/deepseek_v4.py
Acti's MTP-loader patches (~30 lines, three edits in vllm/model_executor/models/deepseek_v4_mtp.py)
Apply each of these or you'll hit the six load errors documented in the model card.
Patch 4a — pass prefix= to e_proj / h_proj
In class DeepSeekV4MultiTokenPredictorLayer.__init__, change both ReplicatedLinear constructions:
self.e_proj = ReplicatedLinear(
config.hidden_size, config.hidden_size,
bias=False, return_bias=False,
quant_config=quant_config,
prefix=f"{prefix}.e_proj", # <-- ADD THIS
)
self.h_proj = ReplicatedLinear(
config.hidden_size, config.hidden_size,
bias=False, return_bias=False,
quant_config=quant_config,
prefix=f"{prefix}.h_proj", # <-- ADD THIS
)
Patch 4b — packed_modules_mapping on DeepSeekV4MTP
Add a class attribute right above __init__:
class DeepSeekV4MTP(nn.Module):
packed_modules_mapping = { # <-- ADD THIS BLOCK
"fused_wqa_wkv": ["wq_a", "wkv"],
"fused_wkv_wgate": ["wkv", "wgate"],
"gate_up_proj": ["w1", "w3"],
}
def __init__(self, *, vllm_config, prefix=""):
...
Patch 4c — .weight_scale (no _inv) in MTP loader
In DeepSeekV4MTP.load_weights find the suffix pick logic and change:
suffix = (
expert_scale_suffix
if _EXPERT_SCALE_RE.search(name)
else ".weight_scale" # <-- was ".weight_scale_inv"
)
If anything goes sideways, sanity-check against the local diff after applying — git diff vllm/model_executor/models/deepseek_v4_mtp.py should show three hunks: (a) prefix=f"{prefix}.e_proj" and prefix=f"{prefix}.h_proj" added in the DeepSeekV4MultiTokenPredictorLayer.__init__ constructor calls, (b) the packed_modules_mapping dict added as a class attribute on DeepSeekV4MTP immediately above its __init__, and (c) the else ".weight_scale_inv" line in DeepSeekV4MTP.load_weights changed to else ".weight_scale".
Build vLLM
cd $HOME/dsv4-local/src/vllm
export CUDA_HOME=$HOME/dsv4-local/conda-cuda
export PATH=/usr/bin:$CUDA_HOME/bin:$PATH
export LD_LIBRARY_PATH=$CUDA_HOME/lib:$LD_LIBRARY_PATH
export CPATH=$CUDA_HOME/include:$CPATH
export TORCH_CUDA_ARCH_LIST=12.0a
export CUDA_ARCH_LIST=120a
export CC=/usr/bin/gcc CXX=/usr/bin/g++ CUDAHOSTCXX=/usr/bin/g++
export NVCC_PREPEND_FLAGS="-ccbin /usr/bin/g++"
pip install -e . # ~20-30 min on a fast box
python -c "import vllm; print(vllm.__version__)"
python -c "from vllm.model_executor.models.deepseek_v4 import DeepseekV4ForCausalLM; print('OK')"
If the build fails: the most common causes are (a) host gcc too new for nvcc 12.9, fixed by exporting CC/CXX above; (b) Triton can't find pyconfig.h, fix by export CPATH=$CUDA_HOME/include:/usr/include/x86_64-linux-gnu/python3.12:$CPATH. Stop and ask the human if those don't fix it — don't pivot to vanilla vLLM.
Phase 5 — model download (143 GB)
mkdir -p $HOME/dsv4-local/models
HF_HUB_ENABLE_HF_TRANSFER=1 hf download \
LordNeel/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8 \
--local-dir $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8
This is the v2 GPTQ release. Expect 30-60 min on a 100 MB/s connection. The 4 base shards (50 GB each) will hash-dedupe against pasta-paul's repo if you've already downloaded that.
Verification:
ls $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/*.safetensors | wc -l # → 5
du -sh $HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8 # → ~146G
python -c "
from safetensors import safe_open
import json
idx = json.load(open('$HOME/dsv4-local/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/model.safetensors.index.json'))
keys = list(idx['weight_map'])
print('total tensors:', len(keys))
print('mtp.0.* tensors:', sum(1 for k in keys if k.startswith('mtp.')))
print('mtp.0.ffn.experts.0.w1.weight_packed present:',
'mtp.0.ffn.experts.0.w1.weight_packed' in keys)
"
Expected: ~102k tensors, 2338 mtp.* tensors, w1.weight_packed True.
Phase 6 — write serve script + env
mkdir -p $HOME/dsv4-local/scripts
cat > $HOME/dsv4-local/scripts/serve_524k.sh <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
ROOT=$HOME/dsv4-local
source "$ROOT/venvs/vllm-dsv4/bin/activate"
export CUDA_VISIBLE_DEVICES=0,1
export CUDA_HOME=$ROOT/conda-cuda
export PATH=/usr/bin:$CUDA_HOME/bin:$PATH
export TRITON_PTXAS_PATH=$CUDA_HOME/bin/ptxas
export LD_LIBRARY_PATH=$CUDA_HOME/lib:${LD_LIBRARY_PATH:-}
export CPATH=$CUDA_HOME/include:${CPATH:-}
export TORCH_CUDA_ARCH_LIST=12.0a
export CUDA_ARCH_LIST=120a
export VLLM_USE_FLASHINFER_SAMPLER=0
export VLLM_ENABLE_DEEPSEEK_V4_SPARSE_MLA_WARMUP=0
export VLLM_ENGINE_READY_TIMEOUT_S=3600
export PYTHONUNBUFFERED=1
# NCCL: validated for Max-Q PCIe topology
export NCCL_DEBUG=WARN
export NCCL_P2P_DISABLE=1
export NCCL_SHM_DISABLE=0
export NCCL_IB_DISABLE=1
# small-msg latency tuning (drops TTFT 154ms -> 91ms on Max-Q at zero decode cost)
export NCCL_PROTO=LL
export NCCL_ALGO=Ring
export NCCL_MIN_NCHANNELS=8
export NCCL_NTHREADS=512
MODEL=$ROOT/models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8
exec vllm serve "$MODEL" \
--served-model-name deepseek-v4-flash deepseek-v4-flash-mtp DSV4-W4A16-FP8 deepseek-ai/DeepSeek-V4-Flash \
--tensor-parallel-size 2 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 524288 \
--max-num-seqs 2 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.93 \
--tokenizer-mode deepseek_v4 \
--tool-call-parser deepseek_v4 --enable-auto-tool-choice \
--reasoning-parser deepseek_v4 \
--trust-remote-code \
--disable-custom-all-reduce \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
--host 0.0.0.0 --port 8000
EOF
chmod +x $HOME/dsv4-local/scripts/serve_524k.sh
Hardware-specific edits to make:
- If running on RTX PRO 6000 Server (with NVLink): you can drop
--disable-custom-all-reduceAND you can drop theNCCL_P2P_DISABLE/NCCL_PROTO=LLblock. CustomAllreduce works on NVLink and is a bit faster. - If running on H200: same as Server (NVLink available); also you'll likely run TP=4 or higher — adjust
--tensor-parallel-sizeandCUDA_VISIBLE_DEVICESaccordingly. Do NOT use TP≥4 with W4A16 quants — there's an upstream MoE scale-sharding bug (vllm-project/vllm#41511). - If running on DGX Spark / GB10: same as Max-Q (no NVLink). Keep all flags.
Phase 7 — start the server, wait for ready
setsid bash $HOME/dsv4-local/scripts/serve_524k.sh > $HOME/dsv4-local/serve.log 2>&1 &
echo "started pid=$!"
# Wait for readiness (≤5 min cold)
until grep -q "Application startup complete" $HOME/dsv4-local/serve.log 2>/dev/null; do
if grep -qE "(Traceback|RuntimeError|Engine core init.*failed|Worker proc.*died)" $HOME/dsv4-local/serve.log; then
echo "FAILED — see $HOME/dsv4-local/serve.log"
tail -50 $HOME/dsv4-local/serve.log
exit 1
fi
sleep 2
done
curl -fsS http://127.0.0.1:8000/health -o /dev/null && echo "health OK"
Expected timing on a 2× Max-Q: weight load ~30 s, KV cache + warmup ~30 s, graph capture ~5 s, total ~75 s wall.
Failure modes to recognize and remediate
| Symptom in log | Cause | Fix |
|---|---|---|
NotImplementedError: DeepSeekV4 requires FP8 attention quantization |
Loaded a non-FP8-attn variant (e.g. Intel/AutoRound) | Use the LordNeel repo, NOT a different DSV4 quant |
Unable to find matching target for '' |
Patch 4a (e_proj/h_proj prefix) missing | Apply Patch 4a |
KeyError: attn.fused_wqa_wkv.weight_scale_inv (mtp_block) |
Patch 4b (packed_modules_mapping on DeepSeekV4MTP) missing |
Apply Patch 4b |
ValueError: Unable to find matching target for model.layers.43.e_proj |
Out-of-date config.json (ignore list missing layer 43 prefixes) | Re-pull the model — that bug is fixed in v2 |
Hangs at gpu_worker.py:619 CUDA graph pool memory then loops shm_broadcast.py:681 |
CustomAllreduce deadlock on Max-Q | Add --disable-custom-all-reduce |
KeyError: experts.w13_weight |
Loaded the v1 RTN ckpt by mistake | Pull v2 (current main branch on Hub) |
Phase 8 — smoke tests
# Non-streaming chat
curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" -H "Authorization: Bearer EMPTY" \
--data '{
"model":"deepseek-v4-flash","stream":false,"max_tokens":64,"temperature":0.0,
"messages":[{"role":"user","content":"Say OK."}]
}' | python -c "import json,sys; d=json.load(sys.stdin); print(d['choices'][0]['message'].get('content') or '(empty content — check reasoning_content)')"
# Streaming
curl -sS -X POST http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" -H "Authorization: Bearer EMPTY" \
--data '{
"model":"deepseek-v4-flash","stream":true,"max_tokens":32,"temperature":0.0,
"messages":[{"role":"user","content":"hi"}]
}' | head -20
If content is empty even on simple prompts, that's the deepseek_v4 reasoning-parser putting output into reasoning_content. Read both fields:
msg = response["choices"][0]["message"]
text = msg.get("content") or msg.get("reasoning_content") or ""
Phase 9 — hand-off
Once Phase 8 passes, the agent should:
- Print the listen address:
http://<host>:8000/v1 - Print the served model names:
deepseek-v4-flash,deepseek-v4-flash-mtp,DSV4-W4A16-FP8,deepseek-ai/DeepSeek-V4-Flash - Note that the API key is the literal string
EMPTY(vLLM's default for unauthenticated) - Print expected single-stream decode TPS for the configured profile (524k → ~85; 128k → ~111)
- Tell the human about reasoning_parser and
--reasoning-parser deepseek_v4(their client may need to read bothcontentandreasoning_content)
Do NOT proceed to fine-tuning, hot-swapping, or modifying the running server without explicit human approval.
Phase 10 — optional: GPU clock unlock (sudo, +0-1% TPS)
If on Max-Q and the human wants the last 1%:
# REQUIRES SUDO. Stop and ask the human first.
sudo nvidia-smi -pm 1
sudo nvidia-smi -pl 325 -i 0; sudo nvidia-smi -pl 325 -i 1
sudo nvidia-smi -lgc 2700 -i 0; sudo nvidia-smi -lgc 2700 -i 1
# revert: sudo nvidia-smi -rgc -i 0,1; sudo nvidia-smi -pl 300 -i 0,1
Honest expectation: Max-Q firmware caps the achievable boost at ~2325 MHz under sustained 99% SM load regardless of the lock value. The +1% gain is real but small. Skip this unless the human asks for it.
Quick reference: where things live in the reference layout
$HOME/dsv4-local/
├── conda-cuda/ # CUDA 12.9 toolkit (no system change)
├── miniforge3/ # conda installer
├── venvs/vllm-dsv4/ # the only Python venv you need
├── src/vllm/ # patched jasl/vllm fork
├── models/DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8/
└── scripts/serve_524k.sh
Questions an agent should ask the human (only when triggered)
| Trigger | Ask |
|---|---|
| < 2× 96 GB GPUs found | "Hardware doesn't fit a 524k TP=2 profile. Want to try smaller profiles or stop?" |
| Build fails with cuda/nvcc errors after the documented fixes | "Build broken. Want me to pivot to vanilla vLLM (won't load this model) or stop?" |
| Phase 7 hangs > 5 min with no log progress | "Server isn't responding; want me to dump py-spy stacks (needs sudo) or kill and retry with --enforce-eager?" |
| Phase 10 (sudo clock unlock) | always ask before running |
| User asks about TP > 2 | "Upstream W4A16 MoE scale-sharding bug — issue #41511. TP=4 will OOM/error. Want to stop or proceed anyway?" |