Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
LordNeel
/
DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8
like
16
Text Generation
Safetensors
English
Chinese
doi:10.57967/hf/8744
vllm
deepseek_v4
deepseek-v4-flash
mixture-of-experts
Mixture of Experts
compressed-tensors
w4a16
gptq
fp8-block
mtp
speculative-decoding
multi-token-prediction
acti
blackwell
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
DeepSeek-V4-Flash-Acti-MTP-W4A16-FP8
156 GB
Ctrl+K
Ctrl+K
1 contributor
History:
19 commits
LordNeel
Add measured MTP draft-acceptance metrics (77.7%
@
128k
, 84.9% @262k prod, 93-96% code) + acceptance chart; update manifest
f561982
verified
12 days ago
charts
Add measured MTP draft-acceptance metrics (77.7% @128k, 84.9% @262k prod, 93-96% code) + acceptance chart; update manifest
12 days ago
docker
docker: end-to-end verified — runtime base is cudnn-devel (nvcc) + python3.12-dev + gcc for Triton JIT; image launches vLLM and serves /v1/models + /v1/chat/completions
2 months ago
.gitattributes
Safe
1.52 kB
initial commit
2 months ago
AGENTS.md
Safe
16.6 kB
Add AGENTS.md — end-to-end runbook for AI agents to set up + serve this model
2 months ago
MTP_QUANT_MANIFEST.json
Safe
1.4 kB
Add measured MTP draft-acceptance metrics (77.7% @128k, 84.9% @262k prod, 93-96% code) + acceptance chart; update manifest
12 days ago
MTP_QUANT_MANIFEST.md
Safe
1.04 kB
Add measured MTP draft-acceptance metrics (77.7% @128k, 84.9% @262k prod, 93-96% code) + acceptance chart; update manifest
12 days ago
README.md
Safe
21.1 kB
Add measured MTP draft-acceptance metrics (77.7% @128k, 84.9% @262k prod, 93-96% code) + acceptance chart; update manifest
12 days ago
config.json
Safe
12.9 kB
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
generation_config.json
Safe
174 Bytes
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
model-00001-of-00004.safetensors
Safe
50 GB
xet
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
model-00002-of-00004.safetensors
Safe
50 GB
xet
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
model-00003-of-00004.safetensors
Safe
50 GB
xet
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
model-00004-of-00004.safetensors
Safe
2.48 GB
xet
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
model-mtp-w4a16.safetensors
Safe
3.55 GB
xet
v2: real GPTQ on MTP routed experts (replaces v1 RTN). Added MTP_QUANT_MANIFEST is_final_quality_preserving=true.
2 months ago
model.safetensors.index.json
Safe
8.68 MB
v2: real GPTQ on MTP routed experts (replaces v1 RTN). Added MTP_QUANT_MANIFEST is_final_quality_preserving=true.
2 months ago
recipe.yaml
Safe
1.97 kB
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
tokenizer.json
Safe
10.1 MB
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago
tokenizer_config.json
Safe
397 Bytes
Initial release: pasta-paul base + Acti MTP block (RTN INT4 routed experts, FP8_BLOCK attn)
2 months ago