Instructions to use WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp") model = AutoModelForCausalLM.from_pretrained("WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp
- SGLang
How to use WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp with Docker Model Runner:
docker model run hf.co/WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp
Experimental — most likely broken (renamed to
-expon 2026-08-12). This variant was produced by the second cascade step (REAM160 → REAM128 → REAM96). Severe multi-turn degradation was observed: generation typically collapses into gibberish from the second request onward. No quality evidence was ever collected for this release. Kept for archival reference only; do not use for inference. A rebuilt REAM96 (single-step prune from the base model, new calibration) will be published at the original repo idWaveCut/DeepSeek-V4-Flash-0731-REAM96-111B.
DeepSeek V4 Flash 0731 REAM96-111B
This release is a calibration-guided, training-free cascade from the committed REAM128 checkpoint. It reduces every routed MoE layer from 128 to 96 experts while preserving 43 transformer layers, one shared expert, and top-k 6 routing. The logical parameter count is 111,163,180,631.
Lineage and precision
- Source:
WaveCut/DeepSeek-V4-Flash-0731-REAM128-146B@84dc02df6c844568738da6bc425ea4772a7c56f4. - Source checkpoint identity:
f96d633ad1841c18f8f3cdd960cd64c7533aee46. - Source
NATIVE_READY.jsonSHA-256:8c32c76316574c75f18b45be69bd3b54c0c547531249b02f424a66485dc4820e. - Run:
20260811T072216Z. - Build code revision:
47b3fe27862537b0a9504a135d9728fba54e0596. - Container:
sha256:0a360022e8de4375af99430f84e8b38951acc397252163a37ceac7204d01be35. - Routed expert weights retain the source-native E2M1 FP4 layout with E8M0 scales.
- Attention projections and unchanged tensors retain the source FP8/mixed layout.
- REAM accumulation is deterministic FP32; BF16 is transient compute, not the primary checkpoint format.
- MTP is disabled (
num_nextn_predict_layers: 0). - Checkpoint: 63,982,995,880 bytes across 44 Safetensors shards,
config.json, andmodel.safetensors.index.json. - Committed
NATIVE_READY.jsonSHA-256:1a0aef0c8d4f736516ec1268b08d3ad698bef16f8e0ee90739fc0a0032831527.
REAM methodology
REAM96 does not apply gradient updates or add training examples to the checkpoint. Fresh layerwise saliency statistics, expert groups, merges, router mappings, and activations were computed against the immutable REAM128 source. The REAM128 importance matrix was not reused.
Calibration composition: 3,072 samples x 512 tokens.
| Slice | Samples | Share | Pinned sources and focus |
|---|---|---|---|
| Code | 614 | 20.0% | sahil2801/CodeAlpaca-20k and bigcode/self-oss-instruct-sc2-exec-filter-50k; instruction following and executable code |
| Math | 307 | 10.0% | openai/gsm8k and AI-MO/NuminaMath-CoT; arithmetic and mathematical reasoning |
| Agent/tool | 1,075 | 35.0% | NousResearch/hermes-function-calling-v1; function selection and structured tool calls |
| Russian/Cyrillic | 1,076 | 35.0% | ZeroAgency/ru-big-russian-dataset; Russian and Cyrillic language coverage |
At each layer, REAP saliency is the routed activation norm multiplied by router probability. The 96 highest-saliency source experts become centroids; the remaining 32 source experts are assigned using output and router-gate similarity. Each group is merged by saliency-weighted deterministic FP32 accumulation, router rows are remapped, and top-k 6 routing is preserved.
Validation and integrity
- Build hardware: 2x NVIDIA GeForce RTX 4090 24 GB.
- Structural gate: 43 layers, 96 routed experts, one shared expert, top-k 6, and no MTP layers.
- Native release gate: all 46 checkpoint artifacts match the committed byte sizes and SHA-256 values.
- The native checkpoint became terminal-ready after 21 h 34 min of the run, including recoveries and durability checks.
- Comparative heldout PPL/NLL: not measured for this cascade.
- Semantic generation suite: not run for this cascade.
- Runtime benchmark: not run for this cascade.
Artifact and lineage checks prove structural integrity; they are not quality or performance evidence.
Loading and runtime compatibility
The configuration and tokenizer can be inspected from the public repository:
from transformers import AutoConfig, AutoTokenizer
repo_id = "WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B"
config = AutoConfig.from_pretrained(repo_id, trust_remote_code=False)
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=False)
print(config.num_hidden_layers, config.n_routed_experts, config.num_experts_per_tok)
Stock AutoModelForCausalLM, vLLM, and SGLang generation is not claimed for
this packed E2M1/E8M0 checkpoint. A runtime must implement the 96-expert
DeepSeek V4 topology and its native FP4/FP8 storage.
Limitations
- Quality and throughput must not be inferred from the REAM128 release; this cascade has no PPL, NLL, semantic, or benchmark result.
- This is a specialized mixed-precision checkpoint, not a conventional BF16/F16 Transformers checkpoint.
- Hardware memory requirements depend on the runtime, sharding strategy, context length, and KV-cache configuration.
License
MIT License, following the source checkpoint.
- Downloads last month
- 112
Model tree for WaveCut/DeepSeek-V4-Flash-0731-REAM96-111B-exp
Base model
deepseek-ai/DeepSeek-V4-Flash-0731