Instructions to use Fileportz/DeepSeek-V4.1-Flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Fileportz/DeepSeek-V4.1-Flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Fileportz/DeepSeek-V4.1-Flash")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Fileportz/DeepSeek-V4.1-Flash", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Fileportz/DeepSeek-V4.1-Flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Fileportz/DeepSeek-V4.1-Flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Fileportz/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Fileportz/DeepSeek-V4.1-Flash
- SGLang
How to use Fileportz/DeepSeek-V4.1-Flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Fileportz/DeepSeek-V4.1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Fileportz/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Fileportz/DeepSeek-V4.1-Flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Fileportz/DeepSeek-V4.1-Flash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Fileportz/DeepSeek-V4.1-Flash with Docker Model Runner:
docker model run hf.co/Fileportz/DeepSeek-V4.1-Flash
File size: 1,759 Bytes
a4fe5ed | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 | #!/usr/bin/env bash
#
# Run the reference inference on a converted checkpoint.
#
# ./run.sh /path/to/DeepSeek-V4.1-Exp-TP8
# ./run.sh /path/to/DeepSeek-V4.1-Exp-TP8 examples/example_harmony.json
# MP=4 ./run.sh /path/to/DeepSeek-V4.1-Exp-TP4
#
# Paths inside an example are resolved from this directory, so run it from anywhere.
set -euo pipefail
cd "$(dirname "$0")"
CKPT_PATH="${1:-${CKPT_PATH:-}}"
INPUT_FILE="${2:-${INPUT_FILE:-examples/example_harmony.json}}"
MP="${MP:-8}"
CONFIG="${CONFIG:-config.json}"
usage() {
echo "usage: $0 <checkpoint-dir> [input-file]" >&2
echo >&2
echo " checkpoint-dir holds model0-mp${MP}.safetensors .. model$((MP - 1))-mp${MP}.safetensors," >&2
echo " as produced by convert.py --model-parallel ${MP}" >&2
echo " input-file TXT or JSON prompts (default: examples/example.txt)" >&2
echo >&2
echo " MP=${MP} CONFIG=${CONFIG} override with environment variables" >&2
exit 1
}
[ -n "${CKPT_PATH}" ] || usage
if [ ! -d "${CKPT_PATH}" ]; then
echo "error: checkpoint directory not found: ${CKPT_PATH}" >&2
usage
fi
missing=0
for rank in $(seq 0 $((MP - 1))); do
if [ ! -f "${CKPT_PATH}/model${rank}-mp${MP}.safetensors" ]; then
missing=$((missing + 1))
fi
done
if [ "${missing}" -ne 0 ]; then
echo "error: ${CKPT_PATH} is missing ${missing} of the ${MP} shards MP=${MP} needs" >&2
echo " expected model0-mp${MP}.safetensors .. model$((MP - 1))-mp${MP}.safetensors" >&2
usage
fi
[ -f "${INPUT_FILE}" ] || { echo "error: input file not found: ${INPUT_FILE}" >&2; usage; }
torchrun --nproc-per-node "${MP}" generate.py \
--ckpt-path "${CKPT_PATH}" \
--config "${CONFIG}" \
--input-file "${INPUT_FILE}"
|