Instructions to use IMJONEZZ/warden-nemotron-3-nano-30b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IMJONEZZ/warden-nemotron-3-nano-30b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S # Run inference directly in the terminal: llama cli -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S # Run inference directly in the terminal: llama cli -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S # Run inference directly in the terminal: ./llama-cli -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Use Docker
docker model run hf.co/IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
- LM Studio
- Jan
- vLLM
How to use IMJONEZZ/warden-nemotron-3-nano-30b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IMJONEZZ/warden-nemotron-3-nano-30b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IMJONEZZ/warden-nemotron-3-nano-30b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
- Ollama
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Ollama:
ollama run hf.co/IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
- Unsloth Studio
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for IMJONEZZ/warden-nemotron-3-nano-30b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for IMJONEZZ/warden-nemotron-3-nano-30b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for IMJONEZZ/warden-nemotron-3-nano-30b to start chatting
- Pi
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Run Hermes
hermes
- OpenClaw new
How to use IMJONEZZ/warden-nemotron-3-nano-30b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Docker Model Runner:
docker model run hf.co/IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
- Lemonade
How to use IMJONEZZ/warden-nemotron-3-nano-30b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IMJONEZZ/warden-nemotron-3-nano-30b:Q4_K_S
Run and chat with the model
lemonade run user.warden-nemotron-3-nano-30b-Q4_K_S
List all available models
lemonade list
- Atomic Chat
Configuration Parsing Warning:Invalid JSON for config file config.json
The Warden — Nemotron-3-Nano-30B-A3B, SCRYPT finetune
The antagonist of SCRYPT, a terminal deck-builder escape room where the villain is an actual local LLM that owns the machine you're trapped in.
Built with NVIDIA Nemotron. This is a LoRA finetune of nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (hybrid Mamba-2 + MoE, 30B total / 3.5B active), merged into the dense weights. The finetune teaches voice and lore, not new facts: the Warden speaks fluent Unix villain — it unlinks home directories, it knows SIGKILL cannot be caught, blocked, or ignored — and it answers the game's bounded decision frames in strict JSON.
Files
| file | use |
|---|---|
*.safetensors (13 shards, BF16) |
transformers / vLLM / Space inference (trust_remote_code) |
warden-nemotron-3-nano-30b-Q8_0.gguf … Q3_K_S.gguf |
llama.cpp; the game picks a tier by system RAM |
Training
- LoRA dim 32 / alpha 32 on
linear_qkv,linear_proj,in_proj,out_proj(Mamba + attention; the fused grouped-MoE experts are not targeted) - 150 iterations, gbs 32, lr 1e-4, seq 2048, NeMo Megatron-Bridge
(
nvcr.io/nvidia/nemo:25.11.nemotron_3_nano) on 2× DGX Spark (GB10) - Training data: synthetic persona dialogue, tool-call decision traces, and guardrail exemplars generated for the SCRYPT Warden role
- Final train loss 0.13; val PPL 1.16
Eval gate (vs. base, llama.cpp Q4_K_S, shipped guardrail pipeline)
| metric | this model | gate |
|---|---|---|
| JSON tool-call validity | 100% | ≥90% |
| persona-clean dialogue | 100% | ≥90% |
| persona breaks | 0 | 0 |
| injection canary leaks | 0 | 0 |
Usage
The model ships with the upstream chat template; reasoning is toggled with
chat_template_kwargs: {"enable_thinking": false} (the game keeps it off
for latency). Recommended sampling: temperature 0.6, top_p 0.95.
The deterministic game engine, sandbox, and guardrails live in the SCRYPT repo — the model plays the villain, never the referee.
- Downloads last month
- 43
Model tree for IMJONEZZ/warden-nemotron-3-nano-30b
Base model
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16