Instructions to use donghanasd/Agents-A1-TQ3_4S-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF # Run inference directly in the terminal: llama cli -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF # Run inference directly in the terminal: llama cli -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf donghanasd/Agents-A1-TQ3_4S-GGUF # Run inference directly in the terminal: ./llama-cli -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf donghanasd/Agents-A1-TQ3_4S-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Use Docker
docker model run hf.co/donghanasd/Agents-A1-TQ3_4S-GGUF
- LM Studio
- Jan
- Ollama
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with Ollama:
ollama run hf.co/donghanasd/Agents-A1-TQ3_4S-GGUF
- Unsloth Desktop
- Pi
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "donghanasd/Agents-A1-TQ3_4S-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with Docker Model Runner:
docker model run hf.co/donghanasd/Agents-A1-TQ3_4S-GGUF
- Lemonade
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull donghanasd/Agents-A1-TQ3_4S-GGUF
Run and chat with the model
lemonade run user.Agents-A1-TQ3_4S-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default donghanasd/Agents-A1-TQ3_4S-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use donghanasd/Agents-A1-TQ3_4S-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "donghanasd/Agents-A1-TQ3_4S-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF# Run inference directly in the terminal:
llama cli -hf donghanasd/Agents-A1-TQ3_4S-GGUFUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf donghanasd/Agents-A1-TQ3_4S-GGUF# Run inference directly in the terminal:
./llama-cli -hf donghanasd/Agents-A1-TQ3_4S-GGUFBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf donghanasd/Agents-A1-TQ3_4S-GGUF# Run inference directly in the terminal:
./build/bin/llama-cli -hf donghanasd/Agents-A1-TQ3_4S-GGUFUse Docker
docker model run hf.co/donghanasd/Agents-A1-TQ3_4S-GGUFAgents-A1 โ TQ3_4S GGUF
A TQ3_4S (~Q3) mixed-precision quantized GGUF of
InternScience/Agents-A1
(a 35B 13 GB**.
Built to run on a single 16 GB GPU (e.g. RTX 5060 Ti).qwen35moe Mixture-of-Experts agent model), weighing in at **
Quantized from the official F16 GGUF
(InternScience/Agents-A1-F16-GGUF, 69 GB)
using the same mixed K-quant + TQ3_4S tensor layout as the
YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S
recipe (both are the same qwen35moe architecture).
Files
| File | Size | Notes |
|---|---|---|
Agents-A1-TQ3_4S.gguf |
~13 GB | Mixed K-quant + TQ3_4S layout |
Quantization recipe
Base ftype Q3_K_M with per-tensor overrides (only the SSM special tensors are
TQ3_4S; everything else is a mixed K-quant layout):
token_embd = Q5_K
output = Q4_K
attn_* / *_shexp / ssm_out = Q6_K
ffn_gate_exps / ffn_up_exps = Q2_K
ffn_down_exps = Q3_K (Q4_K on blocks 21, 28, 38)
ssm_alpha / ssm_beta = TQ3_4S
Verified output layout (733 tensors):
Q2_K x80 7.05 GB (expert gate/up)
Q3_K x37 4.27 GB (expert down)
Q6_K x250 1.15 GB (attention / shared-expert / ssm_out)
Q4_K x4 0.74 GB (output + down_exps blk 21/28/38)
Q5_K x1 0.35 GB (token_embd)
TQ3_4S x60 ~0 GB (ssm_alpha / ssm_beta)
F32 x301 0.09 GB (norms / biases)
Produced with llama-quantize:
llama-quantize \
--token-embedding-type Q5_K \
--output-tensor-type Q4_K \
--tensor-type-file tensor_types.txt \
Agents-A1-F16.gguf \
Agents-A1-TQ3_4S.gguf \
Q3_K_M
No MTP: unlike the Qwen3.6-35B reference, Agents-A1 has no native MTP (
nextn) draft head, so there is no--spec-type draft-mtpspeculative decoding with this model.
Tooling
Built with the turbo-tan/llama.cpp-tq3
fork, which adds the TQ3_4S quantization type.
โ ๏ธ Version note: this file was produced with a recent fork build where the
TQ3_4Sggml type id is 46. You must run it with an up-to-date build of the same fork โ older builds (type id 45) will fail to load it.
Running with llama-server
llama-server \
--host 0.0.0.0 --port 8080 \
--model Agents-A1-TQ3_4S.gguf \
--jinja \
-ngl 99 \
-fa on \
-ctk q8_0 -ctv q8_0 \
--batch-size 2048 --ubatch-size 512 \
--ctx-size 64000 \
--parallel 1 -np 1 \
--reasoning on --reasoning-format auto \
--warmup --perf \
--threads 4 --threads-batch 8 \
--cache-ram 16000 --ctx-checkpoints 32
Notes:
--ctx-size 64000is roughly the empirical max on 16 GB; lower it on OOM.-ctk q8_0 -ctv q8_0quantizes the KV cache to fit more context.- To disable reasoning: replace
--reasoning onwith--reasoning off --reasoning-budget 0.
Sources / Attribution
- Base weights:
InternScience/Agents-A1-F16-GGUF/InternScience/Agents-A1 - Recipe reference:
YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S - Tooling:
turbo-tan/llama.cpp-tq3
License and usage follow the base model โ see
InternScience/Agents-A1.
- Downloads last month
- 50
We're not able to determine the quantization variants.
Model tree for donghanasd/Agents-A1-TQ3_4S-GGUF
Base model
InternScience/Agents-A1
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf donghanasd/Agents-A1-TQ3_4S-GGUF# Run inference directly in the terminal: llama cli -hf donghanasd/Agents-A1-TQ3_4S-GGUF