Agents-A1 โ€” TQ3_4S GGUF

A TQ3_4S (~Q3) mixed-precision quantized GGUF of InternScience/Agents-A1 (a 35B qwen35moe Mixture-of-Experts agent model), weighing in at **13 GB**. Built to run on a single 16 GB GPU (e.g. RTX 5060 Ti).

Quantized from the official F16 GGUF (InternScience/Agents-A1-F16-GGUF, 69 GB) using the same mixed K-quant + TQ3_4S tensor layout as the YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S recipe (both are the same qwen35moe architecture).

Files

File Size Notes
Agents-A1-TQ3_4S.gguf ~13 GB Mixed K-quant + TQ3_4S layout

Quantization recipe

Base ftype Q3_K_M with per-tensor overrides (only the SSM special tensors are TQ3_4S; everything else is a mixed K-quant layout):

token_embd                       = Q5_K
output                           = Q4_K
attn_* / *_shexp / ssm_out       = Q6_K
ffn_gate_exps / ffn_up_exps      = Q2_K
ffn_down_exps                    = Q3_K   (Q4_K on blocks 21, 28, 38)
ssm_alpha / ssm_beta             = TQ3_4S

Verified output layout (733 tensors):

Q2_K    x80    7.05 GB   (expert gate/up)
Q3_K    x37    4.27 GB   (expert down)
Q6_K   x250    1.15 GB   (attention / shared-expert / ssm_out)
Q4_K     x4    0.74 GB   (output + down_exps blk 21/28/38)
Q5_K     x1    0.35 GB   (token_embd)
TQ3_4S  x60   ~0   GB   (ssm_alpha / ssm_beta)
F32    x301    0.09 GB   (norms / biases)

Produced with llama-quantize:

llama-quantize \
  --token-embedding-type Q5_K \
  --output-tensor-type Q4_K \
  --tensor-type-file tensor_types.txt \
  Agents-A1-F16.gguf \
  Agents-A1-TQ3_4S.gguf \
  Q3_K_M

No MTP: unlike the Qwen3.6-35B reference, Agents-A1 has no native MTP (nextn) draft head, so there is no --spec-type draft-mtp speculative decoding with this model.

Tooling

Built with the turbo-tan/llama.cpp-tq3 fork, which adds the TQ3_4S quantization type.

โš ๏ธ Version note: this file was produced with a recent fork build where the TQ3_4S ggml type id is 46. You must run it with an up-to-date build of the same fork โ€” older builds (type id 45) will fail to load it.

Running with llama-server

llama-server \
  --host 0.0.0.0 --port 8080 \
  --model Agents-A1-TQ3_4S.gguf \
  --jinja \
  -ngl 99 \
  -fa on \
  -ctk q8_0 -ctv q8_0 \
  --batch-size 2048 --ubatch-size 512 \
  --ctx-size 64000 \
  --parallel 1 -np 1 \
  --reasoning on --reasoning-format auto \
  --warmup --perf \
  --threads 4 --threads-batch 8 \
  --cache-ram 16000 --ctx-checkpoints 32

Notes:

  • --ctx-size 64000 is roughly the empirical max on 16 GB; lower it on OOM.
  • -ctk q8_0 -ctv q8_0 quantizes the KV cache to fit more context.
  • To disable reasoning: replace --reasoning on with --reasoning off --reasoning-budget 0.

Sources / Attribution

License and usage follow the base model โ€” see InternScience/Agents-A1.

Downloads last month
50
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for donghanasd/Agents-A1-TQ3_4S-GGUF

Quantized
(1)
this model