--- library_name: gguf base_model: InternScience/Agents-A1-F16-GGUF tags: - gguf - qwen3 - qwen35moe - moe - mixture-of-experts - agent - quantized - llama.cpp - tq3_4s --- # Agents-A1 — TQ3_4S GGUF A **TQ3_4S (~Q3) mixed-precision quantized GGUF** of [`InternScience/Agents-A1`](https://huggingface.co/InternScience/Agents-A1) (a ~35B `qwen35moe` Mixture-of-Experts agent model), weighing in at **~13 GB**. Built to run on a single **16 GB GPU** (e.g. RTX 5060 Ti). Quantized from the official F16 GGUF ([`InternScience/Agents-A1-F16-GGUF`](https://huggingface.co/InternScience/Agents-A1-F16-GGUF), 69 GB) using the same mixed K-quant + `TQ3_4S` tensor layout as the [`YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S`](https://huggingface.co/YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S) recipe (both are the same `qwen35moe` architecture). ## Files | File | Size | Notes | |------|------|-------| | `Agents-A1-TQ3_4S.gguf` | ~13 GB | Mixed K-quant + TQ3_4S layout | ## Quantization recipe Base ftype `Q3_K_M` with per-tensor overrides (only the SSM special tensors are `TQ3_4S`; everything else is a mixed K-quant layout): ``` token_embd = Q5_K output = Q4_K attn_* / *_shexp / ssm_out = Q6_K ffn_gate_exps / ffn_up_exps = Q2_K ffn_down_exps = Q3_K (Q4_K on blocks 21, 28, 38) ssm_alpha / ssm_beta = TQ3_4S ``` Verified output layout (733 tensors): ``` Q2_K x80 7.05 GB (expert gate/up) Q3_K x37 4.27 GB (expert down) Q6_K x250 1.15 GB (attention / shared-expert / ssm_out) Q4_K x4 0.74 GB (output + down_exps blk 21/28/38) Q5_K x1 0.35 GB (token_embd) TQ3_4S x60 ~0 GB (ssm_alpha / ssm_beta) F32 x301 0.09 GB (norms / biases) ``` Produced with `llama-quantize`: ```bash llama-quantize \ --token-embedding-type Q5_K \ --output-tensor-type Q4_K \ --tensor-type-file tensor_types.txt \ Agents-A1-F16.gguf \ Agents-A1-TQ3_4S.gguf \ Q3_K_M ``` > **No MTP:** unlike the Qwen3.6-35B reference, Agents-A1 has **no native MTP > (`nextn`) draft head**, so there is no `--spec-type draft-mtp` speculative > decoding with this model. ## Tooling Built with the [`turbo-tan/llama.cpp-tq3`](https://github.com/turbo-tan/llama.cpp-tq3) fork, which adds the `TQ3_4S` quantization type. > ⚠️ **Version note:** this file was produced with a recent fork build where the > `TQ3_4S` ggml type id is **46**. You must run it with an **up-to-date build of > the same fork** — older builds (type id 45) will fail to load it. ## Running with `llama-server` ```bash llama-server \ --host 0.0.0.0 --port 8080 \ --model Agents-A1-TQ3_4S.gguf \ --jinja \ -ngl 99 \ -fa on \ -ctk q8_0 -ctv q8_0 \ --batch-size 2048 --ubatch-size 512 \ --ctx-size 64000 \ --parallel 1 -np 1 \ --reasoning on --reasoning-format auto \ --warmup --perf \ --threads 4 --threads-batch 8 \ --cache-ram 16000 --ctx-checkpoints 32 ``` Notes: - `--ctx-size 64000` is roughly the empirical max on 16 GB; lower it on OOM. - `-ctk q8_0 -ctv q8_0` quantizes the KV cache to fit more context. - To disable reasoning: replace `--reasoning on` with `--reasoning off --reasoning-budget 0`. ## Sources / Attribution - **Base weights:** [`InternScience/Agents-A1-F16-GGUF`](https://huggingface.co/InternScience/Agents-A1-F16-GGUF) / [`InternScience/Agents-A1`](https://huggingface.co/InternScience/Agents-A1) - **Recipe reference:** [`YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S`](https://huggingface.co/YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S) - **Tooling:** [`turbo-tan/llama.cpp-tq3`](https://github.com/turbo-tan/llama.cpp-tq3) License and usage follow the base model — see [`InternScience/Agents-A1`](https://huggingface.co/InternScience/Agents-A1).