How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf skt/A.X-K2-GGUF:IQ4_XS
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default skt/A.X-K2-GGUF:IQ4_XS
Run Hermes
hermes
Quick Links

A.X K2 GGUF

A.X Logo

๐Ÿค— Original model | ๐Ÿค— Collection | ๐Ÿ–ฅ๏ธ Github | ๐Ÿ“„ Technical Report

This repository contains a quantized GGUF build of skt/A.X-K2.

A.X K2 is a Mixture-of-Experts model with 688B parameters, 33B of them active per token, released in block-scaled FP8. The file here is converted from that checkpoint and quantized to the format in Model files.

The original model card covers the architecture, training data, benchmarks, intended use and limitations. This card only adds what is specific to the GGUF build.

Model files

File Size Bits-per-weight
A.X-K2-IQ4_XS.gguf 345 GiB 4.30
hf download skt/A.X-K2-GGUF --include "A.X-K2-IQ4_XS.gguf" --local-dir .
MODEL=$PWD/A.X-K2-IQ4_XS.gguf     # absolute, so it survives cd into a source tree

Tensors left at full precision include output.weight, token_embd.weight, the Gated Norm projections (*norm_gate_a/b.weight), the sparse-attention indexer projection (*indexer.proj.weight), and the MoE router (*ffn_gate_inp.weight).

Run with llama.cpp

Official llama.cpp does not support A.X K2 yet, so build from the A.X-K2 fork. It is upstream b10236 plus A.X K2 support, and nothing else:

git clone -b axk2-b10236 https://github.com/cys4/llama.cpp.git
cd llama.cpp

cmake -B build -DGGML_CUDA=ON     # CUDA; omit -DGGML_CUDA=ON for a CPU-only build
cmake --build build -j

For anything the examples do not cover, see the llama.cpp documentation - it all applies here, as long as you build the fork above.

CLI

./build/bin/llama-cli -m "$MODEL" --temp 0.6 --top-p 0.95 -st -p "๋Œ€ํ•œ๋ฏผ๊ตญ์˜ ์ˆ˜๋„๋Š”?" \
  --reasoning on     # thinking mode; --reasoning off for non-thinking

--temp and --top-p control the sampling: lower temperature is more deterministic, and top-p caps the cumulative probability of the token pool.

-st runs a single turn: llama-cli answers the prompt and exits. Without it, the CLI stays open for interactive chat.

Server

./build/bin/llama-server -m "$MODEL" --temp 0.6 --top-p 0.95 \
  --host 0.0.0.0 --port 8080 \
  --reasoning on     # thinking mode; --reasoning off for non-thinking

The server exposes an OpenAI-compatible API at http://localhost:8080/v1. --host 0.0.0.0 lets other machines connect (the default is 127.0.0.1 only), and --port picks the port (8080 is already the default).

curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
  "messages": [{"role": "user", "content": "๋Œ€ํ•œ๋ฏผ๊ตญ์˜ ์ˆ˜๋„๋Š”?"}]
}'

Run with vLLM

Note: GGUF support in vLLM is experimental and under-optimized upstream, positioned mainly as a way to reduce memory footprint. For best GGUF performance use llama.cpp above.

The same GGUF file can also be served with vLLM. Like llama.cpp above, this needs a custom build: official vLLM does not load A.X K2 GGUFs, so install vLLM from the A.X-K2 vLLM fork, branch axk2-v0.23.0_gguf, which carries the A.X K2 GGUF loading support:

git clone -b axk2-v0.23.0_gguf https://github.com/cys4/vllm_axk2.git
cd vllm_axk2
VLLM_USE_PRECOMPILED=1 pip install -e .

The fork is python-only on top of upstream vLLM, so VLLM_USE_PRECOMPILED=1 reuses the matching precompiled wheel and no CUDA build is needed.

The vllm_hf_config/ folder in this repository carries the config and tokenizer vLLM needs: the original config.json with quantization_config removed (the FP8 declaration would conflict with GGUF loading) plus the unmodified tokenizer and chat template.

hf download skt/A.X-K2-GGUF --include "vllm_hf_config/*" --local-dir "$(dirname "$MODEL")"
CONFIG="$(dirname "$MODEL")/vllm_hf_config"

vllm serve "$MODEL" \
  --hf-config-path "$CONFIG" --tokenizer "$CONFIG" -tp 8 \
  --host 0.0.0.0 --port 8000 \
  --default-chat-template-kwargs '{"enable_thinking": true}'     # thinking mode; false for non-thinking

The server exposes the same OpenAI-compatible API at http://localhost:8000/v1, so query it like llama-server above, with the port changed. vllm_hf_config/ also ships generation_config.json, which carries the sampling defaults (temperature, top_p).

Contact

For questions about A.X K2 โ€” including model behavior, deployment, and licensing โ€” contact the A.X team at a.x@sk.com. Please send reports of vulnerabilities, harmful outputs, suspected misuse, or copyright infringement claims to the same address.

Citation

If you use A.X K2 in your research, please cite the technical report:

@techreport{axk2-2026,
      title={A.X K2 Technical Report},
      author={SK Telecom},
      year={2026},
      institution={SK Telecom},
      url={https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf},
}
Downloads last month
113
GGUF
Model size
690B params
Architecture
axk2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for skt/A.X-K2-GGUF

Base model

skt/A.X-K2
Quantized
(2)
this model

Collection including skt/A.X-K2-GGUF