Safetensors
GGUF
English
Chinese
multilingual
qwen3
qwen3.6
reasoning
coding
academic-writing
uncensored
rys
mtp
ik-llama
conversational
Instructions to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Use Docker
docker model run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with Ollama:
ollama run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
- Unsloth Desktop
- Pi
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with Docker Model Runner:
docker model run hf.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
- Lemonade
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Run and chat with the model
lemonade run user.Qwen3.6-27B-AEON-RYS-15-20-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 7,409 Bytes
21159c4 307ca95 21159c4 3d62660 307ca95 3d62660 58293d9 3d62660 204bd92 58293d9 130dd95 3d62660 e9027a0 c5fde11 3d62660 eacd0f9 bcceae3 eacd0f9 3d62660 5ce6af3 c5fde11 5ce6af3 c5fde11 5ce6af3 eacd0f9 bcceae3 eacd0f9 3d62660 e9027a0 3d62660 21159c4 3d62660 21159c4 3d62660 ecef92b 3d62660 ecef92b 3d62660 21159c4 3d62660 21159c4 3d62660 21159c4 3d62660 21159c4 966c506 eacd0f9 966c506 eacd0f9 966c506 3d62660 21159c4 3d62660 21159c4 3d62660 21159c4 3d62660 21159c4 3d62660 e30fded c5fde11 e30fded c5fde11 e30fded | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 | ---
license: apache-2.0
language:
- en
- zh
- multilingual
tags:
- gguf
- qwen3
- qwen3.6
- reasoning
- coding
- academic-writing
- uncensored
- rys
base_model:
- Qwen/Qwen3.6-27B
---
# Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL GGUF
Hyper-focused Q4NL RYS release for:
- programming
- technical reasoning
- academic-style writing
This release is built from:
- AEON source model:
`https://huggingface.co/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored`
Use this model with:
- custom `ik-llama` fork, specialized and tuned for this exact model:
`https://github.com/noonr48/qwen36-aeon-ik-llama`
Side note (tool calling):
some prompts can trigger repeated *identical* tool calls in one assistant turn (especially when the tool result is empty / slow).
Update to the latest `ik-llama` fork version: it now deduplicates identical `tool_calls` server-side.
## At a glance
- released Q4_NL GGUF:
[`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL-ik-llama-custom-mixed.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL-ik-llama-custom-mixed.gguf)
- BF16 GGUF reference:
[`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf)
- HF-format BF16 safetensors:
[`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors)
- intended runtime:
custom `ik-llama`
- compression:
`54G` BF16 -> `16G` IQ4_NL
- mixed validation snapshot:
`0.7299` BF16 -> `0.7244` IQ4_NL
- overall performance change:
`-0.0055` absolute, about `-0.75%` relative
- public codename:
`MaxThinkCoder`
- project focus:
get as much capability as possible out of a Q4-class RYS model for hard English-first work
- not a generic chat release
- not a stock `llama.cpp` release
## BF16 GGUF
A BF16 GGUF is also included for people who want the unquantized GGUF-side artifact from the same released `15,20` RYS branch:
[`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf)
Use this if you want a GGUF reference build, local conversion/testing, or to compare quantization behavior against the released `IQ4_NL` GGUF. For normal inference, the `IQ4_NL` file is the practical target. For Transformers/LoRA/SFT workflows, use the [`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors) folder instead.
Size note:
- BF16 GGUF: about `54G`
- IQ4_NL GGUF: about `16G`
## BF16 vs released custom IQ4_NL
This is the critical compression result for the released custom `15,20` branch:
- BF16 size:
`54G`
- released IQ4_NL size:
`16G`
- mixed 4-probe mean:
`0.7299` BF16 -> `0.7244` IQ4_NL
- net performance change:
`-0.0055` absolute, about `-0.75%` relative
Probe-level snapshot:
| probe | BF16 | IQ4_NL |
|---|---:|---:|
| `math_16` | `0.8421` | `0.7897` |
| `eq_16` | `0.7123` | `0.7111` |
| `math_4` | `0.4851` | `0.5170` |
| `gsm8k_5` | `0.8800` | `0.8800` |
Practical read:
- the released Q4 model is roughly `70%` smaller on disk
- the mixed validation snapshot stayed under a `1%` overall drop
- `eq_16` and `gsm8k_5` were effectively flat
- `math_4` did not regress in this tiny probe
- the real measurable hit was mostly on `math_16`
## Speed snapshot
Exact comparison hardware:
- `6x NVIDIA GeForce RTX 5060 Ti`
| runtime | tested file | ctx | np | KV | decode tok/s | prompt tok/s | note |
|---|---|---:|---:|---|---:|---:|---|
| patched upstream-style `llama.cpp` | same internal standard-typed comparison file | `4096` | `1` | `f16` | `22.51` | `187.18` | internal comparison only |
| custom `ik-llama` fork | released custom-mixed file | `409600` | `2` | `f32/f32` | `39.37` | `164.98` | actual deployment target |
## Why there is no `llama.cpp` file in this release
We did build and benchmark an internal standard-typed comparison artifact.
We are not releasing it as a public `llama.cpp` file.
Why:
- the main model this project is about is the custom mixed GGUF, which needs the forked `ik-llama` runtime
- even the internal standard-typed path was only validated on a patched upstream-style `llama.cpp`, not clean stock mainline
- since users would still need a special runtime path anyway, we did not think it was worth shipping a second public file that suggests plain stock `llama.cpp` support
So the intended reading is simple:
- this repo releases the `ik-llama`-targeted model
- if you want plain stock `llama.cpp`, this is not that release
## Hyper-focused project
This was a deliberately narrow project.
The target was not “best general chat model”.
The target was:
- strongest Q4-class English-first model we could get for coding, reasoning, and academic work
- using the AEON uncensored branch as the source
- using the custom `ik-llama` path because prior RYS experiments suggested that path preserved quality better than standard `llama.cpp`-style quantization
## Imatrix calibration profile
The quantization was deliberately biased toward reasoning and technical work.
Heuristic calibration breakdown:
- `math_reasoning`: `5,688` chunks, `1,706,070` chars (`36.0%`)
- `code_technical`: `3,518` chunks, `1,343,392` chars (`28.4%`)
- `experiment_docs`: `808` chunks, `224,169` chars (`4.7%`)
- `writing_chat`: `387` chunks, `164,097` chars (`3.5%`)
- `other`: `5,139` chunks, `1,249,396` chars (`26.4%`)
Practical read:
- heavy focus on reasoning math, code, technical prose, and experiment artifacts
- very little emphasis on generic social chat
## RYS choice
This release came from the AEON-derived `15,20` RYS branch.
That was the practical release target because it quantized cleanly and held up as the best balanced candidate for this experiment.
## Use case
Recommended:
- coding
- technical reasoning
- academic-style writing
- long-context English work
Not recommended as a generic safe-default chat model.
This branch came from an uncensored source path.
## BF16 safetensors for fine-tuning
The original HF-format BF16 checkpoint for the released `15,20` RYS branch is included here:
[`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors)
Use the files in that folder for Transformers-based work such as LoRA, SFT, continued training, or conversion into another training format. Use the GGUF file in the repo root for `ik-llama` inference.
Loading example:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF"
subfolder = "bf16-safetensors"
tok = AutoTokenizer.from_pretrained(repo_id, subfolder=subfolder, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
subfolder=subfolder,
torch_dtype="bfloat16",
device_map="auto",
trust_remote_code=True,
)
```
RYS note for anyone rebuilding or modifying this checkpoint: the config is part of the model. The BF16 folder keeps the corrected hybrid-stack metadata for the `15,20` insert, including `text_config.num_hidden_layers = 69` and a 69-entry `text_config.layer_types` list. Do not change the layer count without remapping `layer_types` to the same layer order as the tensors.
|