--- license: apache-2.0 language: - en - zh - multilingual tags: - gguf - qwen3 - qwen3.6 - reasoning - coding - academic-writing - uncensored - rys base_model: - Qwen/Qwen3.6-27B --- # Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL GGUF Hyper-focused Q4NL RYS release for: - programming - technical reasoning - academic-style writing This release is built from: - AEON source model: `https://huggingface.co/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored` Use this model with: - custom `ik-llama` fork, specialized and tuned for this exact model: `https://github.com/noonr48/qwen36-aeon-ik-llama` Side note (tool calling): some prompts can trigger repeated *identical* tool calls in one assistant turn (especially when the tool result is empty / slow). Update to the latest `ik-llama` fork version: it now deduplicates identical `tool_calls` server-side. ## At a glance - released Q4_NL GGUF: [`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL-ik-llama-custom-mixed.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-IQ4_NL-ik-llama-custom-mixed.gguf) - BF16 GGUF reference: [`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf) - HF-format BF16 safetensors: [`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors) - intended runtime: custom `ik-llama` - compression: `54G` BF16 -> `16G` IQ4_NL - mixed validation snapshot: `0.7299` BF16 -> `0.7244` IQ4_NL - overall performance change: `-0.0055` absolute, about `-0.75%` relative - public codename: `MaxThinkCoder` - project focus: get as much capability as possible out of a Q4-class RYS model for hard English-first work - not a generic chat release - not a stock `llama.cpp` release ## BF16 GGUF A BF16 GGUF is also included for people who want the unquantized GGUF-side artifact from the same released `15,20` RYS branch: [`Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/blob/main/Qwen3.6-27B-AEON-RYS-MaxThinkCoder-BF16.gguf) Use this if you want a GGUF reference build, local conversion/testing, or to compare quantization behavior against the released `IQ4_NL` GGUF. For normal inference, the `IQ4_NL` file is the practical target. For Transformers/LoRA/SFT workflows, use the [`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors) folder instead. Size note: - BF16 GGUF: about `54G` - IQ4_NL GGUF: about `16G` ## BF16 vs released custom IQ4_NL This is the critical compression result for the released custom `15,20` branch: - BF16 size: `54G` - released IQ4_NL size: `16G` - mixed 4-probe mean: `0.7299` BF16 -> `0.7244` IQ4_NL - net performance change: `-0.0055` absolute, about `-0.75%` relative Probe-level snapshot: | probe | BF16 | IQ4_NL | |---|---:|---:| | `math_16` | `0.8421` | `0.7897` | | `eq_16` | `0.7123` | `0.7111` | | `math_4` | `0.4851` | `0.5170` | | `gsm8k_5` | `0.8800` | `0.8800` | Practical read: - the released Q4 model is roughly `70%` smaller on disk - the mixed validation snapshot stayed under a `1%` overall drop - `eq_16` and `gsm8k_5` were effectively flat - `math_4` did not regress in this tiny probe - the real measurable hit was mostly on `math_16` ## Speed snapshot Exact comparison hardware: - `6x NVIDIA GeForce RTX 5060 Ti` | runtime | tested file | ctx | np | KV | decode tok/s | prompt tok/s | note | |---|---|---:|---:|---|---:|---:|---| | patched upstream-style `llama.cpp` | same internal standard-typed comparison file | `4096` | `1` | `f16` | `22.51` | `187.18` | internal comparison only | | custom `ik-llama` fork | released custom-mixed file | `409600` | `2` | `f32/f32` | `39.37` | `164.98` | actual deployment target | ## Why there is no `llama.cpp` file in this release We did build and benchmark an internal standard-typed comparison artifact. We are not releasing it as a public `llama.cpp` file. Why: - the main model this project is about is the custom mixed GGUF, which needs the forked `ik-llama` runtime - even the internal standard-typed path was only validated on a patched upstream-style `llama.cpp`, not clean stock mainline - since users would still need a special runtime path anyway, we did not think it was worth shipping a second public file that suggests plain stock `llama.cpp` support So the intended reading is simple: - this repo releases the `ik-llama`-targeted model - if you want plain stock `llama.cpp`, this is not that release ## Hyper-focused project This was a deliberately narrow project. The target was not “best general chat model”. The target was: - strongest Q4-class English-first model we could get for coding, reasoning, and academic work - using the AEON uncensored branch as the source - using the custom `ik-llama` path because prior RYS experiments suggested that path preserved quality better than standard `llama.cpp`-style quantization ## Imatrix calibration profile The quantization was deliberately biased toward reasoning and technical work. Heuristic calibration breakdown: - `math_reasoning`: `5,688` chunks, `1,706,070` chars (`36.0%`) - `code_technical`: `3,518` chunks, `1,343,392` chars (`28.4%`) - `experiment_docs`: `808` chunks, `224,169` chars (`4.7%`) - `writing_chat`: `387` chunks, `164,097` chars (`3.5%`) - `other`: `5,139` chunks, `1,249,396` chars (`26.4%`) Practical read: - heavy focus on reasoning math, code, technical prose, and experiment artifacts - very little emphasis on generic social chat ## RYS choice This release came from the AEON-derived `15,20` RYS branch. That was the practical release target because it quantized cleanly and held up as the best balanced candidate for this experiment. ## Use case Recommended: - coding - technical reasoning - academic-style writing - long-context English work Not recommended as a generic safe-default chat model. This branch came from an uncensored source path. ## BF16 safetensors for fine-tuning The original HF-format BF16 checkpoint for the released `15,20` RYS branch is included here: [`bf16-safetensors/`](https://huggingface.co/jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF/tree/main/bf16-safetensors) Use the files in that folder for Transformers-based work such as LoRA, SFT, continued training, or conversion into another training format. Use the GGUF file in the repo root for `ik-llama` inference. Loading example: ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "jackasda211233/Qwen3.6-27B-AEON-RYS-15-20-GGUF" subfolder = "bf16-safetensors" tok = AutoTokenizer.from_pretrained(repo_id, subfolder=subfolder, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( repo_id, subfolder=subfolder, torch_dtype="bfloat16", device_map="auto", trust_remote_code=True, ) ``` RYS note for anyone rebuilding or modifying this checkpoint: the config is part of the model. The BF16 folder keeps the corrected hybrid-stack metadata for the `15,20` insert, including `text_config.num_hidden_layers = 69` and a 69-entry `text_config.layer_types` list. Do not change the layer count without remapping `layer_types` to the same layer order as the tensors.