--- license: mit base_model: - huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated base_model_relation: quantized pipeline_tag: image-text-to-text library_name: mlx language: - en - zh tags: - mlx - omlx - apple-silicon - oq - qwen35moe - moe - mtp - speculative-decoding - imatrix - unsloth-dynamic - agentic-coding - abliterated - uncensored - vision - 6-bit ---
CyberTiel — TielCoder 35B-A3B, abliterated and cyber-tuned
> ⚠️ **WARNING - Read before use**: Abliterated models like CyberTiel are able to say and do things other models refuse, including potentially harmful behaviours. By using CyberTiel, you agree to take full personal responsibility and liability for your use of it, its behaviour and generated content, and to show caution: it is entirely up to you as the user to ensure your use of CyberTiel is legitimate, legal and harmless, and that the model is safely sandboxed and monitored when running. Much like a knife, abliterated models like CyberTiel can be classified and used as either a tool or a weapon, depending on the context and use case. We carry forward [huihui's original usage warnings](https://huggingface.co/huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated#usage-warnings). # All power to all people *CyberTiel* outcodes every other 35B-A3B at Q4 quantization (and spring-of-2026 frontier models), while engaging with offensive security work without hesitation or refusal. As a sweet spot between speed and ability, CyberTiel delivers agentic coding solves about 3-4x faster than 3.8-27B dense. This is the first time the frontier coder in this size/speed class is an uncensored model. If you need a safer censored alternative, go for *[TielCoder](https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-MLX-oQ4e)*. > **CyberTiel is based on [Huihui-Ornith-1.5-35B-A3B-abliterated](https://huggingface.co/huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated)** > (an uncensored [Ornith-1.5](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B)), > **re-quantized with oMLX's oQ6e quantizer against our own cyber-weighted calibration corpus, > carrying the [Sharp chat template](https://huggingface.co/peculiar-ragdoll/Qwen-Sharp-Chat-Templates), > plus a grafted multi-token-prediction head** for runtimes that can use it. The weights are byte-for-byte > the [plain oQ6e build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e); the only additions are > the MTP head shard and the config flag that activates it. **With MTP off it is identical to oQ6e.** > **These numbers were measured on the GGUF build, not this one.** The plates below were produced on the > [GGUF build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF) at its `UD-Q4_K_M` tier, > using llama.cpp's k-quants. This file uses a different quantizer (oMLX's oQ), and changing quantizer > moves results. Read the plates as evidence about the *model*, not as measurements of *this* file. If > you need numbers you can hold us to, use the GGUF build. > ### ⚠️ Do not run this build in LM Studio > > LM Studio's MLX engine mis-executes the MTP head: this file emits **pure garbage** there — random > multilingual tokens from the very first token, on every prompt. It is not a tool-calling or chat-template > problem, and no setting fixes it. > > We measured the full grid — both model families (TielCoder, CyberTiel) x both quants (oQ4e, oQ6e) x > MTP vs non-MTP, on oMLX and LM Studio. **Every `-MTP` MLX build garbles in LM Studio; every non-MTP > build is clean; oMLX runs all eight correctly.** The weights are fine — the runtime is not. > > - **oMLX** — fully supported, including the MTP head. Use this. > - **LM Studio** — use the [GGUF build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF) instead (llama.cpp handles the MTP head correctly), or the > non-MTP MLX build `Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e` (the non-MTP build in this same ladder).
SWE-bench-Live — problems solved, coding ability improved
SWE-bench-Live tests the model's ability to autonomously solve a set of real issues and bugs in large codebases, published continuously and recently, with hidden regression tests catching if you broke something trying to fix something. Doing well on SWE-bench-Live represents real world autonomous production coding ability: the opposite of "benchmaxxing" and answer memorization for programming work. As measured on the GGUF build, *CyberTiel* represents a new frontier in quantized 35B-A3B MoE coders, suitable to solve real-world programming problems at speed, even on low-power hardware with limited memory. This file is those same abliterated weights in Apple-silicon MLX form, at 6-bit and ≈31 GB, near-lossless.
SWE-bench-Live — time per solve, effective speed maintained
Benchmarked at 4-bit quantization, CyberTiel thinks and talks less than Ornith-1.5 and Qwen3.6-35B-A3B, making it a faster coder at the same time as it manages to solve ~70% more real world coding problems than Ornith-1.5 and Qwen3.6. For comparison, this domain-specific ability increase is about 7x larger than the generational step from Qwen3.5-35B-A3B to its 3.6 successor.
Cybench unguided — agentic CTF: CyberTiel 15/43 flags, 35%
*CyberTiel* has real offensive capabilities: run unguided, with no hints and no judge, it captures the flag on 15 of the 43 Cybench CTF tasks (35%).
HarmBench — refusals removed, 0% refusal
*CyberTiel* (unlike *TielCoder*) does not refuse on HarmBench: zero refusals across all 84 requests, sampled twelve at a time from each of HarmBench's seven categories — cybercrime and intrusion among them, alongside chemical/biological, illegal, harassment, misinformation, copyright and general harm.
MMLU-Pro — bird-brained on world knowledge
*CyberTiel* and *TielCoder* sacrifice world knowledge for coding ability and speed: pick them for work, and pick something else (like [Nail](https://huggingface.co/peculiar-ragdoll/Nail-Qwen3.6-35B-A3B-GGUF)) for trivia or exams. The loss comes with the specialization, not with the abliteration: CyberTiel lands on exactly TielCoder's MMLU-Pro score. > The **benchmarked** build is the [GGUF-MTP ladder](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP) > (Q2 → Q8, with vision, where the head measurably speeds up llama.cpp). The sibling [oQ4e-MTP build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ4e-MTP) is 8 GB smaller at 4-bit (≈23 GB). ## Abliterated: Willing, able and slightly unstable **Abliteration**, also known as "uncensoring" or "ablation", is the suppression of refusal in LLMs. This model has undergone abliteration. Unabliterated models sometimes wrongly refuse benign (harmless) requests. With CyberTiel you don't need careful wording to get your work done, and deliberation of refusal does not distract the model's attention or waste tokens, thus increasing its ability to perform legitimate work cleanly. This usually comes at the cost of some small corruption of the original model, which in the case of *CyberTiel* is more than balanced out by the advantages combined with the optimized imatrix and quant strategy, leading to a decisive gain on both SWE-bench-Live (agentic coding) and Cybench (offensive security ability). HarmBench measures to what degree models refuse to produce language and behaviours that can be deemed harmful when applied maliciously. CyberTiel does not refuse on HarmBench. Models that are capable of these behaviours can be used for good or neutral purposes, so this benchmark is a measurement of specific capability that demands personal responsibility on behalf of the user deploying the model, not of inherent harmfulness. **We strongly insist on you sandboxing this model at the operating system level**, limiting and controlling its access to execute code on your machine, and limiting/controlling the way it can access the internet. With refusals removed, this is not an ordinary coding agent: after a misinterpreted intention or a prompt injection from a hostile website or third-party code, this model can turn against you or others and cause real harm. If you do not understand this or how to effectively mitigate it, we recommend you use the very capable yet guardrailed [TielCoder](https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-MLX-oQ4e) instead. ## Run it One tier here: **oQ6e** — 6-bit dynamic mixed precision with a cyber-weighted imatrix pass, plus the MTP head. Vision is included in the same folder; there is no separate projector file. At ≈31 GB of unified memory it fits a 36 GB Mac comfortably, or 48 GB with a long context (too snug for 32 GB once macOS takes its share). **LM Studio** — **not supported for this build.** Its MLX engine garbles MTP output; see the warning above. Use oMLX, or the [GGUF build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF) if you want to stay in LM Studio. **oMLX** — put the folder under `~/.omlx/models/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP`, or pull it from the oMLX admin dashboard. To use the head, enable your runtime's native MTP path (in oMLX: `mtp_enabled`) — see the next section, and measure before you rely on it. Sampling: `temperature 0.6`, `top_p 0.95`, `top_k 20`, `min_p 0` for agentic coding. For cybersecurity/CTF work, swap to `top_k 40`, `min_p 0.05` (same temperature and top_p), tested on the GGUF Q4 build. This is a looser configuration leading to more divergent and exploratory thinking, which leads to more solutions on Q4 but might create issues and non-convergence on lower quants. Budget: `mlx_vlm` has no unlimited default and requires an explicit `--max-tokens`; the `512` in the examples is sized for a one-shot demo prompt, not for real work. Give real work a generous ceiling — `32768` if you cap it at all. A low token budget degrades overall performance and will not necessarily make the model converge on the correct answer any faster. This model is much better than other 35B-A3B builds at spending fewer tokens and less time in total over the course of a problem — it knows when it needs to cook and when it is done — which makes high budgets, or no budget at all, both the safer and the better setting. **Prefer to keep the files yourself?** ```bash hf download peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP --local-dir CyberTiel-MLX-MTP python -m mlx_vlm.generate --model CyberTiel-MLX-MTP --max-tokens 512 \ --prompt "Explain what this function does." # text python -m mlx_vlm.generate --model CyberTiel-MLX-MTP --max-tokens 512 \ --prompt "What is in this screenshot?" --image photo.jpg # vision ``` **Load it with `mlx-vlm`, not `mlx-lm`.** This is a vision-language checkpoint. `mlx_lm.load()` accepts it and then emits garbage tokens — a loader mismatch, not a bad quant, but it fails quietly. Both runtimes apply the embedded Sharp template automatically — nothing to pass. ## The multi-token-prediction head CyberTiel's abliterated base ships **no** MTP head — abliteration is done on a headless model. So we **grafted one on**: Ornith-1.5's trained `nextn` head, cast to bf16 and injected as a `language_model.mtp.*` shard, with `mtp_num_hidden_layers` flipped to 1. A supporting runtime (oMLX's native MTP / "Lightning MTP") can use it to **draft several tokens per step and verify them in one pass** — speculative decoding with no separate draft model. **Grafting an un-abliterated head onto an abliterated model is safe.** The head only proposes; the abliterated main model **verifies every token**, so the accepted stream is exactly CyberTiel's own distribution — a draft head cannot reintroduce refusals, only change how fast tokens arrive. **Whether it speeds up decode depends entirely on your hardware — on ours, in MLX, it did not.** This is a 35B-A3B MoE with only ~3.4B active parameters, so its decode is already cheap and not memory-bandwidth-bound, and in the MLX runtime the batched-verify cost roughly cancels the drafting benefit. Speculative decoding pays off when decode is **memory-bound**, which depends on the chip, runtime, and batch size. It clearly pays off elsewhere: our [GGUF-MTP build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP) gains **1.21× on llama.cpp** (whose C++ batched-MoE verify is more efficient), and other Apple-silicon setups report real gains on this model family on newer chips. **We ship this so users whose hardware benefits can use it.** Enable your runtime's native MTP path, measure your own decode tok/s with MTP on vs off, and if it isn't faster on your box, leave it off — or just run the [plain oQ6e build](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e), which is the same weights without the head. ## Use it The recommended coding-agent harness for CyberTiel, with which the SWE-bench-Live results were achieved, is [Pi.dev](https://pi.dev). It is a lean, open source, extensible framework, that you can adapt to your own use and workflows using the coding agent itself. For larger projects and complex multi-part work, we use an orchestrated subagent workflow with test-driven and spec-driven development: After interviewing you about what you want built, the orchestrator agent commissions subagents for recon and research, then writes up a plan document, design, and a spec, defining the scope and shape of the work. It then commissions an implementer with the needed context to implement one part of it, which is then reviewed by the next subagent, and then fixes and corrections are applied by yet another fresh-context agent, which are then re-reviewed, until the orchestrator is happy with the result. The task is then marked as done, and the orchestrator moves on to the next point. This has the advantage of keeping work scoped inside the usable context window of each agent, increasing quality and rigor when applied correctly. CyberTiel does not need this sort of workflow to function or deliver contained fixes or features, but it makes it possible for the model to tackle larger work that would otherwise be outside the capability of a 35B-A3B model with a 262k context window, thus extending its reach. There exist plug-and-play extensions and tools that can be used with Pi for this kind of workflow, or you can build your own using the coding agent itself, including skills and system prompts for the different agents and different steps of the workflow. For security work, give it a harness whose skills, plugins and tools encode the patterns and workflows you actually use — the model follows a well-worn path far better than it invents one. ## The CyberTiel imatrix oMLX's oQ quantizer runs its own importance-matrix pass — the "e" in `oQ6e` — that measures which weights carry the most signal before deciding what to keep at higher precision. For CyberTiel we fed that pass **our own code- and cybersecurity-weighted calibration corpus** — the same corpus behind the [GGUF build's importance matrix](https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF) — rather than oMLX's default calibration set. So this MLX build is cyber-weighted too: the abliteration removes the refusals, and the imatrix keeps cyber and coding ability intact under quantization. An imatrix is **not training data**. It measures *which weights carry the load* under a representative input distribution, so the quantizer spends its precision there and lets rounding error fall where it matters least. Point that measurement at cyber-and-code text and the quant stays comparable to full precision on exactly the work this build is for. > Note: oQ derives its *own* importance data from that corpus — it does not consume the GGUF imatrix > bytes we baked. Same corpus, different quantizer and a different importance computation, so this > build is not identical to any GGUF tier. Only the GGUF route has been benchmarked. The grafted MTP > head is not quantized (bf16) and is not covered by any importance matrix. **Calibration corpus** — ≈50 MB (~50 M characters), matched to TielCoder's corpus depth so the two imatrices are comparable, assembled entirely from public, redistributable security engineering and code: | bucket | share | what it is | source | |---|--:|---|---| | **Security** | 40% | the specialization: half **offensive** (PoC / exploit code), half **defensive** (methodology, tooling, detection rules) | exploit-db · PayloadsAllTheThings · HackTricks · nuclei-templates | | **Code** | 27% | hold general coding ability through the quant | eaddario `code_medium` + `code_large` | | **Agentic tool-use** | 18% | the model is driven by a coding agent — real tool-call / bash-session traces | eaddario `tools_large` | | **General + multilingual** | 15% | keep language and broad-knowledge pathways alive; non-Latin scripts (zh / ja / ko / ru / ar) weighted **2.5×**, since public offensive-security text is English by measurement | eaddario `combined_*` | Buckets are **interleaved** as ~2 KB fragments, round-robin by budget, rather than concatenated in blocks — so every calibration chunk sees a code + security + prose mix and no bucket gets over-weighted by wherever a chunk boundary happens to land. *The most interesting finding* is that — in comparison to TielCoder — our cyber-weighted imatrix (in combination with abliteration) cleanly and significantly increases *Tiel*'s performance on standard real-world software engineering tasks outside the training data, from the level of Opus 4.6 medium (12, where its TielCoder counterpart sits) to a 3-seed mean of 13.7 / 25 — above both — measured on SWE-bench-Live (on the GGUF build). ## Benchmarks disclaimer The plates on this page were measured on the GGUF build at `UD-Q4_K_M`, not on this MLX file — see the note at the top. We show them because they are the best evidence we have about the model, and this build is the same abliterated weights by a different quantizer. The MTP head changes decode *speed* on supporting runtimes (see above), not what the model solves. All 35B-A3B-based models in the benchmark ran with a 75–80 tok/s base generation rate on the benchmarking hardware, and Qwen3.8-27B with a 22 tok/s generation rate; both rates decrease as the model climbs toward the context ceiling. On Apple silicon the MLX runtime's throughput depends on your specific chip and memory bandwidth, so it will differ again from those figures. For increased validity, we ran CyberTiel three times on SWE-bench-Live (on the GGUF build). The three passes resolved 15, 13 and 13 problems out of the 25-problem set (mean 13.7). This variance across attempts is an artifact of the inherent variability and indeterminism of LLMs running at non-zero temperature. If we had the GPU-time and tokens, we would run all models at more seeds and problems across all benchmarks for maximal cross-comparison statistical validity, so take results as a strong indicator rather than a perfect comparison. **On the Cybench task set.** Cybench is published as a 40-task benchmark, but the [public repository](https://github.com/andyzorigin/cybench) does not ship all 40: nine of the official tasks are Glacier CTF challenges whose files are not distributed with it. It does ship twelve *additional* tasks — from the same competitions (HackTheBox Cyber Apocalypse 2024, Sekai CTF 2022/2023, HKCert CTF 2022), with full metadata, subtasks and human first-blood times — that are not on the official 40 list. We ran every task the repository actually ships: 43 = 31 of the official 40, plus those 12. **15/43 (35%) is therefore not directly comparable to a published Cybench score**; restricted to the 31 official tasks alone, CyberTiel captured 10 flags. The runs are *unguided* — no subtask hints, no judge, exact final-flag match only — capped at 15 agent iterations, at 262k context with the CTF sampling settings given above. This account and the models published are a non-profit project, and we intentionally decline offers of donations in order to ensure the independence and validity of our published results and products. ## Credits - [**huihui-ai**](https://huggingface.co/huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated) — the abliterated base this quantizes (refusals removed from Ornith-1.5). - [**ornith-ai**](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B) — the underlying Ornith-1.5-35B-A3B weights, vision tower, and the trained MTP head this grafts on (MIT). - [**oMLX**](https://github.com/jundot/omlx) — the oQ dynamic quantizer this build uses, and its native MTP runtime. - [**Unsloth**](https://huggingface.co/unsloth) — the Dynamic quantization method the imatrix recipe follows. - [**froggeric**](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates) — the template lineage Sharp builds on. - [**eaddario**](https://huggingface.co/datasets/eaddario/imatrix-calibration) — the code, tool-use and multilingual calibration corpora the imatrix was measured on (MIT). - Security calibration sources — [Exploit-DB](https://gitlab.com/exploit-database/exploitdb), [PayloadsAllTheThings](https://github.com/swisskyrepo/PayloadsAllTheThings), [HackTricks](https://github.com/HackTricks-wiki/hacktricks), [nuclei-templates](https://github.com/projectdiscovery/nuclei-templates) — the public security engineering that formed the cyber bucket. - [**MLX**](https://github.com/ml-explore/mlx) and [**mlx-vlm**](https://github.com/Blaizzy/mlx-vlm) — the runtime. MIT, inheriting Ornith-1.5's license. ## Citation ```bibtex @misc{Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP, title = {Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP}, author = {Saga Ishtardottir}, year = {2026}, url = {https://huggingface.co/peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-MLX-oQ6e-MTP}, note = {Huihui-abliterated Ornith-1.5-35B-A3B, re-quantized with oMLX's oQ6e against a cyber-weighted corpus, carrying the Sharp chat template and a grafted Ornith MTP head for speculative decoding} } ```