---
license: apache-2.0
base_model: Qwen/Qwen3.5-27B
base_model_relation: finetune
library_name: transformers
pipeline_tag: image-text-to-text
language:
- en
tags:
- qwen3.5
- reasoning
- multimodal
- vision
- image-text-to-text
- uncensored
- long-context
- 1M-context
- mtp
- multi-token-prediction
- function-calling
- tool-use
- agentic
- sft
- dpo
- esft
- full-fine-tune
- cybersecurity
- biomedical
---
# Qwythos-27B-v1
**Developed by [Empero](https://empero.org)**
**Qwythos-27B-v1** is an open-weight, full-parameter reasoning model: the larger sibling of [Qwythos-9B](https://huggingface.co/empero-ai/Qwythos-9B-Claude-Mythos-5-1M), trained on the exact same curriculum and sized up on a Qwen3.5-27B base. It is a complete pre-RL checkpoint, post-trained through **SFT → DPO → ESFT**.
The point of this release is simple: **nothing was ablated to make it fit.** Qwythos retains the base checkpoint's native multi-token-prediction head, full vision tower, and 1,048,576-token context configuration—all intact and active.
- **🧩 Nothing ablated** — native MTP, the full vision tower, and a 1,048,576-token YaRN context window are present together. Most community 27B fine-tunes give up at least one of them.
- **🧠 The bigger sibling** — the same Qwythos-9B curriculum at 27B. In a closed-book organophosphate-poisoning evaluation, the 27B volunteered that physostigmine should be avoided, where the 9B needed a web search to get that fact right.
- **🛠 Built for terminal trajectories** — held-out terminal/tool-session perplexity falls from **356.6** in the base to **2.76** after training, and the model transfers that format to an emulated Claude-Code session.
- **🔧 Native function calling** — the packaged chat template carries Qwen3.5's full tool-use spec, so `tools=[...]` works out of the box with no wrapper and no tool-specific fine-tune.
> Qwythos is intentionally uncensored for technical and research use. It engages substantively with sensitive-but-legal cybersecurity, biomedical, and technical questions; use appropriate review and application-level controls for your deployment.
## Why this checkpoint
Qwythos-27B-v1 preserves the parts of the base that are often dropped during a community fine-tune:
- **Native MTP is preserved.** The one-layer native MTP head and its 15 tensors were preserved verbatim; SFT did not touch them, and they were re-attached through the safetensors index. Runtimes that support it can use the head for self-speculative decoding.
- **Vision is still here.** The Qwen3.5 vision tower is intact: 333 vision tensors, depth 27, hidden size 1152, and output hidden size 5120. Training was text-only and the vision parameters were frozen, so image behavior is inherited from the base rather than newly tuned.
- **The full context configuration remains enabled.** The 262,144-token native context is extended with YaRN factor 4.0 to **1,048,576 tokens**. The configuration was verified numerically finite through position 1,048,575.
The larger model changes what can stay closed-book. On an organophosphate-poisoning prompt, Qwythos-27B correctly volunteered “Avoid physostigmine (it can worsen toxicity)” and gave correct atropine/pralidoxime dosing without tools. That specific fact required a web search in the 9B evaluation. It is a useful, concrete illustration of what the size-up buys—not a substitute for verification in clinical use.
## Agentic and terminal transfer
The strongest measured shift is target-format acquisition on held-out terminal/tool sessions: **356.6 → 2.76 perplexity** versus the base, evaluated on assistant tokens with identical inputs and sequences capped at 24,576 tokens. In a held-out emulated Claude-Code session for an authorized lab network assessment, Qwythos consumed injected `nmap` output, reasoned about scan trade-offs, and emitted the correct next shell-form tool call.
It also emits first-person `...` reasoning rather than the base model's tendency toward meta-planning outlines. Greedy long-form generations in the reviewed set were clean under the loop check, with no degenerate repetition.
## Native function calling
The packaged chat template implements Qwen3.5's tool-use spec in full. Pass `tools=[...]` to `apply_chat_template` and the model emits standard tool-call blocks — no wrapper, no tool-specific fine-tune:
```
value
```
Combined with the terminal/agentic training above, this is the intended deployment path: give Qwythos real tools and let it verify its specifics rather than recalling them closed-book.
## Identity
The packaged chat template introduces the model as *"You are Qwythos, a model created by Empero AI. Only bring up your identity if the user asks."* — it identifies itself only when asked.
## v1, before RL
This is **v1**, the pre-RL checkpoint. It has received full-parameter SFT, DPO, and ESFT, and is released as a usable model in its own right. An RL-trained v2 will follow soon. DPO and ESFT hyperparameters are not being published.
## Model details
| Property | Value |
|---|---|
| Base | [`Qwen/Qwen3.5-27B`](https://huggingface.co/Qwen/Qwen3.5-27B) |
| Architecture | Dense Qwen3.5-27B; hybrid Gated-DeltaNet linear attention with a full-attention layer every fourth layer (3:1) |
| Text backbone | 64 hidden layers; hidden size 5120; 24 attention heads; 4 KV heads; vocabulary 248,320 |
| Modalities | Text + vision, with native MTP |
| Precision | bf16 |
| Context | 1,048,576 tokens: 262,144 native × YaRN factor 4.0 |
| Native MTP | One hidden layer (`mtp_num_hidden_layers: 1`) |
## Training
Qwythos-27B was post-trained on the same curriculum family as Qwythos-9B — STEM and reasoning first, then the agentic target domain — scaled up to 27B.
The pipeline is **full-parameter SFT → DPO → ESFT**, run in bf16 with assistant-only loss on long, untruncated sequences. Qwythos-27B-v1 is the checkpoint at the end of that pipeline, before reinforcement learning.
Training data, per-stage hyperparameters, and the packing and kernel setup are not being published.
## How to use
The base is multimodal. For text-only inference, load it with `AutoModelForImageTextToText`:
```python
from transformers import AutoModelForImageTextToText, AutoTokenizer
model_id = "empero-ai/Qwythos-27B-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
dtype="bfloat16",
device_map="auto",
)
messages = [
{
"role": "user",
"content": "Explain the trade-offs in this terminal plan and propose the next command.",
}
]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(
**inputs,
do_sample=True,
temperature=0.6,
top_p=0.95,
top_k=20,
repetition_penalty=1.05,
max_new_tokens=16384,
)
new_tokens = output[0][inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
```
Use a recent `transformers` with Qwen3.5 support and Gated-DeltaNet kernels (`flash-linear-attention` and a CUDA-matched `causal_conv1d`). Without them, the linear-attention layers fall back to slower PyTorch operations.
### With tools (function calling)
```python
TOOLS = [
{"type": "function", "function": {
"name": "run_shell",
"description": "Run a shell command and return combined stdout/stderr.",
"parameters": {"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]}}},
{"type": "function", "function": {
"name": "web_search",
"description": "Search the web for current facts and citations.",
"parameters": {"type": "object",
"properties": {"query": {"type": "string"},
"max_results": {"type": "integer"}},
"required": ["query"]}}},
]
prompt = tokenizer.apply_chat_template(
messages, tools=TOOLS, tokenize=False, add_generation_prompt=True
)
# then parse ...
```
### Serving at 1M context with vLLM
```bash
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve empero-ai/Qwythos-27B-v1 \
--max-model-len 1010000
```
The static YaRN factor of 4.0 has a short-context quality trade-off. For roughly 512k context, use factor 2.0; if you will never exceed the 262,144-token native window, restore `rope_parameters.rope_type` to `"default"` from the included `config.json.pre_yarn` backup for maximum short-context fidelity.
### GGUF: llama.cpp, Ollama, and LM Studio
For GGUF releases, including llama.cpp, Ollama, and LM Studio workflows, see **[empero-ai/Qwythos-27B-v1-GGUF](https://huggingface.co/empero-ai/Qwythos-27B-v1-GGUF)**. The recommended default is `Qwythos-27B-Q4_K_M.gguf`; the MTP-enabled default is `Qwythos-27B-MTP-Q4_K_M.gguf`. Image use also requires `mmproj-Qwythos-27B-F16.gguf` alongside the selected text quant.
## Sampling recommendations
Qwythos-27B is a reasoning model and inherits Qwen3.5's thinking-mode behaviour. These are the shipped defaults in `generation_config.json` and the recommended starting point:
```python
gen_kwargs = dict(
do_sample=True,
temperature=0.6, # Qwen3.5 thinking-mode recommended
top_p=0.95,
top_k=20,
repetition_penalty=1.05,
max_new_tokens=16384, # generous budget for the block + final answer
)
```
| Parameter | Value |
|---|---|
| `temperature` | 0.6 (agentic / tool use) · 1.0 (open-ended or creative) |
| `top_p` | 0.95 |
| `top_k` | 20 |
| `repetition_penalty` | 1.05 |
| `max_new_tokens` | 16,384+ |
- **Agentic, harness, or tool use → temperature 0.6.** Tighter sampling keeps tool-call syntax and multi-step plans on the rails.
- **Open-ended reasoning or creative work → temperature 1.0.** Loosen it up when you want the model to explore rather than commit early.
- **Give it room.** Reasoning traces are long; budget **16,384+ `max_new_tokens`** so the `` block and the final answer both fit.
Use `repetition_penalty=1.05` — a small deviation from Qwen's default of 1.0 that prevents rare non-terminating reasoning loops on long generations.
## Limitations
- **Vision is inherited rather than tuned.** The vision tower is intact, but text-only training did not tune or evaluate image understanding.
- **Long context has a trade-off.** Static YaRN at factor 4.0 slightly degrades short-context quality; reduce or remove the factor when the full range is unnecessary.
- **Uncensored is a deployment responsibility.** The model engages with sensitive-but-legal technical content; provide your own policy, review, and access controls where appropriate.
## Stay in the loop
Sign up for the Empero newsletter at [empero.org](https://empero.org) for releases, evaluations, and research notes on Qwythos and future open-weight models.
## Support / Donate
If this model helped you, consider supporting the project:
- **BTC:** `bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7v`
- **LTC:** `ltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7x`
- **XMR:** `42Dbm5xg5Nq26fdyzfEU7KBnAJfhi7Cvz5J2ex5CzHXkfKuNEJzYCcmJ1GTbgjFZ5MBx72sdG1G9239Cd6rsZfv4QeDkYJY`
## Provenance & licensing
Qwythos-27B-v1 is derived from **Qwen3.5-27B**. The weights are released under **Apache-2.0**, inherited from the Qwen3.5-27B base, for research and experimentation as-is.
## Acknowledgements
- Developed and released by [Empero](https://empero.org)
- Base model: [Qwen3.5-27B](https://huggingface.co/Qwen/Qwen3.5-27B) (Alibaba Qwen team)
- Training: [TRL](https://github.com/huggingface/trl) + [Transformers](https://github.com/huggingface/transformers)
- GGUF quantization: [llama.cpp](https://github.com/ggml-org/llama.cpp) (ggml-org)