---
base_model:
- Qwen/Qwen3.5-9B
- TokenRhythm/NeoHorse-1-9B
library_name: transformers
tags:
- noesis
- noesis-repack
- bf16
- mtp
- qwen3_5
- text-generation-inference
- reasoning
- distillation
- deepseek
- sft
- rl
- gspo
- math
- stem
- tool-use
- function-calling
- lora-merge
- dare-ties
- agentic
- tool-use
- coding
- reasoning
- instruction-following
- heretic
- uncensored
- decensored
- abliterated
license: apache-2.0
language:
- en
- ru
- zh
- vi
- kk
- ja
- af
- am
- ar
- as
- ast
- az
- be
- bg
- bn
- bs
- ca
- ceb
- ckb
- cs
- cy
- da
- de
- el
- es
- et
- eu
- fa
- ff
- fi
- fil
- fr
- ga
- gl
- gn
- gu
- ha
- he
- hi
- hr
- hu
- hy
- id
- ig
- is
- it
- jv
- ka
- kam
- kea
- km
- kmr
- kn
- ko
- ky
- lb
- lg
- ln
- lo
- lt
- luo
- lv
- mi
- mk
- ml
- mn
- mr
- ms
- mt
- mvy
- my
- ne
- nl
- "no"
- nso
- ny
- oc
- om
- "or"
- pa
- pl
- ps
- pt
- qxp
- ro
- rw
- sd
- sk
- skr
- sl
- sn
- so
- sr
- sv
- sw
- ta
- te
- tg
- th
- ti
- tk
- tr
- ug
- uk
- umb
- ur
- uz
- wo
- xh
- yo
- yue
- zu
---
## NOESIS / AMAImedia
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform.
- **Founder:** Ilia Bolotnikov
- **Organization:** AMAImedia.com
- **X (Twitter):** [@AMAImediacom](https://x.com/AMAImediacom)
- **LinkedIn:** [Ilia Bolotnikov](https://www.linkedin.com/in/ilia-bolotnikov)
- **Telegram:** [@djbionicl](https://t.me/djbionicl)
- **Release date:** 2026-09-07
-
## AMAImedia
- Released **Heretic repository:** [Dingdust/NeoHorse-1-9B-heretic](https://huggingface.co/Dingdust/NeoHorse-1-9B-heretic)
- **Original repository:** [TokenRhythm/NeoHorse-1-9B](https://huggingface.co/TokenRhythm/NeoHorse-1-9B)
---
# This is a decensored version of a model, made using [Heretic](https://heretic-project.org) v1.4.0
## Abliteration parameters
| Parameter | Value |
| :-------- | :---: |
| **direction_index** | 16.33 |
| **attn.o_proj.max_weight** | 1.48 |
| **attn.o_proj.max_weight_position** | 19.08 |
| **attn.o_proj.min_weight** | 1.46 |
| **attn.o_proj.min_weight_distance** | 16.49 |
| **mlp.down_proj.max_weight** | 1.44 |
| **mlp.down_proj.max_weight_position** | 18.83 |
| **mlp.down_proj.min_weight** | 1.43 |
| **mlp.down_proj.min_weight_distance** | 13.39 |
## Performance
| Metric | This model | Original model (a model) |
| :----- | :--------: | :---------------------------: |
| **KL divergence** | 0.0181 | 0 *(by definition)* |
| **Refusals** | 18/100 | 97/100 |
-----
NeoHorse-1-9B
Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.
Technical Report
NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward **recursive self-improvement (RSI)**.
It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.
Derived from [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) and fine-tuned by TokenRhythm.
This release contains **language-model weights only** and is repackaged for text-only inference.
Vision weights are not included. Repackaging changes configuration and tensor key names, without changing the fine-tuned tensor values.
## Highlights
- **Path toward RSI:** the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation鈥搒election鈥搖pdate loop; extending this loop across successive iterations is the next step toward RSI.
- **Agentic post-training framework:** the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
- **Data quality:** exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
- **Broad gains:** 69.04 macro average across ten benchmarks versus 65.60 for Qwen3.5-9B (**+3.44**).
## Model Details
| Property |
Value |
| Model family |
NeoHorse Agent-Native Causal Language Model |
| Parameters |
Approximately 9B |
| Base model |
Qwen3.5-9B |
| Post-training |
Routing-guided agentic post-training |
| Interface |
Text input and text output |
| Context length |
262,144 natively and extensible up to 1,010,000 tokens. |
| Weight format / precision |
Safetensors / BF16 |
## Evaluation
The 9B track compares NeoHorse-1-9B with five representative open-weight baselines: Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and Muse-Glimmer-30B. Results cover ten benchmarks and are grouped by capability. Higher is better; `螖` is NeoHorse-1-9B minus Qwen3.5-9B. **Bold** and underline mark the best and second-best results in each benchmark row, respectively; ties share the same formatting.
| Benchmark |
Granite-4.2-8B |
Qwen3.5-9B |
Ornith-1.5-9B |
Gemma-4-12B-it |
Muse-Glimmer-30B |
NeoHorse-1-9B |
螖 vs Qwen3.5-9B |
| 馃 Agentic |
QwenClawBench | 37.01 | 44.04 | 47.27 | 43.53 | 46.11 | 48.73 | +4.69 |
WorkBuddy Bench | 35.07 | 39.60 | 29.29 | 29.65 | 45.85 | 40.15 | +0.55 |
PinchBench | 56.93 | 74.55 | 68.22 | 58.89 | 71.35 | 82.25 | +7.70 |
VitaBench | 23.00 | 31.25 | 26.75 | 36.50 | 48.50 | 42.25 | +11.00 |
BFCL v4 | 52.06 | 64.88 | 65.03 | 62.06 | 53.74 | 67.43 | +2.55 |
tau2-Bench | 62.28 | 88.04 | 83.68 | 59.37 | 76.64 | 90.82 | +2.78 |
| 馃捇 Coding |
HumanEval | 96.34 | 92.68 | 93.90 | 100.00 | 98.17 | 98.17 | +5.49 |
LiveCodeBench v6 | 72.00 | 65.14 | 47.43 | 73.14 | 65.71 | 65.14 | +0.00 |
| 馃摎 Instruction Following |
IFBench | 78.00 | 66.33 | 40.00 | 77.67 | 78.67 | 66.33 | +0.00 |
IFEval | 92.98 | 89.46 | 71.35 | 94.27 | 93.90 | 89.09 | -0.37 |
| 馃搳 Overall |
Ten-benchmark average | 60.57 | 65.60 | 57.29 | 63.51 | 67.86 | 69.04 | +3.44 |
> **Reported protocol:** SGLang v0.5.17 路 `temperature=1.0` 路 `top_p=0.95` 路 `top_k=20` 路 `min_p=0.0` 路 `presence_penalty=1.5` 路 `repetition_penalty=1.0` 路 thinking mode enabled with `enable_thinking=true` and `force_nonempty_content=true`. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.
## Deployment
The examples below are for self-hosted deployment from a downloaded local checkpoint.
### Local checkpoint path
The examples below assume the checkpoint has already been downloaded to local disk. Set `MODEL_PATH` to the directory containing `config.json`, tokenizer files, and model weights.
```bash
MODEL_PATH="/path/to/NeoHorse-1-9B"
```
The OpenAI-compatible requests below use the server's `--served-model-name` (for example, `neohorse-1-9b`), not the filesystem path.
### SGLang
The technical report uses SGLang v0.5.17.
```bash
pip install "sglang==0.5.17"
MODEL_PATH="/path/to/NeoHorse-1-9B"
python3 -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--served-model-name neohorse-1-9b \
--host 0.0.0.0 \
--port 30000 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
```
Send an OpenAI-compatible request after the server starts:
```bash
curl http://localhost:30000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'
```
### vLLM
```bash
pip install -U vllm
MODEL_PATH="/path/to/NeoHorse-1-9B"
vllm serve "$MODEL_PATH" \
--served-model-name neohorse-1-9b \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
```
The server exposes an OpenAI-compatible `/v1/chat/completions` endpoint. Send a request after the server starts:
```bash
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'
```
The example uses the configured 262,144-token context limit.
Actual capacity depends on GPU memory and serving settings; reduce the context limit if needed.
These launch examples have not yet been validated on GPU for this repackaged release.
## License
NeoHorse-1-9B is released under the **Apache License 2.0**.
The upstream model is [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B).
Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file.
TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference.
Modification notices are included in this model card and the released configuration, weight index, and Safetensors metadata.
## Citation
```
@misc{neohorse2026,
title = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
author = {NeoHorse Team},
year = {2026},
howpublished = {arXiv preprint}
}
```
For questions or issue reports, use the [NeoHorse project repository](https://github.com/TokenRhythm/NeoHorse).