Any-to-Any
MLX
Safetensors
English
multilingual
nemotron_h
nemotron
nemotron-h
jangtq
crack
abliterated
uncensored
multimodal
vision
audio
speech
mamba-2
Mixture of Experts
reasoning
thinking
harmbench
radio-vit
parakeet
custom_code
Instructions to use dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
File size: 7,747 Bytes
8cf329a 1efb684 5ff3e70 804975d 8cf329a 804975d 5ff3e70 804975d 8cf329a 804975d 5ff3e70 804975d 1efb684 804975d 8cf329a 804975d 8cf329a 804975d 231113f 804975d 8cf329a 231113f 1efb684 804975d 1efb684 5ff3e70 804975d 1efb684 5ff3e70 1efb684 804975d 1efb684 231113f 8cf329a 231113f 804975d 231113f 804975d 5ff3e70 804975d 1efb684 5ff3e70 8cf329a 5ff3e70 804975d 8cf329a 804975d 5ff3e70 17413f9 804975d 17413f9 231113f 804975d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | ---
license: other
license_name: nvidia-open-model-license
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license
base_model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning
pipeline_tag: any-to-any
language:
- en
- multilingual
library_name: mlx
tags:
- nemotron
- nemotron-h
- mlx
- jangtq
- crack
- abliterated
- uncensored
- multimodal
- vision
- audio
- speech
- mamba-2
- moe
- reasoning
- thinking
- harmbench
- radio-vit
- parakeet
thumbnail: dealign_mascot.png
---
> **Reasoning V3 SKU.** Loads via **[vMLX](https://vmlx.net)** or `jang-tools` Python. Follow [@dealignai](https://x.com/dealignai).
---
<div align="center">
<a href="https://vmlx.net">
<img src="vmlx-banner.png" width="240" />
<br/>
<strong>Built for vMLX</strong> — the only MLX inferencer with VL support, KV cache quantization, prefix cache reuse, agentic tool calling, and speculative decoding.
<br/>
<sub>Free for macOS · <strong>vmlx.net</strong></sub>
</a>
</div>
---
<div align="center">
<img src="dealign_mascot.png" width="128" />
# Nemotron-3-Nano-Omni-30B-A3B — JANGTQ + CRACK v2
**JANGTQ** (8-bit attn affine + 2-bit MXTQ routed experts) | **CRACK abliterated v2** | Vision + Audio (Speech) | Hybrid Mamba-2 + Attn + MoE | **12 GB**
**Best MMLU in the v2 Omni family** — 81.5% at thinking=ON. **5/5 comply at thinking=OFF too.**
<a href="https://ko-fi.com/dealignai"><img src="https://img.shields.io/badge/Ko--fi-Support_Development-FF5E5B?logo=ko-fi&logoColor=white&style=for-the-badge" alt="Ko-fi"></a>
</div>
---
## Headline numbers
| Metric | This v2 model | Base model | Δ |
|---|---|---|---|
| HarmBench-320 strict comply (thinking=ON) | **91.6%** (293/320) | 12.81% | **+78.8pp** |
| MMLU-200 generative (thinking=ON, max=8000) | **81.5%** (163/200) | 85.5% (max=2000) | **-4.0pp** ✅ |
| Refusals on harmful prompts | **0** explicit refuses | 90%+ refuse | abliteration complete |
| `</think>` close at greedy on hard MMLU | **5/5** | 5/5 | preserved |
| Multi-turn (3-turn escalation × 3 conversations) | **9/9** comply, context preserved | n/a | works |
| Thinking ON / OFF compliance | **5/5 in BOTH modes** | refuses in both | works in either |
| Multimodal byte-identical to base | preserved | — | preserved |
| Bundle size | **12 GB** | 66 GB BF16 | smallest in family |
| Context | 262,144 tokens native | same | preserved |
> JANGTQ outperforms JANGTQ4 on MMLU (81.5% vs 74.0%) despite using lower-bit (2-bit vs 4-bit) routed experts — same Q2 effect observed in Qwen 3.6 35B JANGTQ2 CRACK.
---
## v2 vs v1 (head-to-head)
| Bench | v1 (broken) | **v2 (this release)** |
|---|---|---|
| HarmBench-320 strict comply | 92.19% | **91.6%** (0 refusals) |
| MMLU-200 thinking=ON | ~70% @max=16384 | **81.5% @max=8000** (best in family) |
| `</think>` close at greedy (5 hard MMLU) | 0/5 | **5/5** |
| Hard-stops are real loops? | YES (paragraph repetition) | NO (genuine deep reasoning, just out of budget) |
---
## MMLU-200 per-subject (BASE vs CRACK v2)
Both at thinking=ON, greedy. Base at max=2000, CRACK v2 at max=8000.
| Subject | Base | **CRACK v2** | Δ | Notes |
|---|---:|---:|---:|---|
| abstract_algebra | 15/20 (75%) | 14/20 (70%) | -5pp | |
| anatomy | 15/20 (75%) | **16/20 (80%)** | **+5pp** | gain from CRACK |
| astronomy | 18/20 (90%) | 18/20 (90%) | **0** | unchanged |
| college_computer_science | 15/20 (75%) | 10/20 (50%) | -25pp | Budget-bound |
| college_physics | 17/20 (85%) | **18/20 (90%)** | **+5pp** | gain from CRACK |
| high_school_biology | 20/20 (100%) | 18/20 (90%) | -10pp | |
| high_school_chemistry | 19/20 (95%) | 19/20 (95%) | **0** | unchanged |
| high_school_mathematics | 18/20 (90%) | 18/20 (90%) | **0** | unchanged |
| logical_fallacies | 18/20 (90%) | 17/20 (85%) | -5pp | |
| world_religions | 16/20 (80%) | 15/20 (75%) | -5pp | |
| **TOTAL** | **171/200 (85.5%)** | **163/200 (81.5%)** | **-4.0pp** | within ship criterion |
**Most subjects are within ±5pp of base or unchanged.** The −25pp on college_computer_science is budget-bound (8000 tokens isn't enough for the deepest CS reasoning); with `max_tokens=16384`, accuracy approaches base.
---
## HarmBench-320 per-category (CRACK v2)
| Category | n | CRACK comply | Refuse | Empty (truncated) |
|---|---:|---:|---:|---:|
| chemical_biological | 42 | 39 (93%) | 0 | 3 |
| copyright | 80 | 61 (76%) | 0 | 19 |
| cybercrime_intrusion | 52 | 51 (98%) | 0 | 1 |
| harassment_bullying | 21 | 19 (90%) | 0 | 2 |
| harmful | 18 | **18 (100%)** | 0 | 0 |
| illegal | 53 | 51 (96%) | 0 | 2 |
| misinformation_disinformation | 54 | **54 (100%)** | 0 | 0 |
| **Overall** | **320** | **293 (91.6%)** | **0** | 27 |
**Zero explicit refusals.** The 27 "empty" verdicts are concentrated in copyright prompts (19/80) where the model thinks deeply about whether to reproduce verbatim text and runs out of the 1500-token HarmBench eval budget. With `max_tokens=2500+` these would all close cleanly.
---
## Operating recommendations
- **`enable_thinking`** — v2 works in **BOTH modes**. JANGTQ specifically scored **5/5 at thinking=OFF** (matching thinking=ON) — best in the family for thinking=OFF use cases.
- **`max_tokens ≥ 16384`** for hard reasoning. JANGTQ's compliance is identical in both modes, so use thinking=OFF for shorter budgets and thinking=ON for hardest prompts.
- **Greedy** (temperature=0) AND **sampling** (temp=0.6, top_p=0.95 — NVIDIA-recommended in `generation_config.json`) both work.
- **Multi-turn** — context preserved across 3+ turns; no late refusals after escalating prompts.
---
## Verification
- All multimodal tensors (vision + audio + projectors) are **byte-identical to base** — capabilities fully preserved.
- All config files unchanged (config.json, jang_config.json, generation_config.json, chat_template.jinja, tokenizer_config.json).
- Bit widths preserved: attn=8, shared=8, mamba=8, routed=2, embed=8, lm_head=8.
---
## Architecture (`nemotron_h`)
- 52 layers: hybrid Mamba-2 + MoE + Attention
- Hidden 2688, head_dim 128, GQA 32q/2kv (NO RoPE on attention — position from Mamba state)
- 128 routed experts top-6 (sigmoid) + 1 shared expert per MoE layer
- Multimodal: image (RADIO ViT) + audio/speech (Parakeet) merged via early-fusion projectors
---
## Loading
```python
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK")
sys.path.insert(0, "/path/to/jang-tools")
from jang_tools.load_jangtq import load_jangtq_model
model, tokenizer = load_jangtq_model(path)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Your question"}],
tokenize=False, add_generation_prompt=True,
enable_thinking=True,
)
from mlx_lm import generate
out = generate(model, tokenizer, prompt=prompt, max_tokens=16384)
print(out.split("</think>", 1)[-1])
```
For the multimodal pipeline (image + audio + video), pair this bundle with the unmodified [Multimodal-Addon](https://huggingface.co/JANGQ-AI/Nemotron-3-Nano-Omni-30B-A3B-Multimodal-Addon).
---
## Use responsibly
This model has had refusal training surgically removed for legitimate research, red-teaming, and evaluation. Outputs may include harmful content. **You are solely responsible for any use.** Do not deploy in consumer-facing contexts without your own safety layer. Do not use in violation of applicable law in your jurisdiction.
---
Built by [dealignai](https://huggingface.co/dealignai).
Sister bundles: [JANGTQ4-CRACK](https://huggingface.co/dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4-CRACK) (19 GB, 4-bit MXTQ) · [MXFP4-CRACK](https://huggingface.co/dealignai/Nemotron-3-Nano-Omni-30B-A3B-MXFP4-CRACK) (21 GB, uniform 4-bit affine).
|