Text Generation
Transformers
Safetensors
qwen3_5_moe
image-text-to-text
fp8
fp8-block
compressed-tensors
llm-compressor
qwen3.5
Mixture of Experts
code
reasoning
conversational
Instructions to use shisa-ai/Ornith-1.0-35B-FP8-BLOCK with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shisa-ai/Ornith-1.0-35B-FP8-BLOCK with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="shisa-ai/Ornith-1.0-35B-FP8-BLOCK") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shisa-ai/Ornith-1.0-35B-FP8-BLOCK") model = AutoModelForMultimodalLM.from_pretrained("shisa-ai/Ornith-1.0-35B-FP8-BLOCK", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use shisa-ai/Ornith-1.0-35B-FP8-BLOCK with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "shisa-ai/Ornith-1.0-35B-FP8-BLOCK" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shisa-ai/Ornith-1.0-35B-FP8-BLOCK", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/shisa-ai/Ornith-1.0-35B-FP8-BLOCK
- SGLang
How to use shisa-ai/Ornith-1.0-35B-FP8-BLOCK with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "shisa-ai/Ornith-1.0-35B-FP8-BLOCK" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shisa-ai/Ornith-1.0-35B-FP8-BLOCK", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "shisa-ai/Ornith-1.0-35B-FP8-BLOCK" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shisa-ai/Ornith-1.0-35B-FP8-BLOCK", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use shisa-ai/Ornith-1.0-35B-FP8-BLOCK with Docker Model Runner:
docker model run hf.co/shisa-ai/Ornith-1.0-35B-FP8-BLOCK
File size: 5,252 Bytes
e4e8a1d 1991f8f e4e8a1d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | ---
library_name: transformers
license: mit
license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
pipeline_tag: text-generation
base_model: deepreinforce-ai/Ornith-1.0-35B
base_model_relation: quantized
tags:
- fp8
- fp8-block
- compressed-tensors
- llm-compressor
- qwen3.5
- moe
- code
- reasoning
---
<img width="600px" src="assets/ornith_logo.png">
# Ornith-1.0-35B-FP8-BLOCK
This repository contains a public `FP8_BLOCK` / `compressed-tensors`
quantization of [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B).
The original model is Ornith-1.0-35B, a Qwen3.5 MoE-family reasoning and
agentic coding model released by DeepReinforce. This repository only changes
the checkpoint representation; it does not introduce additional fine-tuning.
This model is a faithful quant of the original Ornith-1.0-35B and therefore **does not** have an MTP head, however, we have also uploaded a [shisa-ai/Ornith-1.0-35B-FP8-BLOCK-MTP](https://huggingface.co/shisa-ai/Ornith-1.0-35B-FP8-BLOCK-MTP) model that grafts the Qwen 3.6 35B-A3B official MTP head onto the model. This is the best-performing MTP and based on our testing it gives a >20% throughput uplift for coding workloads (tested on RTX PRO 6000 and vLLM v0.23.0).
## Quantization Summary
- Source model: [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B)
- Quantized model: `shisa-ai/Ornith-1.0-35B-FP8-BLOCK`
- Quantization tool: `llm-compressor` model-free PTQ
- Quantization format: `compressed-tensors`
- Scheme: `FP8_BLOCK`
- Calibration data: none; this is data-free/model-free PTQ
- Weight quantization: static FP8, symmetric, block strategy, `128x128` blocks
- Activation quantization: dynamic FP8, symmetric, group strategy, group size `128`
- Target modules: `Linear`
- `compressed-tensors` metadata version recorded in `config.json`: `0.15.1.a20260406`
The local quantization environment used for this artifact reported:
- `llmcompressor=0.10.1.dev67+ga1cec6fa`
- `compressed-tensors=0.15.1a20260406`
- `transformers=5.5.0`
The definitive machine-readable quantization metadata is in
[`config.json`](config.json).
## Unquantized / Ignored Paths
The quantization run intentionally skipped the following module patterns:
```text
re:.*lm_head$
re:.*embed_tokens$
re:.*visual.*
re:.*mlp\.gate$
re:.*mlp\.shared_expert_gate$
re:.*linear_attn.*
re:^mtp.*
```
Practical implications:
- `lm_head` and token embeddings remain unquantized.
- MoE router/gating paths remain unquantized.
- Qwen3.5 linear-attention/Gated-DeltaNet paths remain unquantized.
- Vision tower weights remain unquantized in this artifact.
- No MTP tensors were present in the source checkpoint used here, so this model
card does not advertise MTP/speculative decoding support.
## Example Usage
Use a recent runtime that supports both `Qwen3_5MoeForConditionalGeneration`
and `compressed-tensors` `FP8_BLOCK` checkpoints.
### vLLM
```bash
vllm serve shisa-ai/Ornith-1.0-35B-FP8-BLOCK \
--served-model-name Ornith-1.0-35B-FP8-BLOCK \
--tensor-parallel-size 8 \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
```
Adjust `--tensor-parallel-size`, `--max-model-len`, and memory settings for your
hardware. Long-context serving still requires substantial KV-cache memory even
though the checkpoint weights are compressed.
### OpenAI-Compatible Client
```python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="Ornith-1.0-35B-FP8-BLOCK",
messages=[
{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
],
temperature=0.6,
top_p=0.95,
max_tokens=1024,
)
message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
```
Ornith is a reasoning model. With vLLM's `--reasoning-parser qwen3`, reasoning
tokens are surfaced separately as `reasoning_content`; final answers remain in
`content`.
## Quality and Performance Notes
This upload documents the quantized checkpoint format and provenance. It does
not publish a new retained benchmark table for the quantized model. For
production use, validate quality and end-to-end serving latency against the
BF16 source model on your target workload and runtime.
For original source-model capabilities, benchmark descriptions, and training
context, see the upstream card:
[`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B).
## License and Attribution
The source model is MIT licensed. This derivative quantized artifact keeps the
source license metadata and links to the upstream license file.
If you use the source model, cite the original Ornith release:
```bibtex
@misc{ornith-35b,
title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
url = {https://deep-reinforce.com/ornith_1_0.html},
author = {{DeepReinforce Team}},
year = {2026}
}
```
|