---
library_name: transformers
language:
- en
tags:
- falcon-h1r
- unsloth
license: other
license_name: falcon-llm-license
license_link: https://falconllm.tii.ae/falcon-terms-and-conditions.html
base_model:
- tiiuae/Falcon-H1R-7B
---
> [!NOTE]
> Includes Unsloth **chat template fixes**!
For `llama.cpp`, use `--jinja`
>
Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.
# Falcon-H1R-7B
This repository presents **Falcon-H1R-7B**, a reasoning-specialized model built on top of [Falcon-H1-7B-Base](https://huggingface.co/tiiuae/Falcon-H1-7B-Base) and trained via cold-start supervised fine-tuning with
long reasoning traces and further enhanced by scaling RL with GRPO. The model demonstrates outstanding performance across various benchmark evaluations, including mathematics, programming, instruction following, and general logic.
## Model Description
- **Developed by:** [Technology Innovation Institute](https://www.tii.ae)
- **Model type:** Causal decoder-only
- **Architecture:** Hybrid (Transformers + Mamba2) architecture
- **Language(s):** English, Multilingual
- **License:** [Falcon-LLM License](https://falconllm.tii.ae/falcon-terms-and-conditions.html)
## Training details
For more details about the training protocol of this model, please refer to the [Falcon-H1R technical blogpost](https://falcon-lm.github.io/blog/falcon-h1r-7b) and [Technical Report](https://github.com/tiiuae/falcon-h1r/blob/main/tech_report.pdf).
# Usage
Currently to use this model, you can either rely on Hugging Face `transformers`, `vLLM` or `SGLang` library.
## Inference
Make sure to install the latest version of `transformers` or `vLLM` or `SGLang`.
```bash
pip install transformers
pip install mamba-ssm[causal-conv1d]
```
For vLLM, make sure to install `vllm=0.11.0`:
```bash
pip install "vllm>=0.11.0"
```
## Sampling Parameters
We recommend using a **temperature** of **0.6** and **top-p** as **0.95** with max new tokens up to 65536.
For supported frameworks, you can adjust the repetition_penalty and presence_penalty parameters to reduce endless repetitions.
For reasoning tasks with continuous batching and requiring higher max new tokens, we recommend to use TP=2.
## 🤗 Transformers
Refer to the snippet below to run H1R models using 🤗 transformers.
Model will generate think content wrapped in a `| Category | Benchmark | Falcon-H1R-7B | Qwen3-8B | DeepSeek-R1-0528-Qwen3-8B | Phi-4-Reasoning-Plus-14B | Apriel-1.5-15b-Thinker | GPT-OSS-20B | Qwen3-32B | Nemotron-H-47B-Reasoning |
|---|---|---|---|---|---|---|---|---|---|
| MATH | AIME24 | 88.1 | 77.9 | 83.3 | 77.2 | 86.2 | 83.3 | 79.4 | 64.6 |
| AIME25 | 83.1 | 65.8 | 75.8 | 71.2 | 80.0 | 84.4 | 71.0 | 51.4 | |
| HMMT25 | 64.9 | 41.0 | 54.3 | 47.7 | 61.0 | 64.8 | 49.8 | 34.2 | |
| AMO-BENCH | 36.3 | 14.1 | 23.3 | 15.0 | 22.2 | 26.0 | 21.3 | 7.0 | |
| MATH500 | 97.4 | 97.4 | 96.8 | 95.4 | 97.2 | 94.8 | 96.8 | 91.4 | |
| Code | LCBv5-v6 | 68.6 | 53.0 | 57.2 | 53.1 | 53.0 | 72.0 | 61.0 | 47.4 |
| SciCode (sub/main) | 28.3 / 3.9 | 28.3 / 6.7 | 22.2 / 2.6 | 29.8 / 7.2 | 31.9 / 8.2 | 34.9 / 6.2 | 36.4 / 9.2 | 26.1 / 4.6 | |
| General | GPQA-D | 61.3 | 61.2 | 61.4 | 67.9 | 68.2 | 61.2 | 67.3 | 56.8 |
| MMLU-Pro | 72.1 | 63.5 | 69.1 | 79.2 | 76.5 | 75.6 | 73.9 | 78.6 | |
| HLE | 11.1 | 4.2 | 5.6 | 5.9 | 12.0 | 9.8 | 8.3 | 4.4 | |
| IFBench | 53.4 | 35.3 | 29.2 | 51.7 | 55.8 | 69.4 | 35.4 | 34.3 | |
| Agentic Workflows | 𝜏²-Bench Telecom | 25.4 | 27.8 | 68.4 | 60.2 | 29.8 | 11.4 | ||
| Terminal-Bench Hard | 4.9 | 2.1 | 1.4 | 2.1 | 9.9 | 9.9 | 2.8 | 1.4 |
| Benchmark | Falcon-H1R-7B | Qwen3-8B | DeepSeek-R1-0528-Qwen3-8B | Nemotron-H-8B | Phi-4-Reasoning-Plus-14B | Qwen3-32B |
|---|---|---|---|---|---|---|
| AIME24 | 96.7 | 80.0 | 90.0 | 53.3 | 86.7 | 86.7 |
| AIME25 | 96.7 | 80.0 | 82.8 | 43.3 | 83.3 | 86.7 |
| GPQA-D | 70.2 | 60.9 | 59.9 | 61.1 | 73.2 | 70.1 |
| AMO-Bench* | 35.9 | 15.4 | 25.6 | 7.7 | 20.5 | 28.2 |