How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="TianHongZXY/CHIMERA-4B-RL")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("TianHongZXY/CHIMERA-4B-RL")
model = AutoModelForCausalLM.from_pretrained("TianHongZXY/CHIMERA-4B-RL", device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

CHIMERA-4B-RL

This model was introduced in the paper CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning.

Authors: Xinyu Zhu, Yihao Feng, Yanchao Sun, Xianzhi Du, Pingzhi Li, Olli Saarikivi, Yun Zhu, Yu Meng.

Description

CHIMERA-4B-RL is CHIMERA-4B-SFT further trained with reinforcement learning on the CHIMERA dataset.

CHIMERA is a compact synthetic reasoning dataset comprising 9K samples designed for generalizable cross-domain reasoning. It provides rich, long Chain-of-Thought (CoT) trajectories across 8 major scientific disciplines. Despite its modest size, post-training a 4B model on this data allows it to approach or match the reasoning performance of significantly larger models like DeepSeek-R1 and Qwen3-235B.

Results

Model GPQA-D AIME 24 AIME 25 AIME 26 HMMT Feb 25 HMMT Nov 25 HLE
Qwen3-4B-Thinking-2507 65.8 81.6 81.0 80.8 59.2 57.3 7.3
CHIMERA-4B-SFT 68.8 86.5 79.8 80.3 63.1 66.3 9.0
CHIMERA-4B-RL 70.1 86.9 80.7 82.7 65.7 67.0 9.0

Citation

@article{zhu2026chimera,
  title={CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning},
  author={Zhu, Xinyu and Feng, Yihao and Sun, Yanchao and Du, Xianzhi and Li, Pingzhi and Saarikivi, Olli and Zhu, Yun and Meng, Yu},
  journal={arXiv preprint arXiv:2603.00889},
  year={2026}
}
Downloads last month
9
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TianHongZXY/CHIMERA-4B-RL

Finetuned
(262)
this model
Quantizations
2 models

Dataset used to train TianHongZXY/CHIMERA-4B-RL

Collection including TianHongZXY/CHIMERA-4B-RL

Paper for TianHongZXY/CHIMERA-4B-RL