--- language: - en license: apache-2.0 base_model: Qwen/Qwen3-30B-A3B-Instruct-2507 datasets: - OpenMOSS-Team/SciJudgeBench tags: - scientific-taste - GRPO pipeline_tag: text-generation library_name: transformers model-index: - name: SciJudge-30B-2605 results: - task: type: text-generation name: Citation-based pairwise paper judgment dataset: name: SciJudgeBench MAIN_1000 type: OpenMOSS-Team/SciJudgeBench split: test metrics: - type: accuracy name: Avg. accuracy value: 82.7 --- # SciJudge-30B-2605 SciJudge-30B-2605 is a Qwen3-30B-A3B-Instruct-2507 MoE model fine-tuned for scientific paper evaluation. Given two papers' titles, abstracts, and publication dates, it predicts which paper has higher citation impact. This release is part of [AI Can Learn Scientific Taste](https://arxiv.org/abs/2603.14473). The companion smaller model is [SciJudge-4B-2605](https://huggingface.co/OpenMOSS-Team/SciJudge-4B-2605), and the benchmark dataset is [SciJudgeBench](https://huggingface.co/datasets/OpenMOSS-Team/SciJudgeBench). Resources: [Project page](https://tongjingqi.github.io/AI-Can-Learn-Scientific-Taste/) and [GitHub repository](https://github.com/tongjingqi/AI-Can-Learn-Scientific-Taste). ## Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "OpenMOSS-Team/SciJudge-30B-2605" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype=torch.bfloat16, device_map="auto", ) messages = [ {"role": "system", "content": "You are a helpful assistant. You first think about the reasoning process in your mind and then provide the user with the answer."}, {"role": "user", "content": "Today is 2025-12-10. Based on the titles, abstracts, and publication dates of the following two papers A and B, determine which paper has a higher citation count.\nShow your reasoning process in tags. And return the final answer in tags. The final answer should contain only 'A' or 'B'.\n\nPaper A:\nTitle: ...\nAbstract: ...\nDate: ...\n\nPaper B:\nTitle: ...\nAbstract: ...\nDate: ..."} ] text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(text, return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.7, top_p=0.8, top_k=20) response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True) print(response) ``` ## Training Details - **Base model:** [Qwen/Qwen3-30B-A3B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507) - **Training method:** GRPO with DAPO loss - **Reward:** external preference reward for citation-based pairwise judgment - **Training data:** 720,341 preference pairs from SciJudgeBench - **Precision:** bfloat16 - **KL coefficient:** 0.03 ## Evaluation Accuracy on the SciJudgeBench `test` split, the 1,000-example MAIN_1000 in-domain evaluation set: | Model | CS | Math | Physics | Others | Avg. | | --- | ---: | ---: | ---: | ---: | ---: | | Qwen3-30B-A3B-Instruct-2507 | 73.39 | 79.90 | 64.41 | 63.94 | 69.7 | | SciJudge-30B-2605 | 83.47 | 89.71 | 81.18 | 77.40 | 82.7 | ## Citation ```bibtex @misc{tong2026ailearnscientifictaste, title={AI Can Learn Scientific Taste}, author={Jingqi Tong and Mingzhe Li and Hangcheng Li and Yongzhuo Yang and Yurong Mou and Weijie Ma and Zhiheng Xi and Hongji Chen and Xiaoran Liu and Qinyuan Cheng and Ming Zhang and Qiguang Chen and Weifeng Ge and Qipeng Guo and Tianlei Ying and Tianxiang Sun and Yining Zheng and Xinchi Chen and Jun Zhao and Ning Ding and Xuanjing Huang and Yugang Jiang and Xipeng Qiu}, year={2026}, eprint={2603.14473}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2603.14473}, } ```