DedeProGames commited on
Commit
e23feb2
·
verified ·
1 Parent(s): 9793d8e

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +205 -0
README.md ADDED
@@ -0,0 +1,205 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - OrionLLM/GRM-2.6-Plus
5
+ pipeline_tag: image-text-to-text
6
+ ---
7
+
8
+ <p align="center">
9
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/685ea8ff7b4139b6845ce395/_66bkNH630dGeIt2Uuctd.png" alt="logo" width="500">
10
+ </p>
11
+ <div align="center">
12
+ <a href="https://huggingface.co/OrionLLM/GRM-2.6-Plus-0628/" style="text-decoration: none;">
13
+ <img src="https://img.shields.io/badge/🤗-HuggingFace-FC926C?style=for-the-badge" alt="HuggingFace">
14
+ </a>
15
+ <a href="https://huggingface.co/collections/OrionLLM/grm-26" style="text-decoration: none;">
16
+ <img src="https://img.shields.io/badge/📚-Collection-3B82F6?style=for-the-badge" alt="Collection">
17
+ </a>
18
+ <a href="https://grape.skinnertopia.com/chat" style="text-decoration: none;">
19
+ <img src="https://img.shields.io/badge/💬-Chat-22C55E?style=for-the-badge" alt="Chat">
20
+ </a>
21
+ <a href="https://www.apache.org/licenses/LICENSE-2.0" style="text-decoration: none;">
22
+ <img src="https://img.shields.io/badge/📜-License-E343BD?style=for-the-badge" alt="License">
23
+ </a>
24
+ </div>
25
+
26
+ ## 1. Introduction
27
+
28
+ GRM-2.6-Plus-0628 is a **27B-parameter reasoning model** and a small update to **GRM-2.6-Plus**, built for **general-purpose AI** and optimized for **difficult, high-complexity tasks**. It is designed to deliver stronger performance for its size while remaining practical, efficient, and accessible for advanced local and research-oriented use.
29
+
30
+ This version improves upon **GRM-2.6-Plus** with a focus on **long-horizon agentic tasks** and the ability to **solve harder problems**, allowing it to better compete head-to-head with frontier models. The model focuses on **structured reasoning**, helping it produce more accurate, coherent, and reliable responses across demanding problems. GRM-2.6-Plus-0628 brings **elite-level reasoning** to complex workloads, making it suitable for users who need a capable model for advanced problem-solving, coding, agents, and everyday intelligence.
31
+
32
+ ## 2. Key Capabilities
33
+
34
+ - **Elite-Level Reasoning for Hard Tasks:** GRM-2.6-Plus-0628 is optimized to handle difficult reasoning workloads with clarity, consistency, and strong step-by-step problem-solving ability.
35
+ - **Improved Long-Horizon Agentic Performance:** This update specifically targets long-horizon agentic workflows, enabling the model to maintain coherence and effectiveness across extended multi-step tasks.
36
+ - **High Performance for Its Size:** With **27B parameters**, the model is designed to deliver excellent capability relative to its scale, balancing strong intelligence with practical deployment.
37
+ - **Advanced Coding and Agentic Use:** GRM-2.6-Plus-0628 is well suited for code generation, structured problem-solving, tool-style workflows, and local agentic applications.
38
+ - **Optimized for Practical Deployment:** The model aims to remain efficient and usable across capable consumer and workstation hardware while offering strong performance for advanced tasks.
39
+
40
+ ## 3. Performance
41
+
42
+ GRM-2.6-Plus-0628 is designed to be a highly capable **27B local AI model** for complex reasoning, coding, everyday chat, and agentic workflows. It focuses on delivering **better performance for its size**, making it a strong option for users who want powerful reasoning without relying only on massive-scale models.
43
+
44
+ Its core strength is **practical intelligence**: elite-level reasoning, strong task understanding, stable responses, and the ability to handle difficult problems across multiple domains.
45
+
46
+ ### Detailed Benchmarks
47
+
48
+ <table>
49
+ <tr>
50
+ <th style="background: rgba(128,128,128,0.1); text-align: center;"> </th>
51
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">GRM-2.6-Plus-0628</th>
52
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">GRM-2.6-Plus</th>
53
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">Qwen3.6-27B</th>
54
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">google/gemma-4-31B-it</th>
55
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">GPT-5.4-Mini</th>
56
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">Claude-4.5-Haiku</th>
57
+ </tr>
58
+ <tr>
59
+ <td align="center" colspan="7" style="background: linear-gradient(90deg, rgba(124,58,237,0.45) 0%, rgba(99,102,241,0.42) 50%, rgba(59,130,246,0.45) 100%); font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>Knowledge &amp; STEM</i></td>
60
+ </tr>
61
+ <tr>
62
+ <td align="center">MMLU-Pro</td>
63
+ <td align="center"><b>88.1</b></td>
64
+ <td align="center">86.8</td>
65
+ <td align="center">86.2</td>
66
+ <td align="center">85.2</td>
67
+ <td align="center">--</td>
68
+ <td align="center">80.0</td>
69
+ </tr>
70
+ <tr>
71
+ <td align="center">MMLU-Redux</td>
72
+ <td align="center"><b>96.4</b></td>
73
+ <td align="center">94.2</td>
74
+ <td align="center">93.5</td>
75
+ <td align="center">93.7</td>
76
+ <td align="center">--</td>
77
+ <td align="center">--</td>
78
+ </tr>
79
+ <tr>
80
+ <td align="center">C-Eval</td>
81
+ <td align="center"><b>92.4</b></td>
82
+ <td align="center">92.0</td>
83
+ <td align="center">91.4</td>
84
+ <td align="center">82.6</td>
85
+ <td align="center">--</td>
86
+ <td align="center">--</td>
87
+ </tr>
88
+ <tr>
89
+ <td align="center">GPQA Diamond</td>
90
+ <td align="center"><b>90.1</b></td>
91
+ <td align="center">88.3</td>
92
+ <td align="center">87.8</td>
93
+ <td align="center">84.3</td>
94
+ <td align="center">88.0</td>
95
+ <td align="center">73.0</td>
96
+ </tr>
97
+ <tr>
98
+ <td align="center">SuperGPQA</td>
99
+ <td align="center"><b>67.5</b></td>
100
+ <td align="center">66.4</td>
101
+ <td align="center">66.0</td>
102
+ <td align="center">65.7</td>
103
+ <td align="center">--</td>
104
+ <td align="center">--</td>
105
+ </tr>
106
+ <tr>
107
+ <td align="center" colspan="7" style="background: linear-gradient(90deg, rgba(124,58,237,0.45) 0%, rgba(99,102,241,0.42) 50%, rgba(59,130,246,0.45) 100%); font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>Reasoning &amp; Coding</i></td>
108
+ </tr>
109
+ <tr>
110
+ <td align="center">LiveCodeBench v6</td>
111
+ <td align="center"><b>86.5</b></td>
112
+ <td align="center">84.8</td>
113
+ <td align="center">83.9</td>
114
+ <td align="center">80.0</td>
115
+ <td align="center">--</td>
116
+ <td align="center">51.1</td>
117
+ </tr>
118
+ <tr>
119
+ <td align="center">HMMT Feb 26</td>
120
+ <td align="center"><b>85.9</b></td>
121
+ <td align="center">84.8</td>
122
+ <td align="center">84.3</td>
123
+ <td align="center">77.2</td>
124
+ <td align="center">--</td>
125
+ <td align="center">--</td>
126
+ </tr>
127
+ <tr>
128
+ <td align="center">AIME26</td>
129
+ <td align="center"><b>95.6</b></td>
130
+ <td align="center">95.1</td>
131
+ <td align="center">94.1</td>
132
+ <td align="center">89.2</td>
133
+ <td align="center">--</td>
134
+ <td align="center">--</td>
135
+ </tr>
136
+ <tr>
137
+ <td align="center" colspan="7" style="background: linear-gradient(90deg, rgba(124,58,237,0.45) 0%, rgba(99,102,241,0.42) 50%, rgba(59,130,246,0.45) 100%); font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>General Agent</i></td>
138
+ </tr>
139
+ <tr>
140
+ <td align="center">SWE-bench Verified</td>
141
+ <td align="center"><b>79.7</b></td>
142
+ <td align="center">77.7</td>
143
+ <td align="center">77.2</td>
144
+ <td align="center">52.0</td>
145
+ <td align="center">--</td>
146
+ <td align="center">73.3</td>
147
+ </tr>
148
+ <tr>
149
+ <td align="center">SWE-bench Pro</td>
150
+ <td align="center"><b>56.1</b></td>
151
+ <td align="center">54.0</td>
152
+ <td align="center">53.5</td>
153
+ <td align="center">35.7</td>
154
+ <td align="center">54.4</td>
155
+ <td align="center">--</td>
156
+ </tr>
157
+ <tr>
158
+ <td align="center">Terminal-Bench 2.0</td>
159
+ <td align="center"><b>62.6</b></td>
160
+ <td align="center">59.8</td>
161
+ <td align="center">59.3</td>
162
+ <td align="center">42.9</td>
163
+ <td align="center">60.0</td>
164
+ <td align="center">41.0</td>
165
+ </tr>
166
+ </table>
167
+
168
+ ## 4. Family
169
+ The GRM-2.6 family is available in various sizes to suit every case.
170
+
171
+ <table>
172
+ <tr>
173
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">Model</th>
174
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">Size</th>
175
+ <th style="background: rgba(128,128,128,0.1); text-align: center;">Domain</th>
176
+ </tr>
177
+ <tr>
178
+ <td align="center">GRM-2.6-Plus-0628</td>
179
+ <td align="center">27B</td>
180
+ <td align="center">Updated model for extremely difficult tasks with improved long-horizon agentic performance</td>
181
+ </tr>
182
+ <tr>
183
+ <td align="center">GRM-2.6-Plus</td>
184
+ <td align="center">27B</td>
185
+ <td align="center">Powerful model for extremely difficult tasks</td>
186
+ </tr>
187
+ <tr>
188
+ <td align="center">GRM-2.6-Opus</td>
189
+ <td align="center">27B</td>
190
+ <td align="center">Merge of GRM-2.6-Plus optimized for difficult terminal and coding tasks</td>
191
+ </tr>
192
+ </table>
193
+
194
+ ## 5. Architecture
195
+ GRM-2.6-Plus-0628 is built on the Qwen3.6 architecture and is optimized for complex tasks, agent environments, and everyday chat.
196
+
197
+ GRM-2.6-Plus-0628 applies the same principle to a stronger, larger foundation, resulting in a model that punches above its weight class on structured reasoning tasks while remaining deployable on consumer hardware.
198
+
199
+ ---
200
+
201
+ <div align="center">
202
+
203
+ **GRM-2.6-Plus-0628** is developed by **[OrionLLM](https://huggingface.co/OrionLLM)** and released under the Apache 2.0 License.
204
+
205
+ </div>