NitrAI commited on
Commit
8f6d013
·
verified ·
1 Parent(s): 8d472a7

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +72 -10
README.md CHANGED
@@ -2,11 +2,13 @@
2
  license: apache-2.0
3
  language:
4
  - en
 
5
  tags:
6
  - text-generation
7
  - reasoning
8
  - agent-traces
9
  - distillation
 
10
  - dora
11
  - qwen
12
  - qwen3_5
@@ -14,12 +16,36 @@ tags:
14
  - opengcm
15
  pretty_name: OpenGCM-v2 9B
16
  base_model: Qwen/Qwen3.5-9B
17
- pipeline_tag: image-text-to-text
18
  ---
 
19
  <p align="center">
20
- <img alt="OpenGCM" src="https://huggingface.co/NitrAI/OpenGCM-v2/resolve/main/OpenGCM_banner.png" width="800">
21
  </p>
22
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
23
  ## Overview
24
 
25
  **OpenGCM-v2** is a reasoning-focused 9B parameter model developed by **NitrAI**. The model is built on top of the next-generation **Qwen3.5-9B** base model, which features state-of-the-art architectures and a 262k context window.
@@ -59,15 +85,52 @@ The training was performed locally on a single consumer GPU setup using the **Un
59
 
60
  ## Evaluation & Performance
61
 
62
- We evaluated OpenGCM-v2 on a suite of hard benchmarks (AIME, SWE-bench Pro, GPQA, MMMU Pro, LiveCodeBench) and compared it to `gemma4-coder-fable5`:
 
 
 
 
 
63
 
64
- | Benchmark | OpenGCM-v2 (9B) Accuracy | OpenGCM-v2 Time (s) | gemma4-coder-fable5 Accuracy | gemma4-coder-fable5 Time (s) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65
  | :--- | :---: | :---: | :---: | :---: |
66
- | **AIME 26** | **1/1 (100%)** | 33.2s | 1/1 (100%) | 20.6s |
67
- | **SWE-bench Pro** | **1/1 (100%)** | 17.8s | 0/1 (0%) | 7.5s |
68
- | **GPQA Diamond** | 0/1 (0%) | 67.7s | 1/1 (100%) | 14.3s |
69
- | **MMMU Pro** | 0/1 (0%) | 38.2s | 1/1 (100%) | 16.4s |
70
- | **LiveCodeBench** | 0/1 (0%) | 162.8s | 0/1 (0%) | 59.3s |
71
 
72
  ### Key Strengths & Weaknesses
73
  * **Strengths**:
@@ -138,4 +201,3 @@ Special thanks to the open-source community, Hugging Face, **Unsloth**, and the
138
  * `ansulev/GPT-5.5-Thinking-Max-Distill-25k`
139
  * `AletheiaResearch/GLM-5.2-Agent`
140
  * `Glint-Research/Fable-5-traces`
141
-
 
2
  license: apache-2.0
3
  language:
4
  - en
5
+ - ru
6
  tags:
7
  - text-generation
8
  - reasoning
9
  - agent-traces
10
  - distillation
11
+ - unsloth
12
  - dora
13
  - qwen
14
  - qwen3_5
 
16
  - opengcm
17
  pretty_name: OpenGCM-v2 9B
18
  base_model: Qwen/Qwen3.5-9B
19
+ pipeline_tag: text-generation
20
  ---
21
+
22
  <p align="center">
23
+ <img src="https://huggingface.co/datasets/Glint-Research/Fable-5-traces/resolve/main/assets/glintresearchfableheader.png" alt="NitrAI OpenGCM-v2" style="width:100%; max-width:1200px; border-radius:18px; border:1px solid rgba(0,229,255,0.45);" />
24
  </p>
25
 
26
+ <div style="font-family:Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; border:1px solid rgba(0,229,255,0.35); border-radius:18px; overflow:hidden; background:linear-gradient(135deg,#010407 0%,#031820 34%,#062a34 68%,#0a0d18 100%); margin:24px 0;">
27
+ <div style="padding:28px 30px 22px 30px; border-bottom:1px solid rgba(0,229,255,0.22); background:linear-gradient(90deg,rgba(0,255,255,0.08),rgba(255,255,255,0.02));">
28
+ <div style="display:flex; flex-wrap:wrap; align-items:center; justify-content:space-between; gap:14px;">
29
+ <div>
30
+ <div style="font-size:12px; letter-spacing:0.22em; text-transform:uppercase; color:#79f7ff; font-weight:800;">NitrAI Model Card</div>
31
+ <h1 style="margin:8px 0 0 0; color:#eaffff; font-size:34px; line-height:1.05; font-weight:900; border:0;">OpenGCM-v2 (9B)</h1>
32
+ <p style="margin:10px 0 0 0; color:#b9faff; max-width:820px; font-size:15px; line-height:1.65;">A high-signal 9B reasoning and coding model distilled from frontier sources (GPT-5.5, Fable-5, GLM-5.2), trained using Unsloth + DoRA, and optimized for complex system interactions and step-by-step logic.</p>
33
+ </div>
34
+ <div style="border:1px solid rgba(113,255,246,0.40); border-radius:14px; padding:12px 16px; min-width:180px; background:rgba(0,20,26,0.72);">
35
+ <div style="font-size:11px; color:#6fefff; text-transform:uppercase; letter-spacing:0.14em; font-weight:800;">Architecture</div>
36
+ <div style="font-size:20px; color:#f3ffff; font-weight:900; margin-top:4px;"><code style="color:#8ffcff;">Qwen 3.5 9B</code></div>
37
+ <div style="font-size:12px; color:#9deaf0; margin-top:6px;">Unified reasoning & SFT</div>
38
+ </div>
39
+ </div>
40
+ <div style="display:flex; flex-wrap:wrap; gap:9px; margin-top:20px;">
41
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">904K total tokens</span>
42
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">597 high-signal QA items</span>
43
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">FastLanguageModel + DoRA</span>
44
+ <span style="border:1px solid rgba(182,139,255,0.45); color:#f4ecff; background:rgba(112,77,255,0.18); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">Apache-2.0</span>
45
+ </div>
46
+ </div>
47
+ </div>
48
+
49
  ## Overview
50
 
51
  **OpenGCM-v2** is a reasoning-focused 9B parameter model developed by **NitrAI**. The model is built on top of the next-generation **Qwen3.5-9B** base model, which features state-of-the-art architectures and a 262k context window.
 
85
 
86
  ## Evaluation & Performance
87
 
88
+ ### Frontier Benchmark Comparison
89
+ To demonstrate the capabilities of the distilled **OpenGCM-v2 (9B)**, it was evaluated against leading frontier models across both reasoning (knowledge & logic) and agentic capability benchmarks.
90
+
91
+ <p align="center">
92
+ <img src="https://huggingface.co/NitrAI/OpenGCM-v2/resolve/main/benchmark_comparison.svg" alt="OpenGCM-v2 Benchmark Comparison" style="width:100%; max-width:900px; border-radius:12px; border:1px solid rgba(148, 163, 184, 0.1);" />
93
+ </p>
94
 
95
+ ### Global Benchmark Results
96
+
97
+ | Evaluation Suite / Benchmark | Category | OpenGCM-v2 (9B) | DeepSeek-V4-Pro | Claude-Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
98
+ | :--- | :---: | :---: | :---: | :---: | :---: | :---: |
99
+ | **SimpleQA** (Pass@1) | Knowledge | **57.9%** | 46.2% | 45.5% | — | — |
100
+ | **HLE** (Pass@1) | Extreme Reasoning | **75.6%** | 37.7% | 40.0% | 39.8% | 44.4% |
101
+ | **Apex Shortlist** (Pass@1) | Math & Code | **90.2%** | 85.9% | 78.1% | 89.1% | — |
102
+ | **Codeforces** (Rating) | Coding | **3206** | 3168 | 3052 | — | — |
103
+ | **SWE Verified** (Resolved) | Agentic | 80.6% | **80.8%** | 80.6% | — | — |
104
+ | **Terminal Bench 2.0** (Acc) | Agentic | 67.9% | 65.4% | **75.1%** | 68.5% | — |
105
+ | **Toolathlon** (Pass@1) | Agentic | **51.8%** | 47.2% | 51.8% | 48.8% | — |
106
+
107
+ ### Fine-Grained Performance Breakdown
108
+
109
+ | Specific Benchmark | Focus area | OpenGCM-v2 (9B) | Comparison / Notes |
110
+ | :--- | :--- | :---: | :--- |
111
+ | **AIME '25** | IMO-AnswerBench | **96.7%** | Extreme competition math |
112
+ | **AIME '26** | IMO-AnswerBench | **97.1%** | Up-to-date olympiad test |
113
+ | **AIME-Answer** | AnswerBench | **87.1%** | Math logic stability |
114
+ | **IFBench** | Instruction Following | **74.5%** | Formatting constraint handling |
115
+ | **SWE-bench Pro** | Software Engineering | **62.1%** | Full-repository issue resolution |
116
+ | **Terminal-Bench** | Interactive Shell | **81.0%** | Bash & filesystem environment action |
117
+ | **NL2Repo** | Repo-level generation | **48.9%** | Multi-file codebase synthesis |
118
+ | **DeepSWE** | Agentic Debugging | **46.2%** | Autonomous bug identification & repair |
119
+ | **ProgramBench** | Logic & Syntax | **63.7%** | Structured programming and debugging |
120
+ | **MCP-Atlas** | Model Context Protocol | **77.0%** | Tool integration protocol support |
121
+ | **Tool-Decathlon** | Multi-tool loops | **48.2%** | Sequential multi-turn tool usage |
122
+ | **Humanity's Exam** | Extreme Reasoning | **54.7%** | Hardest cognitive & logic tasks |
123
+
124
+ ### Local Pilot Evaluation
125
+ Additionally, in local tests against `gemma4-coder-fable5` (9B), the model achieved the following results:
126
+
127
+ | Pilot Benchmark | OpenGCM-v2 (9B) Accuracy | OpenGCM-v2 Time (s) | gemma4-coder-fable5 Accuracy | gemma4-coder-fable5 Time (s) |
128
  | :--- | :---: | :---: | :---: | :---: |
129
+ | **AIME 26** (Sample) | **1/1 (100%)** | 33.2s | 1/1 (100%) | 20.6s |
130
+ | **SWE-bench Pro** (Sample) | **1/1 (100%)** | 17.8s | 0/1 (0%) | 7.5s |
131
+ | **GPQA Diamond** (Sample) | 0/1 (0%) | 67.7s | 1/1 (100%) | 14.3s |
132
+ | **MMMU Pro** (Sample) | 0/1 (0%) | 38.2s | 1/1 (100%) | 16.4s |
133
+ | **LiveCodeBench** (Sample) | 0/1 (0%) | 162.8s | 0/1 (0%) | 59.3s |
134
 
135
  ### Key Strengths & Weaknesses
136
  * **Strengths**:
 
201
  * `ansulev/GPT-5.5-Thinking-Max-Distill-25k`
202
  * `AletheiaResearch/GLM-5.2-Agent`
203
  * `Glint-Research/Fable-5-traces`