Instructions to use zgcagi/ZGCM-1-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zgcagi/ZGCM-1-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zgcagi/ZGCM-1-7B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("zgcagi/ZGCM-1-7B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zgcagi/ZGCM-1-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zgcagi/ZGCM-1-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zgcagi/ZGCM-1-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zgcagi/ZGCM-1-7B
- SGLang
How to use zgcagi/ZGCM-1-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zgcagi/ZGCM-1-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zgcagi/ZGCM-1-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zgcagi/ZGCM-1-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zgcagi/ZGCM-1-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use zgcagi/ZGCM-1-7B with Docker Model Runner:
docker model run hf.co/zgcagi/ZGCM-1-7B
A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
📄 Tech Report · 🤗 Model · 🤗 Data · 📊 Results · 💻 Training Code · 💬 WeChat Community
Introduction
ZGCM-1 is a 7.39B-parameter dense language model trained from scratch, built for mathematical reasoning and tool-assisted search. It combines deliberate internal thinking with active information gathering, supporting 256K-token context and both thinking and direct-response modes in a single model.
The project brings together an efficient hybrid-attention architecture, FP8 training with Muon, progressive long-context mid-training, and general-agentic supervised fine-tuning. Researcher-directed AI agents contribute throughout development, from data curation and cluster operations to evaluation and deployment.
Highlights
- Math and reasoning at 7B scale. ZGCM-1 achieves 97.13% on MATH-500, 75.00% on AIME 2026, and 70.42% on HMMT 2025, with the best average rank across the report's 14 reasoning benchmarks among the seven compared 7B–8B models.
- Search beyond model memory. Multi-step tool use reaches 63.09% on WebWalkerQA, 19.43% on BrowseComp, and 62.00% on Binary Function Search.
- Efficient long context. Gated sliding-window and global attention deliver 3.94× training throughput at 256K compared with full attention in the report's architecture experiments. The report estimates an approximately 4.2× improvement in 16K pretraining time-to-loss from combined architecture, precision, optimizer, and normalization gains.
- An open research recipe. The training repository releases the data-processing, pretraining, mid-training, and SFT workflows, with stage-specific configurations and runtime documentation.
Quickstart
ZGCM-1 ships its own modeling code, so trust_remote_code=True is required.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "zgcagi/ZGCM-1-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="bfloat16",
device_map="auto",
)
messages = [{"role": "user", "content": "Compute 1+1."}]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=4096,
do_sample=True,
temperature=1.0,
top_p=1.0,
)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Evaluation Results
The following results are from the technical report, using the 256K SFT checkpoint in thinking mode. Non-agentic evaluations use temperature 1.0, top-p 1.0, and mean pass@1 over 32 runs unless otherwise specified.
Selected reasoning benchmarks
| Benchmark (%) | ZGCM-1 | DeepSeek-R1-0528-Qwen3-8B | MiniCPM4.1-8B | Qwen3-8B | Olmo 3 7B Think |
|---|---|---|---|---|---|
| MATH-500 | 97.13 | 96.32 | 95.60 | 96.20 | 95.10 |
| AIME 2024 | 80.62 | 83.33 | 83.33 | 80.00 | 71.60 |
| AIME 2025 | 73.33 | 75.21 | 73.33 | 63.33 | 64.60 |
| AIME 2026 | 75.00 | 69.17 | 71.67 | 66.67 | 66.16 |
| HMMT 2025 | 70.42 | 61.50 | 52.50 | 43.33 | 43.89 |
| HMMT 2026 | 59.48 | 51.52 | 46.21 | 45.45 | 43.94 |
Selected rows and models from Table 2; bold marks the best score in each displayed row. The full evaluation covers 20 benchmarks, including code, knowledge, and instruction following.
Agentic search
| Benchmark | ZGCM-1 (%) | Setting |
|---|---|---|
| WebWalkerQA | 63.09 | Web search and page reading |
| BrowseComp | 19.43 | Web search and page reading |
| GAIA (text-only) | 42.52 | Web search and page reading |
| Binary Function Search | 62.00 | 31/50 exact function-entry matches using Ghidra tools |
Source: Tables 3–4. Web research allows up to 64 search-and-read steps. Binary Function Search uses a separate protocol on 50 tasks from 10 held-out projects.
Architecture and Training
| Model specification | ZGCM-1 |
|---|---|
| Architecture | Decoder-only dense Transformer |
| Parameters | 7.39B |
| Layers / hidden size | 32 / 4,096 |
| Attention | 27 gated sliding-window layers + 5 global layers |
| Local window / GQA heads | 128 tokens / 32 query heads, 8 KV heads |
| Maximum context | 262,144 tokens (256K) |
The training recipe described in the report has three main stages:
- Pretraining: approximately 4.19T tokens, combining curriculum-based data mixing with hybrid FP8 precision and Muon optimization.
- Mid-training: approximately 600B tokens with context extended from 16K → 64K → 256K. Interaction traces are reformulated as Markov Decision Process (MDP) state-action transitions to supervise individual decisions.
- Supervised fine-tuning: joint general and agentic training with mixed thinking/direct-response examples, execution-verified trajectories, and assistant-only loss. The report also explores mixed RL for mathematics and code.
Resources
| Resource | Location |
|---|---|
| Model weights | zgcagi/ZGCM-1-7B |
| Data | zgcagi/ZGCM-1-Data |
| Training code | github.com/zgcagi/ZGCM-1 |
For dataset usage, see the ZGCM-1-Data card. Model specifications, evaluation details, and figures are presented in ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search.
Training Code
The complete training pipeline is released at github.com/zgcagi/ZGCM-1, covering data processing, pretraining, mid-training, supervised fine-tuning, and reinforcement learning. Each stage is a top-level directory with its own code, configurations, and runtime documentation.
WeChat Community
Scan the QR code to join the ZGCM-1 community group. Click the image to open it at full size. If the code has expired, please open a Discussion and ask the maintainers for the latest one.
License
ZGCM-1 is released under the MIT License. Third-party licenses and notices bundled with the training code are listed in the training repository; the dataset carries its own terms, given on the ZGCM-1-Data card.
- Downloads last month
- 67


