---
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.6-35B-A3B
tags:
- agent
- agentic
- co-work
- tool-use
- long-context
- mixture-of-experts
- coding
---
Occamy-1.0
Open Pareto-frontier 35B Intelligence for Co-work
Project Website |
Model Weights |
Training Framework
## 1. Model Introduction
Occamy-1.0 is a compact agentic model purpose-built for real-world co-work: long-horizon, stateful tasks that require coordinated use of search, code, tools, files, structured APIs, and productivity software. Starting from the post-trained [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) checkpoint, Occamy concentrates further training on reliable execution, persistent state tracking, recovery, and follow-through rather than relearning general capabilities from scratch.
### Key Features
- **Co-work specialization:** Designed for sustained execution across multi-step professional workflows, not isolated question answering.
- **Compact inference footprint:** A 35B-total, 3B-active Mixture-of-Experts model that keeps long-running agent workloads practical.
- **Long-horizon continuity:** Designed to keep work coherent across tool calls, delegated runs, and history rewrites such as context compaction.
- **Broad agentic capability:** Co-work gains are accompanied by strong tool calling, terminal coding, and instruction following.
- **Execution-grounded training:** Supervised fine-tuning spans general agentic work, long-horizon interaction, software engineering, and tool-call grounding.
- **Open training stack:** The multi-harness reinforcement-learning infrastructure used to train Occamy is released as [Dressage](https://github.com/Accio-Lab/Dressage).
> [!NOTE]
> Occamy is optimized for common co-work workloads, not as a replacement for frontier models on every task. Retrieval-heavy and simulated-user tasks still have headroom, and native browser or desktop visual interaction is not part of the current co-work training interface.
## 2. Model Summary
| Architecture | Mixture-of-Experts causal model with vision encoder |
| Total Parameters | 35B |
| Activated Parameters | 3B |
| Number of Layers | 40 |
| Number of Experts | 256 |
| Activated Experts | 8 routed + 1 shared |
| Base Architecture Context | 262,144 tokens |
| SFT Sequence Length | 131,072 tokens |
| Starting Checkpoint | Qwen3.6-35B-A3B |
| Post-training | Full-parameter SFT, HDPO, model merging, and SAO |
Architecture fields follow the starting checkpoint's published model card. Occamy post-trains the language backbone without changing the architecture; the vision encoder and projector are frozen during SFT. The released checkpoint configuration remains the source of truth for serving limits.
## 3. Evaluation Results
### Full Evaluation
| Benchmark |
35B-A3B Models |
Large-scale Models |
| Occamy-1.0 |
Qwen3.6 35B-A3B |
Agents-A1 |
Nex-N2-mini |
BigBang-1.0 |
Ornith-1.5 |
GPT-5.6 Sol |
Qwen3.8-Max |
DeepSeek V4 Pro (0813) |
GLM-5.2 |
| Co-work |
| Claw-Eval (average) | 82.20 | 69.50 | 69.90 | 66.60 | 63.50 | 64.40 | 81.80 | 83.92 | 81.70 | 81.60 |
| Claw-Eval (Pass³) | 71.40 | 54.80 | 41.70 | 37.00 | 40.20 | 48.70 | 68.90 | 73.68 | 74.50 | 68.30 |
| WildClawBench | 49.16 | 40.40 | 30.73 | 30.31 | 32.87 | 45.91 | 67.20 | 54.42 | 37.30 | 52.14 |
| CommerceAgentBench | 37.38 | 19.60 | 9.30 | 16.80 | 30.80 | 37.40 | 49.50 | 46.30 | 43.30 | 39.30 |
| Business Arena | $79,868 | $44,751 | $33,626 | $13,325 | $56,477 | $66,292 | $168,867 | $89,423 | $40,804 | $55,742 |
| GDPval† | 1,128 | 1,004 | 869 | 999 | 951 | 855 | 1,741 | 1,640 | 1,500 | 1,452 |
| OfficeQA Pro | 48.10 | 39.10 | 23.30 | 46.60 | 43.60 | 59.40 | 74.40 | 69.20 | 51.20 | 66.20 |
| τ³-Bench (Banking) | 37.10 | 11.90 | 7.20 | 25.80 | 10.30 | 21.70 | 46.90 | 54.60 | 44.30 | 37.10 |
| Tool calling |
| AutomationBench (Pass¹) | 27.60 | 7.50 | 2.20 | 5.70 | 14.80 | 18.50 | 45.50 | 43.50 | 32.00 | 28.00 |
| AutomationBench (partial) | 69.10 | 39.40 | 14.70 | 27.90 | 47.40 | 58.00 | 81.20 | 81.20 | 59.70 | 70.00 |
| BFCL v4 | 65.40 | 63.19 | 57.23 | 62.81 | 57.86 | 68.51 | 64.33 | 73.65 | 67.10 | 70.33 |
| VitaBench | 41.75 | 34.25 | 37.00 | 26.25 | 46.00 | 40.25 | 46.75 | 52.25 | 53.50 | 43.75 |
| Coding |
| Terminal-Bench 2.1 | 59.00 | 49.50 | 41.60 | 60.70* | 33.70 | 67.80* | 88.80 | 81.30* | 87.90* | 82.70 |
| Instruction following |
| IFEval | 91.53 | 86.90 | 91.60 | 91.60 | 90.50 | 81.80 | 95.00 | 95.02 | 93.74 | 93.89 |
Within each size group, **bold** denotes the best result and underlining denotes the second-best result. * Official model-card or Artificial Analysis result. † Reproduced on the public task release.
### Cost-Performance
Across Claw-Eval, WildClawBench, AutomationBench, and GDPval, Occamy-1.0 lies near the low-cost knee of the empirical Pareto frontier. Relative to its Qwen3.6-35B-A3B starting checkpoint, it delivers a large aggregate capability gain with only a modest change in measured per-task inference cost. Benchmark scores are equally weighted after per-benchmark min-max normalization, and costs are macro-averaged per task under the frozen pricing protocol used in the report.
## 4. Training Recipe
Occamy uses staged specialization and consolidation:
```text
Qwen3.6-35B-A3B
├─ Marathon Expert: SFT → HDPO ┐
└─ Sprint Expert: SFT ├─ Uniform merge → SAO → Occamy-1.0
```
The Marathon Expert learns sustained execution and accuracy-conditioned efficiency, while the Sprint Expert preserves broader agentic capability. A uniform parameter-space merge combines both experts into one checkpoint with no inference-time routing or ensembling, and a final Single-Rollout Asynchronous Optimization (SAO) stage refines the merged policy on a broad co-work mixture.
The deduplicated SFT union across both experts is:
| Data source | Trajectories | Average length | Tokens |
| --- | ---: | ---: | ---: |
| General agentic | 5,418 | 37.7K | 204.1M |
| Long-horizon interactive agents | 923 | 95.8K | 88.4M |
| Terminal and software engineering | 1,228 | 35.1K | 43.1M |
| Tool-call grounding | 7,429 | 9.1K | 67.7M |
| **Overall** | **14,998** | **26.9K** | **403.3M** |
Training tasks are grounded in executable environments with observable state transitions and task-level grading. The open-source [Dressage](https://github.com/Accio-Lab/Dressage) stack provides multi-harness execution, token-exact trajectory capture, sandbox integration, and multi-segment conversion for reinforcement learning.
## 5. Deployment
Occamy-1.0 keeps the Qwen3.6-35B-A3B architecture, so the [upstream deployment recipe](https://huggingface.co/Qwen/Qwen3.6-35B-A3B#deployment) is the reference serving path. The examples below mirror that recipe with eight-way tensor parallelism and its full context length; adjust both to fit your hardware and confirm them against the released Occamy checkpoint configuration.
### SGLang
The upstream model card recommends [SGLang](https://github.com/sgl-project/sglang) 0.5.10 or newer for the Qwen3.6 architecture.
```bash
python -m sglang.launch_server \
--model-path Accio-Lab/Occamy-1.0 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
```
### vLLM
The upstream model card recommends [vLLM](https://github.com/vllm-project/vllm) 0.19.0 or newer for the Qwen3.6 architecture.
```bash
vllm serve Accio-Lab/Occamy-1.0 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
```
Both commands expose an OpenAI-compatible endpoint at `http://localhost:8000/v1`.
## 6. Model Usage
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Accio-Lab/Occamy-1.0",
messages=[
{
"role": "user",
"content": "Inspect this repository, fix the failing test, and explain the change.",
}
],
max_tokens=32768,
temperature=1.0,
top_p=0.95,
presence_penalty=1.5,
extra_body={
"top_k": 20,
"chat_template_kwargs": {
"enable_thinking": True,
"preserve_thinking": True,
},
},
)
print(response.choices[0].message.content)
```
For multi-turn agent runs, retain the complete assistant message returned by the server, including reasoning content and tool calls, then append tool results using the standard OpenAI chat-completions schema. This preserves the execution context that Occamy relies on across long workflows.
### Agent Frameworks
Occamy was trained and evaluated across multiple harnesses, including [OpenClaw](https://github.com/openclaw/openclaw), [Hermes Agent](https://github.com/NousResearch/hermes-agent), and Accio Work. It can be integrated with other tool-using agent frameworks through the same OpenAI-compatible API.
---
## 7. License
This repository is released under the [Apache License 2.0](LICENSE). See the Hugging Face model card for the terms that apply to the model weights.
---
## 8. Contact Us
For questions or feedback, please open an [issue](https://github.com/Accio-Lab/occamy/issues).