Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

📄 Paper  |  📚 arXiv  |  💻 GitHub

This repository contains a Qwen3.6-35B-A3B agent model trained with Skill2Env, a capability-oriented framework for synthesizing executable environments from skills.

TL;DR

Skills provide reusable domain knowledge, procedures, and tool-use instructions. Skill2Env turns these ingredients into complete executable environments for agent post-training.

  • Capability-oriented synthesis: reusable difficulty patterns target environment understanding, planning, skill usage, long-horizon consistency, and error recovery.
  • Blueprint-guided construction: task blueprints specify objectives, challenges, environment facts, information boundaries, and acceptance criteria, guiding the construction of execution substrates, workspaces, and rubric-based evaluators.
  • Iterative Task Hardening: agent rollouts reveal weaknesses in task design and guide coordinated updates to task blueprints and environments.
  • Supervised fine-tuning: high-scoring trajectories generated in Skill2Env environments provide supervision for agent training.

We construct 2,963 tasks and use 1.5K SFT trajectories to train Qwen3.6-35B-A3B.

Results

Skill2Env improves on its backbone by +8.4 points on average across seven agent benchmarks.

Main results on seven agent benchmarks:

Model Terminal-Bench 2.1 SWE-bench Multilingual SkillsBench Claw-Eval τ³-Banking AutomationBench VitaBench Avg.
Frontier Closed-Source Models
GPT-5.4 78.3 71.7 51.7 60.3 28.5 27.7 47.1 52.2
Claude Opus 4.6 71.2 77.8 50.2 70.4 20.3 25.5 38.3 50.5
Gemini-3.1 Pro 73.8 44.0 60.8 57.8 23.7 28.2 52.7 48.7
Open-Weight Models
DeepSeek-V4-Flash-0731 78.7 76.0 53.8 49.3 30.3 36.3 56.3 54.4
GLM-5.2 77.9 81.7 62.1 65.8 28.9 26.3 50.3 56.1
Kimi-K2.6 65.9 76.7 54.0 62.3 19.9 26.3 43.9 49.9
Qwen3.5-397B-A17B 51.3 66.0 36.5 56.8 16.2 5.5 42.1 39.2
Qwen3.6-35B-A3B 44.9 63.3 32.5 55.8 10.7 10.3 38.9 36.6
Qwen3.8-27B 79.8 73.8 35.6 69.2 33.7 38.5 41.8 53.2
Skill-Based Environment Synthesis
FACET-Terminal-Qwen3.5-27B 47.6 - - - - - - -
Skill2Env (ours / 35B-A3B) 58.4 71.0 46.9 61.3 15.1 17.7 44.8 45.0
Gain over baseline (+13.5) (+7.7) (+14.3) (+5.5) (+4.5) (+7.3) (+5.9) (+8.4)

Avg. is the unweighted mean across the seven benchmarks. Scores are rounded to one decimal place for display. Reported gains are computed from the unrounded scores before rounding.

Evaluation protocol. We use Terminus-2 for Terminal-Bench 2.1, mini-SWE-agent for SWE-bench Multilingual, OpenHands with skills enabled for SkillsBench, and the native agent loop for Claw-Eval. We report success rate for Terminal-Bench 2.1, resolved rate for SWE-bench Multilingual, Avg@3 for SkillsBench, Pass³ for Claw-Eval, Pass¹ for τ³-Banking, pass rate for AutomationBench (v1.0.6), and mean score for VitaBench. FACET results are source-reported.

Model Description

  • Base model: Qwen/Qwen3.6-35B-A3B (35B total parameters, 3B activated, MoE with vision encoder, 262,144-token context)
  • Training data: 1.5K SFT trajectories
  • Format: Hugging Face Transformers (compatible with Transformers, vLLM, SGLang, KTransformers, etc.)

Quickstart

vLLM

uv pip install vllm --torch-backend=auto

vllm serve AllSpark-Research/Skill2Env \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 262144 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3

SGLang

uv pip install sglang[all]

python -m sglang.launch_server --model-path AllSpark-Research/Skill2Env \
  --port 8000 --tp-size 8 --mem-fraction-static 0.8 \
  --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder

Transformers

from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(
    "AllSpark-Research/Skill2Env", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("AllSpark-Research/Skill2Env")

The model has a default context length of 262,144 tokens. We advise maintaining a context length of at least 128K tokens to preserve thinking capabilities.

License

This model is released under the Apache 2.0 license, inherited from the base model.

Citation

If you find this work useful, please cite:

@article{xu2026skill2env,
  title={Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents},
  author={Xu, Weiyi and Yang, Xiaowen and Da, Wen and Xu, Hang and Li, Canwei and You, Hongjie and Dong, Pusen and Zeng, Yucheng and Luo, Zhaokai and Chuan, Mu},
  journal={arXiv preprint arXiv:2609.33772},
  year={2026}
}
Downloads last month
-
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AllSpark-Research/Skill2Env

Finetuned
(330)
this model
Quantizations
1 model

Paper for AllSpark-Research/Skill2Env