--- license: mit language: - en library_name: transformers tags: - reinforcement-learning - post-training - distillation - agentic-coding - composer-2.5 - cursor - kimi-k2 - grpo - dapo - diloco - prime-rl - openenv - trl - verl - monarch - torchforge - research - methodology pretty_name: "Composer 2.5 Replication Framework β€” Research Synthesis" --- # Composer 2.5 Replication Framework > **Repo type:** `model` (methodology). **Status:** Research synthesis + v0.0 spike kickoff (2026-05-25). > **Author:** [Codeseys](https://huggingface.co/Codeseys) > **Goal:** Replicate Cursor's [Composer 2.5](https://cursor.com/blog/composer-2-5) (a post-trained Kimi K2.5 specialised for agentic coding) on **any** HuggingFace base model, using a synthesis of decentralized RL post-training techniques. This repository is the **"paper of the project"** β€” it is the methodology / research / framework specification for an open replication of Cursor's Composer 2.5 system, plus a **novel multi-teacher trace-replay distillation channel** that stacks on top of the Composer recipe. **v0.0 spike progress (2026-05-25):** - 🟒 Spike 001 (kill-switch teacher cost) β€” **VALIDATED**: 150 real OpenRouter calls, $0.98/trace, p95 latency 20.5s. The novel research direction is economically viable. - 🟒 Spike 005 (integrated 3-channel trainer skeleton) β€” **SKELETON-VALIDATED + COMPOSITION-VERIFIED**: 38/38 unit tests passing; the integration architecture claim ("all three channels run simultaneously, ablate cleanly, train without divergence") is empirically verified by 5-step training run on a tiny model. - πŸ“‹ Spikes 002a/002b/003/004 β€” planned, awaiting GPU budget commitment. See [`spikes/README.md`](spikes/README.md) for the 5-stage spike plan, [`docs/INTEGRATION_ARCHITECTURE.md`](docs/INTEGRATION_ARCHITECTURE.md) for the per-framework extension-point analysis, and [`spikes/005-integrated-trainer-skeleton/`](spikes/005-integrated-trainer-skeleton/) for runnable trainer code. --- ## TL;DR β€” what's in here, why it matters Cursor's Composer 2.5 is the strongest case study for "RL post-training of a frontier MoE base produces a model that beats GPT-5.5 on agentic coding while costing 5–10Γ— less to serve." The recipe is **almost entirely post-training** (~85% of compute) and the most important trick is **non-obvious**: a per-turn on-policy distillation loss called *Targeted RL with Textual Feedback*. This repo contains: 1. **`framework/composer-replication-framework.md`** β€” master synthesis: architecture, stack picks, phase plan, open questions. The TL;DR table maps every layer of the system to a concrete software pick with rationale. 2. **`research/01-composer-2.5.md`** β€” Composer 2.5 deep-dive: base model, 5-stage recipe, the secret-sauce hint-distillation loss, results. 3. **`research/02-diloco-family.md`** β€” DiLoCo / OpenDiLoCo / Streaming DiLoCo / PRIME-RL / INTELLECT-1+2 deep-dive: when decentralized training actually helps, when it's premature. 4. **`research/03-monarch-torchforge-openenv.md`** β€” Meta's Monarch actor mesh + TorchForge (paused) + OpenEnv environment standard. What's alive, what to bet on. 5. **`research/04-verl-trl.md`** β€” Algorithm-library deep-dive: GRPO / DAPO / DPO / PRM in TRL vs VeRL, plus the 3D-HybridEngine resharding pattern. 6. **`research/05-trace-replay-distillation.md`** β€” Novelty assessment of the trace-replay multi-teacher distillation idea: prior art (rStar / Math-Shepherd / OmegaPRM / Magpie / MoA), cost analysis, reward-shape options. Each of the five research deep-dives was authored by a **different LLM family** (Gemini 3.1 Pro, DeepSeek V4 Pro, GPT-5, Sonnet 4.6, Kimi K2-Thinking) running in parallel. The synthesis at `framework/composer-replication-framework.md` cross-checks their findings. ## Headline findings ### 1. Composer 2.5's secret sauce is the *targeted hint-distillation loss* β€” and it's published as SDPO The 1T MoE base (Kimi K2.5) and the "Feature Deletion" RL environment are the obvious moves. The non-obvious one β€” Cursor's "Targeted RL with Textual Feedback" β€” turns out to be **mathematically the same as the published SDPO method** (HΓΌbotter et al., ICLR 2026 Workshop, [arXiv:2601.20802](https://arxiv.org/abs/2601.20802) + [code at github.com/siyan-zhao/OPSD](https://github.com/siyan-zhao/OPSD), MIT licensed). Cursor cites this paper directly in the blog's footnote 1. > **The mechanism:** when a 100K-token rollout has a localized error, generate a text hint correcting the error, run forward pass *with* the hint to get "Teacher" logits, run forward pass *without* the hint to get "Student" logits, and apply KL divergence loss to pull Student toward Teacher *only at that turn*. **Same model is both teacher and student** β€” the teacher is just "the model with hint inserted into context." This sidesteps "GRPO on agentic traces is brittle because one bad step poisons 100 good ones." The biggest reproducibility gap is **how the text hints are generated** β€” Cursor never tells. v0.1 will use template-based hints first; v0.2 will explore LLM-driven hint generators. See [`docs/COMPOSER_RECIPE_MAPPING.md`](docs/COMPOSER_RECIPE_MAPPING.md) for the rigorous stage-by-stage mapping of Cursor's blog onto our framework. ### 2. The trace-replay multi-teacher idea is genuinely novel β€” and **economically viable** (verified) Closest precedent is rStar-Math (single-teacher MCTS counterfactuals at training time). **Multi-teacher *frozen-trace replay* with disagreement-as-reward is open territory.** **Spike 001 (βœ… VALIDATED, 2026-05-25)** measured the per-trace cost floor empirically: $0.98/trace ungated (vs. $5 cap), 5x headroom; with VOI gating in v0.1 we project ~$0.30/trace. **Critical distinction:** Composer's hint-distill (SDPO, single model with hint context) and trace-replay-distill (N external teachers) are **two different mechanisms**, not competing implementations. They stack: - **Composer hint-distill (SDPO)** = same-model self-teacher with hint context, pulls student at error sites. ~1 extra forward pass. No API cost. - **Trace-replay-distill** = N external pretrained teachers, pull student at all steps. ~$0.30/trace with VOI gating. Novel. v0.1 runs both. v0.0 (current) tests trace-replay alone vs. plain GRPO to falsify the novel claim cheaply. ### 3. Recommended stack (verified across all 5 reports) | Layer | Pick | Why not the alternative | |---|---|---| | **RL substrate** | [PRIME-RL](https://github.com/PrimeIntellect-ai/prime-rl) | INTELLECT-2 already proved 32B globally distributed; Forge is "development-paused" by Meta | | **Algorithm impl** | [TRL](https://github.com/huggingface/trl) (lift loss math) | Cleanest GRPO + first-class OpenEnv integration | | **Resharding pattern** | [VeRL](https://github.com/volcengine/verl)'s 3D-HybridEngine (reference) | Most battle-tested at 70B+ | | **Environments** | [OpenEnv](https://github.com/meta-pytorch/openenv) + [verifiers](https://github.com/willccbb/verifiers) | HF + Meta backing, MCP RFC landing, Hub-hosted | | **Distributed sync** | Skip DiLoCo for v0.1 | Outer loop only matters when training spans clusters | | **Orchestration** | Ray today, [Monarch](https://github.com/meta-pytorch/monarch) when mature | Forge paused; Monarch K8s story still landing | ## Architecture ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ OpenEnv Environment Hub β”‚ β”‚ (HF Hub, Docker images, MCP tool-calling)β”‚ β”‚ - Anyrun-style code sandbox β”‚ β”‚ - SWE-Gym, SWE-Bench-Verified envs β”‚ β”‚ - "Feature Deletion" auto-grader env β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ rollouts (verifiers protocol) β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ ORCHESTRATOR (CPU) β”‚ β”‚ - Schedules rollouts across inference workers β”‚ β”‚ - Assembles training batches β”‚ β”‚ - Routes hint-distillation pairs (Composer-style) β”‚ β”‚ - Routes trace-replay teacher queries (NOVEL) β”‚ β”‚ - Built on Monarch (future) or Ray (today) β”‚ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”˜ β”‚ rollout requests β”‚ training batches β”‚ teacher queries β–Ό β–Ό β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ INFERENCE POOL β”‚ β”‚ TRAINER (GPU) β”‚ β”‚ TEACHER POOL β”‚ β”‚ (vLLM / SGLang) β”‚ β”‚ - FSDP2 sharded β”‚ β”‚ - Frozen N teachers β”‚ β”‚ - Student policy β”‚ β”‚ - GRPO + DAPO β”‚ β”‚ - HF Inference, β”‚ β”‚ - Auto-resharded β”‚ β”‚ - +Hint distill β”‚ β”‚ OpenRouter, vLLM β”‚ β”‚ via SHARDCAST β”‚ β”‚ KL loss β”‚ β”‚ - Diverse families β”‚ β”‚ - Async tool waits β”‚ β”‚ - +PRM/DPO from β”‚ β”‚ (Anthropic / OpenAI β”‚ β”‚ don't block GPU β”‚ β”‚ trace-replay β”‚ β”‚ / DeepSeek / Qwen) β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ pseudo-gradients (every H steps) β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ OUTER LOOP (DiLoCo, optional) β”‚ β”‚ - Only when training spans β”‚ β”‚ multiple clusters / DCs β”‚ β”‚ - Streaming variant for β”‚ β”‚ bandwidth-limited links β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` Three reward channels feed the trainer: 1. **RLVR** β€” verifiable rewards (tests pass, build succeeds). Ground truth, never skipped. 2. **Composer hint-distill** β€” per-turn KL to a hint-conditioned forward pass. 3. **Trace-replay-distill** β€” per-step preference / process-reward signal from N frozen teachers. The novel contribution is channel (3) β€” no published work systematically replays each step of frozen agentic traces with multiple teachers to harvest step-level supervision. ## Roadmap | Phase | Timeline | Goal | Trained variant repo | Data repo | |---|---|---|---|---| | **v0.0 spike** | 1–2 weeks | Prove trace-replay-DPO beats plain GRPO on Qwen3-7B + SWE-bench-lite | `Codeseys/composer-replication-qwen3-7b-v0` | `Codeseys/composer-replication-traces-v0` | | **v0.1** | 1–2 months | Full Composer recipe (RLVR + hint-distill + trace-replay) on Qwen3-32B + Feature Deletion env. Match Cursor's ~50% SWE-bench-multilingual at 32B scale. | `Codeseys/composer-replication-qwen3-32b-v1` | `Codeseys/composer-replication-traces-v1` | | **v0.2** | 3–6 months | Decentralized scaling: Streaming DiLoCo + SHARDCAST + Monarch. Multi-cluster reproduction of v0.1. | `Codeseys/composer-replication-qwen3-32b-decentralized` | (re-uses v1 data) | Each variant will get its own model repo (LoRA adapters or full fine-tunes) per the [HF multi-artifact research project layout](https://huggingface.co/docs/hub/repositories). This methodology repo will be linked from each variant's README and via an HF Collection once v0.0 produces a result. ## Methodology β€” how this synthesis was produced To minimize single-model bias, the five research deep-dives were generated **in parallel** by five different LLM families via the [`delegate_task` parallel-research pattern](https://huggingface.co/docs/transformers/research): | Topic | Author model | |---|---| | `01-composer-2.5.md` | google/gemini-3.1-pro-preview | | `02-diloco-family.md` | deepseek/deepseek-v4-pro | | `03-monarch-torchforge-openenv.md` | openai/gpt-5 | | `04-verl-trl.md` | anthropic/claude-sonnet-4.6 | | `05-trace-replay-distillation.md` | moonshotai/kimi-k2-thinking | Convergent findings across reports (β‰₯2 independent confirmations): - **GRPO+DAPO is the consensus algorithm** (3/4 reports that compared) - **PRIME-RL is the most production-ready decentralized substrate** (2 reports independently) - **OpenEnv is the env-format winner** (3 reports converge) - **Trace-replay-with-N-teachers is genuinely under-explored** (the trace-replay report's primary finding, corroborated by the absence of it in the 4 other reports) The synthesis at `framework/composer-replication-framework.md` reconciles divergences (e.g., DiLoCo vs single-cluster timing) with explicit rationale. ## Citation If you use this framework or its derivative artifacts (the trained variants, the trace dataset, or the Feature-Deletion environment), please cite: ```bibtex @misc{composer-replication-framework-2026, author = {Codeseys}, title = {Composer 2.5 Replication Framework: A Methodology for Open Replication of Cursor's Agentic Coding Recipe}, year = {2026}, publisher = {HuggingFace}, howpublished = {\url{https://huggingface.co/Codeseys/composer-replication-framework}}, note = {Pre-spike research synthesis. Five-author parallel research with cross-family verification.} } ``` ## License MIT. Use freely; attribution appreciated. Underlying primary sources (Cursor blog, Moonshot K2.5 paper, DeepMind DiLoCo paper, Microsoft rStar paper, etc.) are owned by their respective authors and are cited inline in the research notes. ## Related work / links - [Cursor β€” Introducing Composer 2.5](https://cursor.com/blog/composer-2-5) (Cursor blog, 2026) - [Moonshot AI β€” Kimi K2 Thinking](https://huggingface.co/moonshotai/Kimi-K2-Thinking) - [Prime Intellect β€” PRIME-RL](https://github.com/PrimeIntellect-ai/prime-rl) and [INTELLECT-2 model card](https://huggingface.co/PrimeIntellect/INTELLECT-2) - [Hugging Face β€” TRL](https://github.com/huggingface/trl) - [ByteDance β€” VeRL](https://github.com/volcengine/verl) - [Meta β€” OpenEnv](https://github.com/meta-pytorch/openenv) + [Monarch](https://github.com/meta-pytorch/monarch) - [Microsoft β€” rStar / rStar-Math](https://github.com/microsoft/rStar) - [DeepMind β€” DiLoCo paper](https://arxiv.org/abs/2311.08105) and [Streaming DiLoCo](https://arxiv.org/abs/2501.18512) ## Contact Open a [Discussion](https://huggingface.co/Codeseys/composer-replication-framework/discussions) on this repo for technical questions, corrections, or collaboration interest. The five research notes are open to PRs β€” if you find a misattribution or a missing primary source, send a fix.