--- title: DarwinX emoji: ๐งฌ colorFrom: blue colorTo: indigo sdk: static app_file: index.html pinned: false license: cc-by-4.0 short_description: Evolving Agent Harnesses Through Natural Selection --- # DarwinX: Evolving Agent Harnesses Through Natural Selection Project page for [arXiv:2608.07545](https://arxiv.org/abs/2608.07545). DarwinX treats agent self-evolution as selection over a *population* of harnesses with the base model frozen. Across four benchmarks, one loop adds ~17 points on average: | Benchmark | Result | Gain | What it isolates | | --- | --- | --- | --- | | Terminal-Bench 2.1 (avg@5) | 84.7% | +7.7 on matched base | in-domain evolution | | TerminalWorld (held-out) | 68.3% | +7.3 | disjoint held-out task split | | WebArena-Infinity (audit-clean pass@1) | 93.0% | +49.5 pp | synthetic โ real intent shift | | SWE-bench Verified (zero-shot transfer) | 84.2% | +3.4 | cross-benchmark transfer | ## Page sections `Abstract` ยท `Method` ยท `Results at a glance` ยท `Benchmark detail` ยท `Ablation: what evolution changes` ยท `Limitations` ยท `BibTeX` The ablation section reproduces the paper's framing: it is an *exploratory attribution, not a per-skill causal ablation*, since the skills were co-selected rather than independently randomized. The limitations section is carried over from the paper's Discussion. ## Interactive figures Eight figures are interactive, and each is driven by real run artifacts or the paper's own plotting scripts rather than illustrative numbers. `tools/build_data.py` regenerates `assets/data.js` from the original files, so nothing is hand-transcribed. **1. TerminalWorld merge explorer.** Toggle any subset of the four evolved specialists and watch which of the 41 held-out tasks the selection covers. Built from the `per_task_results` arrays of the five evaluation runs, joined on `task_id` (the run files list the tasks in different orders, so position-based joining would silently misalign them). It surfaces something the paper states only in aggregate. The four specialists solve 24/25/26/27 tasks and their **union is 29**, but the harness recombination actually produced solves **28**: | | tasks | | --- | --- | | best single specialist (D) | 27/41 | | union of all four | 29/41 | | realized merge | 28/41 | | solved by merge, by no specialist | `tw_448247` | | solved by some specialist, not by merge | `tw_449421`, `tw_498533` | | solved by all four | 21/41 | | solved by none | 12/41 | So recombination is not a free set union: it adds a capability no parent had and loses two. These are single attempts and one task is worth 2.4 points, so individual flips sit inside the noise band โ the page says so next to the widget. **2. TB2.1 cluster explorer.** Per-cluster base vs. evolved avg@5, sortable by gain, base rate, cluster size, or name, with exact rates on hover. Numbers from `notes/TB21_RESULTS.md` (paired protocol, 88 tasks). Deltas are carried over as reported rather than recomputed, because the source rounds them from unrounded rates. **3. WebArena-Infinity evolution curve.** All 37 evaluated variants with a best-so-far envelope, from `notes/tw_dynamics.json` โ `wai_adaptive_scores`. Toggle either series. **4. Headline four-benchmark panels.** Replaces the right half of the teaser. From `scripts/gen_summary_figure.py`. The reason to make this one interactive is that the paper's figure has to *state* its truncation in the caption: on a shared axis WebArena-Infinity's +49.5 flattens the two terminal benchmarks into slivers, so each panel is scaled to its own range. Here the reader can switch between per-panel and shared 0โ100 axes and see both framings. The best-prior-agent bar can be hidden, since those systems use different models and effort settings and are context rather than a controlled comparison. Per-panel labels carry the exceptions: on SWE-bench Verified the grey bar is the fix-skill reference, not an unevolved Monet, and there is no prior-agent bar. **5. TerminalWorld specialist bars.** From `notes/tw_dynamics.json` โ `tw_heldout`. Same axis question in miniature: the six arms span 58.5โ68.3, so a 0โ100 track renders them near-identical. Defaults to a 55โ70 axis with the full axis one click away and the truncation named in both notes. Spec. D and the Claude Code reference are both 27/41, so D's bar landing exactly on the dashed reference line is a built-in check that the axis transform is right. **6. TB2.1 compute.** From `scripts/gen_tb21_compute.py` (medians over clean attempts). Switch between turns and tokens; the bars share one scale across both task groups so the newly-solved vs. already-solved contrast is not rescaled away. **7. WAI invalid-trajectory composition.** From `scripts/gen_wai_invalid_composition.py`. Break the 293 โ 17 collapse down by application or by mechanism. Both rows sit on a shared absolute scale by default, which makes the "after" row a near-invisible sliver โ that *is* the finding; normalizing each row to its own total then shows what the remainder consists of. `build_data.py` asserts both decompositions still total 293 and 17. **8. WAI audit dumbbells.** From `scripts/gen_wai_audit_by_app.py`. One line per application from raw pass@1 to audited pass@1, so line length reads directly as how much a harness was leaning on trajectories the audit rejects. Sortable by audit loss. `build_data.py` cross-checks the audited DarwinX column against the per-application table rendered elsewhere on the page. Three figures stay static because they are conceptual diagrams with no underlying data: the teaser's selection schematic, the method overview, and the per-generation operators. The archive lineage tree also stays static for a different reason โ its node/edge data lives in a `state.db` on the cluster, not on this machine, so there is nothing truthful to make interactive yet. Tables are click-to-sort. Numeric columns open descending, text columns AโZ, and `Overall` rows stay pinned to the bottom. ## Local preview ```bash python3 -m http.server 8000 # open http://localhost:8000 ``` ## Regenerate and test ```bash python3 tools/build_data.py # rebuild assets/data.js from the run artifacts # then, with the server running, open: # http://localhost:8000/tools/interaction_test.html # it drives every widget with real click events and writes PASS/FAIL into the page title ``` Headless run: ```bash python3 -m http.server 8000 & "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" --headless --disable-gpu \ --virtual-time-budget=6000 --dump-dom http://localhost:8000/tools/interaction_test.html \ | grep -o '