--- title: DarwinX emoji: ๐Ÿงฌ colorFrom: blue colorTo: indigo sdk: static app_file: index.html pinned: false license: cc-by-4.0 short_description: Evolving Agent Harnesses Through Natural Selection --- # DarwinX: Evolving Agent Harnesses Through Natural Selection Project page for [arXiv:2608.07545](https://arxiv.org/abs/2608.07545). DarwinX treats agent self-evolution as selection over a *population* of harnesses with the base model frozen. Across four benchmarks, one loop adds ~17 points on average: | Benchmark | Result | Gain | What it isolates | | --- | --- | --- | --- | | Terminal-Bench 2.1 (avg@5) | 84.7% | +7.7 on matched base | in-domain evolution | | TerminalWorld (held-out) | 68.3% | +7.3 | disjoint held-out task split | | WebArena-Infinity (audit-clean pass@1) | 93.0% | +49.5 pp | synthetic โ†’ real intent shift | | SWE-bench Verified (zero-shot transfer) | 84.2% | +3.4 | cross-benchmark transfer | ## Page sections `Abstract` ยท `Method` ยท `Results at a glance` ยท `Benchmark detail` ยท `Ablation: what evolution changes` ยท `Limitations` ยท `BibTeX` The ablation section reproduces the paper's framing: it is an *exploratory attribution, not a per-skill causal ablation*, since the skills were co-selected rather than independently randomized. The limitations section is carried over from the paper's Discussion. ## Interactive figures Eight figures are interactive, and each is driven by real run artifacts or the paper's own plotting scripts rather than illustrative numbers. `tools/build_data.py` regenerates `assets/data.js` from the original files, so nothing is hand-transcribed. **1. TerminalWorld merge explorer.** Toggle any subset of the four evolved specialists and watch which of the 41 held-out tasks the selection covers. Built from the `per_task_results` arrays of the five evaluation runs, joined on `task_id` (the run files list the tasks in different orders, so position-based joining would silently misalign them). It surfaces something the paper states only in aggregate. The four specialists solve 24/25/26/27 tasks and their **union is 29**, but the harness recombination actually produced solves **28**: | | tasks | | --- | --- | | best single specialist (D) | 27/41 | | union of all four | 29/41 | | realized merge | 28/41 | | solved by merge, by no specialist | `tw_448247` | | solved by some specialist, not by merge | `tw_449421`, `tw_498533` | | solved by all four | 21/41 | | solved by none | 12/41 | So recombination is not a free set union: it adds a capability no parent had and loses two. These are single attempts and one task is worth 2.4 points, so individual flips sit inside the noise band โ€” the page says so next to the widget. **2. TB2.1 cluster explorer.** Per-cluster base vs. evolved avg@5, sortable by gain, base rate, cluster size, or name, with exact rates on hover. Numbers from `notes/TB21_RESULTS.md` (paired protocol, 88 tasks). Deltas are carried over as reported rather than recomputed, because the source rounds them from unrounded rates. **3. WebArena-Infinity evolution curve.** All 37 evaluated variants with a best-so-far envelope, from `notes/tw_dynamics.json` โ†’ `wai_adaptive_scores`. Toggle either series. **4. Headline four-benchmark panels.** Replaces the right half of the teaser. From `scripts/gen_summary_figure.py`. The reason to make this one interactive is that the paper's figure has to *state* its truncation in the caption: on a shared axis WebArena-Infinity's +49.5 flattens the two terminal benchmarks into slivers, so each panel is scaled to its own range. Here the reader can switch between per-panel and shared 0โ€“100 axes and see both framings. The best-prior-agent bar can be hidden, since those systems use different models and effort settings and are context rather than a controlled comparison. Per-panel labels carry the exceptions: on SWE-bench Verified the grey bar is the fix-skill reference, not an unevolved Monet, and there is no prior-agent bar. **5. TerminalWorld specialist bars.** From `notes/tw_dynamics.json` โ†’ `tw_heldout`. Same axis question in miniature: the six arms span 58.5โ€“68.3, so a 0โ€“100 track renders them near-identical. Defaults to a 55โ€“70 axis with the full axis one click away and the truncation named in both notes. Spec. D and the Claude Code reference are both 27/41, so D's bar landing exactly on the dashed reference line is a built-in check that the axis transform is right. **6. TB2.1 compute.** From `scripts/gen_tb21_compute.py` (medians over clean attempts). Switch between turns and tokens; the bars share one scale across both task groups so the newly-solved vs. already-solved contrast is not rescaled away. **7. WAI invalid-trajectory composition.** From `scripts/gen_wai_invalid_composition.py`. Break the 293 โ†’ 17 collapse down by application or by mechanism. Both rows sit on a shared absolute scale by default, which makes the "after" row a near-invisible sliver โ€” that *is* the finding; normalizing each row to its own total then shows what the remainder consists of. `build_data.py` asserts both decompositions still total 293 and 17. **8. WAI audit dumbbells.** From `scripts/gen_wai_audit_by_app.py`. One line per application from raw pass@1 to audited pass@1, so line length reads directly as how much a harness was leaning on trajectories the audit rejects. Sortable by audit loss. `build_data.py` cross-checks the audited DarwinX column against the per-application table rendered elsewhere on the page. Three figures stay static because they are conceptual diagrams with no underlying data: the teaser's selection schematic, the method overview, and the per-generation operators. The archive lineage tree also stays static for a different reason โ€” its node/edge data lives in a `state.db` on the cluster, not on this machine, so there is nothing truthful to make interactive yet. Tables are click-to-sort. Numeric columns open descending, text columns Aโ€“Z, and `Overall` rows stay pinned to the bottom. ## Local preview ```bash python3 -m http.server 8000 # open http://localhost:8000 ``` ## Regenerate and test ```bash python3 tools/build_data.py # rebuild assets/data.js from the run artifacts # then, with the server running, open: # http://localhost:8000/tools/interaction_test.html # it drives every widget with real click events and writes PASS/FAIL into the page title ``` Headless run: ```bash python3 -m http.server 8000 & "/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" --headless --disable-gpu \ --virtual-time-budget=6000 --dump-dom http://localhost:8000/tools/interaction_test.html \ | grep -o '.*' ``` ## Layout ``` index.html the page: inline CSS + vanilla JS, no build step, no external requests, no web fonts, no chart library assets/data.js generated โ€” real numbers behind the interactive figures assets/*.png web-sized figures (max 2000px wide, 256-color) โ€” 736 KB total assets/full/*.png full-resolution originals for the three dense diagrams, linked from the page via click-to-zoom tools/build_data.py regenerates assets/data.js from the original run artifacts tools/interaction_test.html click-driven test suite for the interactive figures (49 checks) ``` | Figure | Content | | --- | --- | | `teaser.png` | four-benchmark teaser (full figure, click-to-zoom target) | | `teaser_schematic.png` | teaser's left half โ€” the selection schematic shown on the page | | `overview.png` | method overview | | `population.png` | per-generation operators (mutation, classification, merge) | | `tb21_category_gain.png` | per-cluster TB2.1 gain | | `tb21_compute.png` | per-task compute: newly-solved vs already-solved | | `tw_heldout_specialists.png` | held-out generalization vs archive diversity | | `wai_evolution_curve.png` | best-so-far screening score | | `wai_tree.png` | archive lineage tree | | `wai_invalid_composition.png` | invalid trajectories, 293 โ†’ 17 | | `wai_audit_by_app.png` | raw vs audited pass@1 per application | Result figures are rendered from the paper's PDF sources at 200 DPI, then palette-reduced without dithering (RMSE vs. truecolor < 0.5%, which is imperceptible for line art but 4.7x smaller). ## Hosting The frontmatter above configures this directory as a Hugging Face **static Space**. The same files work unchanged as a GitHub Pages site or behind any static file server. One thing to change when hosting: `og:image` and `twitter:image` are relative paths, which most crawlers resolve but X/Twitter wants absolute. Replace them with the deployed absolute URL to get a large social card.