670 MB
6 files
Updated 2 days ago
Name
Size
data
.gitattributes3.4 kB
xet
README.md6.38 kB
xet
dataset-manifest.json83.1 kB
xet
moonshiner-dataset-banner.png1.58 MB
xet
traces.jsonl656 MB
xet
README.md

Claude Fable 5 Agent Traces

Moonshiner — Claude Fable 5 instruction following, tool use, and coding

2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL

Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces.

Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain.

This is an actively growing dataset. More is coming: additional training programs and substantially more sessions will be added to this same repo.

What makes it different

  • All real model trajectories. Every row is a cumulative prefix of a genuine Claude Fable 5 session captured end-to-end. The agent's causal exploration, tool arguments, results, corrections, and final responses are retained.

  • One next step per row. A trajectory with N assistant turns produces N rows. Row k contains the complete context through assistant turn k; that final assistant message is the sole training target.

  • Runtime-normalized. Runtime plumbing, UI decoration, control sequences, and verbose success boilerplate are removed or canonicalized while causal context remains.

  • Independently verified. Coding sessions must pass deterministic tests and protected-file checks. Instruction-following sessions must pass deterministic tool-call, staging, argument, and response-constraint checks. Every retained trajectory also clears independent review.

  • Reasoning-effort step-down. Failed trace attempts proceed through xhigh → medium → low (up to the configured attempt count) and stop at the first judge-accepted trace. If higher reasoning fails a task that lower reasoning succeeds on, the lower-effort trace is retained.

Task mix

High-level training programs, calculated from accepted trajectories using the same program mapping published in the seed catalog:

kind trajectories share row share flavor
Tool calling 640 29.6% 12.4% Select tools, construct grounded arguments, run independent calls together, and stage dependent calls.
Instruction following 545 25.2% 12.4% Honor constraints, formats, corrections, state, context, memory, relevance, and abstention.
Building 286 13.2% 14.0% Implement complete libraries, services, CLIs, workflows, and systems from specifications.
Debugging 230 10.6% 10.5% Diagnose failures, repair defects, resolve compiler/runtime issues, and preserve regressions.
Project & integration 138 6.4% 11.6% Coordinate multi-file, repository-scale, migration, and integration work.
Clarification 101 4.7% 1.7% Recognize missing information, ask only when required, and never invent parameters.
Feature development 81 3.7% 5.3% Extend working systems while preserving existing behavior.
Error recovery 70 3.2% 1.9% Recover from tool failures, partial results, retries, and idempotency hazards.
Seed authoring 49 2.3% 29.1% Accepted trajectories grouped by their explicit published dataset category.
Refactoring & performance 21 1.0% 1.0% Restructure safely and improve measured performance without behavior drift.

Languages (current drop)

English, Python, TypeScript, Go, PowerShell, Java, Rust, C#, Ruby, Bash, Zsh, JavaScript, C, C++

Schema

Each row:

column type contents
task string stable task id
lang string English (en) or primary programming language
category string detailed recipe category
split string trajectory-disjoint train or val partition
assistant_step int 1-based target assistant turn
assistant_steps int assistant turns in the source trajectory
target_message_index int index of the final assistant target
n_messages int cumulative message count through the target
messages list of objects cumulative context ending at the target

messages is native JSON.

Layout

The Hugging Face viewer reads the validated active Parquet shards listed in dataset-manifest.json. The equivalent canonical traces.jsonl is also published for direct download and conversion. It currently contains 11,235 cumulative next-step rows derived from 2,161 accepted trajectories over disjoint train and validation tasks.

When training from the cumulative view, supervise only the final assistant message in each row. Supervising every assistant span would repeatedly overweight early steps because those spans recur as context in later prefixes.

Intended use

Supervised fine-tuning of instruction-following, tool-calling, and coding agents, plus analysis of multi-step planning, parallel calls, tool selection, state tracking, build-test-fix loops, and verification-driven completion.

Provenance

Generated with Claude Fable 5 (anthropic/claude-fable-5), then filtered by deterministic verification and explicit Codex acceptance review. Codex supplies no demonstration content; it only judges or requests replacement of candidate traces. Provider credentials, user keys, and host-identifying data are scrubbed before publication.

License

CC BY 4.0 — free for training, research, commercial products, modification, redistribution, and inclusion in other datasets or corpora, with attribution.

Suggested attribution:

Claude Fable 5 Agent Traces — https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces

Total size
670 MB
Files
6
Last updated
Jul 26
Pre-warmed CDN
US EU US EU

Contributors