Buckets:

144 GB
7,520 files
Updated 23 minutes ago
Name
Size
code
processed
raw
stats
README.md11.6 kB
xet
STATUS.md2.7 kB
xet
README.md

SkillsStorage — agent-skills corpus (SKILL.md bundles)

Corpus of agent skills for training a skill-retrieval embedding model. Built from the "Skills Data Ledger" (2026-09-27). Bucket: hf://buckets/Mercity/SkillsStorage. Live progress: STATUS.md (regenerated every 30 min while jobs run).

A skill is a bundle, not a file: a folder holding SKILL.md plus any scripts/, references/, assets/, templates, evals, etc. Every source is normalised to two tables that join on skill_uid: skills (one row per bundle, with the SKILL.md text) and files (every other file of the bundle).

Layout

raw/                               untouched source data (re-process from here)
  hf/<owner>__<dataset>/           server-side copies of 20 Hugging Face datasets (GitSkills, ClawHub dump, other skill sets, benchmarks)
  vendor/                          31 official/curated GitHub repos, one tar.gz each (whole repo minus .git) + manifest.json
  clawhub_api/                     ClawHub public API: listing.jsonl (all public skills) + zips/ (bundles newer than the HF dump)
  github_delta/<run>/              GitHub crawl shards (one row per file) + repos.jsonl (per-repo outcome) + done.txt
seeds/ (local)                     repo seed lists used by the crawl (data/seeds in the code repo)
processed/
  gitskills/skills/                1,877,981 rows — every distinct SKILL.md content of the Jul-2026 GitHub census + repo metadata
  gitskills/files/                 5,858,945 rows — bundle files from GitSkills' artifact_siblings (text when GitSkills fetched it)
  gitskills/occurrences/           3,797,117 rows — every SKILL.md occurrence incl. verbatim copies (repo, path, file_sha)
  clawhub/skills.parquet           80,961 ClawHub skills (HF dump 2026-09-28) + ClawScan / VirusTotal / SkillSpector verdicts
  clawhub/files.parquet            291,125 bundle files of those skills
  clawhub/api_listing.parquet      81,082 public ClawHub skills: downloads, installs, stars, versions, topics
  clawhub/api_delta_*.parquet      1,659 skills created / re-versioned after the dump (+21,810 bundle files)
  github_delta/<run>/{skills,files}/   GitHub crawl, same layout; runs below
  canonical/                       one row per distinct SKILL.md content + near-dup clusters (see below)
  index/locator.parquet            skill_uid -> table file holding its rows (see "Rebuilding a skill folder")
code/                              every script that produced the above (see "Pipeline")

GitHub crawl runs (processed/github_delta/<run>):

run what repos
pass1_sourcegraph popular repos (Sourcegraph slice) with SKILL.md that are not in GitSkills 8,386 (complete)
pass2_new repos created since July 2026 (GitHub topic search) + skills.sh / claudemarketplaces / Anthropic-marketplace repos 29,296 (complete)
backfill_gitskills GitSkills repos with multi-file skills, re-fetched so their bundles are complete (GitSkills' own folder listings are truncated for 236k skills and ~1/3 of its sibling files lack text) 136,218 (in progress)

Totals (exact, 2026-10-02)

level count
SKILL.md bundle records collected (all sources, incl. copies / re-crawled versions) 4,440,161
distinct SKILL.md contents across all GitHub sources (git blob SHA) 2,500,256
… of which in GitSkills (Jul 2026) 1,842,600
… new since July (not in GitSkills): pass1 49,301 · pass2 158,777 · backfill 466,154 (new skills + edited versions in known repos) 657,656
ClawHub skills whose name+description is not on GitHub 37,122
distinct (name, description) pairs, all sources 1,863,211
bundle files (all sources, incl. GitSkills' partial listings) 21.55M

Computed by code/dedup_count.py → stats/dedup_count.json.

Canonical skills table + near-duplicate clusters (steps 1–3, 2026-10-02)

processed/canonical/skills/skills-{0..f}.parquet — 2,576,774 rows, one per distinct SKILL.md content (git blob SHA), merged across every source, sorted by sha (5,000-row row groups, so a lookup by sha reads one). Each row points at the most complete bundle for that content (crawl/backfill bundle with all files > GitSkills; then most stars; non-fork; freshest) — join canon_bundle_uid to the source files tables for the bundle's other files.

Key columns: sha, skill_md (full text), name, description, frontmatter_valid, body_chars, canon_source, canon_bundle_uid, canon_repo, canon_skill_dir, n_files, completeness, canon_stars, n_records, n_repos, sources, n_copies_jul (copies on GitHub in July), cjk_ratio, flags, train_ok. License columns (canon_license, license_tier, release_ok) are reference metadata only — everything here is open.

flag / tier rows
train_ok (name + description present, not stub, not CJK, not template spam, not security-flagged) 2,076,916
flag_cjk (CJK ratio > 10%; out of scope by decision) 253,831
flag_stub (body < 100 chars, description < 20 chars, or placeholder text) 224,618
flag_template_desc (one description reused by ≥50 contents with ≥20 different names) 22,648
flag_security (ClawHub: ClawScan malicious/suspicious, static malicious, or VirusTotal hits) 511
valid frontmatter 2,301,806

Canonical source of the chosen bundle: backfill 1,263,340 · GitSkills 1,039,506 · pass2 153,556 · ClawHub 77,727 · pass1 42,645.

processed/canonical/clusters.parquet — MinHash near-duplicate clusters (128 perms, word 5-gram shingles, LSH 16×8, verified Jaccard ≥ 0.8, union-find): sha, cluster_id (sha of the representative), cluster_size, is_rep (representative = most July copies, then stars, then completeness). 2,576,765 skills → 1,939,987 clusters (1,703,691 singletons; 8,644 clusters ≥ 10; largest 3,685). Bottom line (exact): 2,576,774 distinct skills → 2,076,916 pass the train_ok filters → those collapse into 1,576,972 near-duplicate clusters = unique, training-ready skills. (Non-CJK: 2,322,943 skills / 1,719,877 clusters.) See stats/final_unique.json.

Signatures: processed/canonical/minhash/ (re-cluster at another threshold without recomputing). Summaries: stats/canonical_summary.json, stats/near_dup_summary.json.

Schemas

skills (all sources): skill_uid, source, repo / owner+slug, skill_dir, skill_md (full SKILL.md text), name, description, frontmatter_valid, body_chars, cjk_ratio, bundle counts/bytes, plus source extras:

  • gitskills: file_sha (git blob SHA), location_class, has_scripts, has_references, sibling_count, composition_truncated, commit dates, repo_license, repo_stars, repo_forks, repo_is_fork, repo_language, …
  • github_delta: commit, file_sha (git blob SHA), n_bundle_files_with_content, fetched_at, via (git | codeload)
  • clawhub: version, license (MIT-0), scanner verdicts (clawscan_verdict, virustotal_*, skillspector_*, …)

files (all sources): skill_uid, rel_path (path inside the skill folder), size, blob_sha/sha256, content (text), content_bin (small binaries), has_content, skipped (why content is absent: bundle_cap >500 files or >8 MB per bundle, repo_cap >100 MB per repo, too_large >1 MB file, binary_large binary >256 KB, not_fetched_by_gitskills, …). Nothing is dropped silently: files without content are still listed with path, size and reason.

skill_uid: gh:{repo}:{skill_dir} (GitSkills, Jul 2026), gh:{repo}:{skill_dir}@{commit12} (our crawl), clawhub:{owner}/{slug}@{version}. A SKILL.md at the repo root has skill_dir = "." (bundle = whole repo, capped). file_sha/blob_sha are git blob SHAs in both GitSkills and our crawl, so exact copies dedup across sources.

Fill pass (fill-repo_cap-*.parquet in each run's files/): the 175,685 files first skipped with repo_cap (100 MB-per-repo memory guard) were fetched afterwards by exact git blob SHA (or by path at the crawled commit for codeload-path rows) — 175,685/175,685 recovered. Those rows have the same skill_uid + rel_path as the original repo_cap placeholder rows: when both exist, use the row with content.

Rebuilding a skill folder

A bundle = its skills row (SKILL.md text) + its files rows (each at rel_path inside the folder), joined on skill_uid.

python code/rebuild_skill.py <sha | skill_uid> [--out DIR] [--with-children]
  • sha (from processed/canonical/skills) rebuilds the canonical bundle for that SKILL.md content; a skill_uid rebuilds that exact bundle (repo + folder + commit / ClawHub version).
  • processed/index/locator.parquet (skill_uid, table, path, n_rows, sorted by skill_uid; built by code/build_locator.py) says which table file holds a bundle's rows, so a rebuild reads one or two small files.
  • Every written file is checked against its git blob SHA, or sha256 for ClawHub (byte-exact). Files listed without content (> 1 MB, binary

    256 KB, bundle caps, not fetched by GitSkills) go to _MISSING.tsv with the reason. When a path has both a repo_cap placeholder and a fill-pass row, the row with content wins.

  • Nested skills: a folder with its own SKILL.md inside another skill's folder is a separate bundle. In our crawl its files belong to the deepest SKILL.md folder only; GitSkills' listings of the parent may also include them. --with-children rebuilds the parent with every nested skill (same commit) in place.
  • Text is stored as UTF-8. Non-UTF-8 text files (rare) are exact only in the raw crawl shards (raw/github_delta/<run>/, content as bytes); the checker reports them as mismatches.
  • Bulk export (e.g. all cluster representatives): stream the files tables in order and group by skill_uid instead of rebuilding one bundle at a time.

Sources and decisions

  • Bucket A (no account, bulk): GitSkills, ClawHub dump + API delta, other HF skill sets (raw only; mostly overlap GitHub), benchmarks, vendor repos.
  • Bucket B (no account, crawl): GitHub passes above. Anonymous git is throttled at times; the crawler falls back to codeload tarballs (results verified identical to git, incl. blob SHAs).
  • Out of scope by decision: Chinese-only registries (Tencent SkillHub, ModelScope). CJK-heavy skills elsewhere are kept and flagged with cjk_ratio.

License notes (reference only — all data and models here are open)

  • GitHub content keeps its origin repo's license (~55% of repos have none): filter on repo_license / the repo.
  • anthropics/skills docx/pdf/pptx/xlsx skills are source-available, not OSS; trailofbits/skills is CC-BY-SA-4.0.
  • ClawHub is MIT-0 throughout. Scanner-flagged (malicious/suspicious) skills are kept but flagged.

Pipeline (code/)

All long jobs run as systemd services via run_job.sh (memory caps, auto-restart, resumable); watchdog.sh relaunches dead jobs and writes STATUS.md; memguard.sh freezes low-priority jobs when RAM is low. mirror_hf.py (bucket A copies) · vendor_repos.py · clawhub_api.py + build_clawhub_api.py · build_clawhub.py · build_gitskills.py · github_topic_seeds.py + build_pass2_seeds.py · github_bundles.py (crawler) · build_github_delta.py + ship_delta.py (convert → upload → delete) · crawl_chain.sh · status.py.

Secrets

The HF token lives only in a local, git-ignored .env. Nothing in this bucket contains it.

Total size
144 GB
Files
7,520
Last updated
Oct 9
Pre-warmed CDN
US EU US EU

Contributors