Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| code | 32 items | ||
| processed | 3,786 items | ||
| raw | 3,696 items | ||
| stats | 4 items | ||
| README.md | 11.6 kB xet | 97e0df67 | |
| STATUS.md | 2.7 kB xet | ccf06890 |
SkillsStorage — agent-skills corpus (SKILL.md bundles)
Corpus of agent skills for training a skill-retrieval embedding model. Built from the
"Skills Data Ledger" (2026-09-27). Bucket: hf://buckets/Mercity/SkillsStorage.
Live progress: STATUS.md (regenerated every 30 min while jobs run).
A skill is a bundle, not a file: a folder holding SKILL.md plus any scripts/, references/,
assets/, templates, evals, etc. Every source is normalised to two tables that join on skill_uid:
skills (one row per bundle, with the SKILL.md text) and files (every other file of the bundle).
Layout
raw/ untouched source data (re-process from here)
hf/<owner>__<dataset>/ server-side copies of 20 Hugging Face datasets (GitSkills, ClawHub dump, other skill sets, benchmarks)
vendor/ 31 official/curated GitHub repos, one tar.gz each (whole repo minus .git) + manifest.json
clawhub_api/ ClawHub public API: listing.jsonl (all public skills) + zips/ (bundles newer than the HF dump)
github_delta/<run>/ GitHub crawl shards (one row per file) + repos.jsonl (per-repo outcome) + done.txt
seeds/ (local) repo seed lists used by the crawl (data/seeds in the code repo)
processed/
gitskills/skills/ 1,877,981 rows — every distinct SKILL.md content of the Jul-2026 GitHub census + repo metadata
gitskills/files/ 5,858,945 rows — bundle files from GitSkills' artifact_siblings (text when GitSkills fetched it)
gitskills/occurrences/ 3,797,117 rows — every SKILL.md occurrence incl. verbatim copies (repo, path, file_sha)
clawhub/skills.parquet 80,961 ClawHub skills (HF dump 2026-09-28) + ClawScan / VirusTotal / SkillSpector verdicts
clawhub/files.parquet 291,125 bundle files of those skills
clawhub/api_listing.parquet 81,082 public ClawHub skills: downloads, installs, stars, versions, topics
clawhub/api_delta_*.parquet 1,659 skills created / re-versioned after the dump (+21,810 bundle files)
github_delta/<run>/{skills,files}/ GitHub crawl, same layout; runs below
canonical/ one row per distinct SKILL.md content + near-dup clusters (see below)
index/locator.parquet skill_uid -> table file holding its rows (see "Rebuilding a skill folder")
code/ every script that produced the above (see "Pipeline")
GitHub crawl runs (processed/github_delta/<run>):
| run | what | repos |
|---|---|---|
pass1_sourcegraph |
popular repos (Sourcegraph slice) with SKILL.md that are not in GitSkills | 8,386 (complete) |
pass2_new |
repos created since July 2026 (GitHub topic search) + skills.sh / claudemarketplaces / Anthropic-marketplace repos | 29,296 (complete) |
backfill_gitskills |
GitSkills repos with multi-file skills, re-fetched so their bundles are complete (GitSkills' own folder listings are truncated for 236k skills and ~1/3 of its sibling files lack text) | 136,218 (in progress) |
Totals (exact, 2026-10-02)
| level | count |
|---|---|
| SKILL.md bundle records collected (all sources, incl. copies / re-crawled versions) | 4,440,161 |
| distinct SKILL.md contents across all GitHub sources (git blob SHA) | 2,500,256 |
| … of which in GitSkills (Jul 2026) | 1,842,600 |
| … new since July (not in GitSkills): pass1 49,301 · pass2 158,777 · backfill 466,154 (new skills + edited versions in known repos) | 657,656 |
| ClawHub skills whose name+description is not on GitHub | 37,122 |
| distinct (name, description) pairs, all sources | 1,863,211 |
| bundle files (all sources, incl. GitSkills' partial listings) | 21.55M |
Computed by code/dedup_count.py → stats/dedup_count.json.
Canonical skills table + near-duplicate clusters (steps 1–3, 2026-10-02)
processed/canonical/skills/skills-{0..f}.parquet — 2,576,774 rows, one per distinct SKILL.md content (git blob SHA),
merged across every source, sorted by sha (5,000-row row groups, so a lookup by sha reads one). Each row points at the most complete bundle for that content (crawl/backfill bundle with
all files > GitSkills; then most stars; non-fork; freshest) — join canon_bundle_uid to the source files tables for
the bundle's other files.
Key columns: sha, skill_md (full text), name, description, frontmatter_valid, body_chars, canon_source,
canon_bundle_uid, canon_repo, canon_skill_dir, n_files, completeness, canon_stars, n_records, n_repos,
sources, n_copies_jul (copies on GitHub in July), cjk_ratio, flags, train_ok.
License columns (canon_license, license_tier, release_ok) are reference metadata only — everything here is open.
| flag / tier | rows |
|---|---|
train_ok (name + description present, not stub, not CJK, not template spam, not security-flagged) |
2,076,916 |
flag_cjk (CJK ratio > 10%; out of scope by decision) |
253,831 |
flag_stub (body < 100 chars, description < 20 chars, or placeholder text) |
224,618 |
flag_template_desc (one description reused by ≥50 contents with ≥20 different names) |
22,648 |
flag_security (ClawHub: ClawScan malicious/suspicious, static malicious, or VirusTotal hits) |
511 |
| valid frontmatter | 2,301,806 |
Canonical source of the chosen bundle: backfill 1,263,340 · GitSkills 1,039,506 · pass2 153,556 · ClawHub 77,727 · pass1 42,645.
processed/canonical/clusters.parquet — MinHash near-duplicate clusters (128 perms, word 5-gram shingles,
LSH 16×8, verified Jaccard ≥ 0.8, union-find): sha, cluster_id (sha of the representative), cluster_size, is_rep
(representative = most July copies, then stars, then completeness). 2,576,765 skills → 1,939,987 clusters
(1,703,691 singletons; 8,644 clusters ≥ 10; largest 3,685). Bottom line (exact): 2,576,774 distinct skills → 2,076,916 pass the train_ok filters → those collapse into
1,576,972 near-duplicate clusters = unique, training-ready skills. (Non-CJK: 2,322,943 skills / 1,719,877 clusters.)
See stats/final_unique.json.
Signatures: processed/canonical/minhash/ (re-cluster at
another threshold without recomputing). Summaries: stats/canonical_summary.json, stats/near_dup_summary.json.
Schemas
skills (all sources): skill_uid, source, repo / owner+slug, skill_dir, skill_md (full SKILL.md text),
name, description, frontmatter_valid, body_chars, cjk_ratio, bundle counts/bytes, plus source extras:
- gitskills:
file_sha(git blob SHA),location_class,has_scripts,has_references,sibling_count,composition_truncated, commit dates,repo_license,repo_stars,repo_forks,repo_is_fork,repo_language, … - github_delta:
commit,file_sha(git blob SHA),n_bundle_files_with_content,fetched_at,via(git | codeload) - clawhub:
version,license(MIT-0), scanner verdicts (clawscan_verdict,virustotal_*,skillspector_*, …)
files (all sources): skill_uid, rel_path (path inside the skill folder), size, blob_sha/sha256, content (text),
content_bin (small binaries), has_content, skipped (why content is absent: bundle_cap >500 files or >8 MB per bundle,
repo_cap >100 MB per repo, too_large >1 MB file, binary_large binary >256 KB, not_fetched_by_gitskills, …).
Nothing is dropped silently: files without content are still listed with path, size and reason.
skill_uid: gh:{repo}:{skill_dir} (GitSkills, Jul 2026), gh:{repo}:{skill_dir}@{commit12} (our crawl),
clawhub:{owner}/{slug}@{version}. A SKILL.md at the repo root has skill_dir = "." (bundle = whole repo, capped).
file_sha/blob_sha are git blob SHAs in both GitSkills and our crawl, so exact copies dedup across sources.
Fill pass (fill-repo_cap-*.parquet in each run's files/): the 175,685 files first skipped with repo_cap
(100 MB-per-repo memory guard) were fetched afterwards by exact git blob SHA (or by path at the crawled commit for
codeload-path rows) — 175,685/175,685 recovered. Those rows have the same skill_uid + rel_path as the original
repo_cap placeholder rows: when both exist, use the row with content.
Rebuilding a skill folder
A bundle = its skills row (SKILL.md text) + its files rows (each at rel_path inside the folder), joined on skill_uid.
python code/rebuild_skill.py <sha | skill_uid> [--out DIR] [--with-children]
sha(fromprocessed/canonical/skills) rebuilds the canonical bundle for that SKILL.md content; askill_uidrebuilds that exact bundle (repo + folder + commit / ClawHub version).processed/index/locator.parquet(skill_uid,table,path,n_rows, sorted byskill_uid; built bycode/build_locator.py) says which table file holds a bundle's rows, so a rebuild reads one or two small files.- Every written file is checked against its git blob SHA, or sha256 for ClawHub (byte-exact). Files listed without content (> 1 MB, binary
256 KB, bundle caps, not fetched by GitSkills) go to
_MISSING.tsvwith the reason. When a path has both arepo_capplaceholder and a fill-pass row, the row with content wins. - Nested skills: a folder with its own SKILL.md inside another skill's folder is a separate bundle. In our crawl its
files belong to the deepest SKILL.md folder only; GitSkills' listings of the parent may also include them.
--with-childrenrebuilds the parent with every nested skill (same commit) in place. - Text is stored as UTF-8. Non-UTF-8 text files (rare) are exact only in the raw crawl shards
(
raw/github_delta/<run>/,contentas bytes); the checker reports them as mismatches. - Bulk export (e.g. all cluster representatives): stream the
filestables in order and group byskill_uidinstead of rebuilding one bundle at a time.
Sources and decisions
- Bucket A (no account, bulk): GitSkills, ClawHub dump + API delta, other HF skill sets (raw only; mostly overlap GitHub), benchmarks, vendor repos.
- Bucket B (no account, crawl): GitHub passes above. Anonymous git is throttled at times; the crawler falls back to codeload tarballs (results verified identical to git, incl. blob SHAs).
- Out of scope by decision: Chinese-only registries (Tencent SkillHub, ModelScope). CJK-heavy skills elsewhere are kept and flagged with
cjk_ratio.
License notes (reference only — all data and models here are open)
- GitHub content keeps its origin repo's license (~55% of repos have none): filter on
repo_license/ the repo. anthropics/skillsdocx/pdf/pptx/xlsx skills are source-available, not OSS;trailofbits/skillsis CC-BY-SA-4.0.- ClawHub is MIT-0 throughout. Scanner-flagged (malicious/suspicious) skills are kept but flagged.
Pipeline (code/)
All long jobs run as systemd services via run_job.sh (memory caps, auto-restart, resumable); watchdog.sh relaunches
dead jobs and writes STATUS.md; memguard.sh freezes low-priority jobs when RAM is low.
mirror_hf.py (bucket A copies) · vendor_repos.py · clawhub_api.py + build_clawhub_api.py · build_clawhub.py ·
build_gitskills.py · github_topic_seeds.py + build_pass2_seeds.py · github_bundles.py (crawler) ·
build_github_delta.py + ship_delta.py (convert → upload → delete) · crawl_chain.sh · status.py.
Secrets
The HF token lives only in a local, git-ignored .env. Nothing in this bucket contains it.
- Total size
- 144 GB
- Files
- 7,520
- Last updated
- Oct 9
- Pre-warmed CDN
- US EU US EU