Spaces:
Running
Running
metadata
license: other
license_name: proprietary
license_link: https://github.com/szl-holdings/a11oy/blob/main/LICENSE
tags:
- benchmark
- receipts
- governance
- mathcomp
- mirror-not-canonical
pretty_name: A11oy staged test-results schema
A11oy test-results — staged dataset schema
This directory defines the future SZLHOLDINGS/a11oy-test-results dataset
layout. It is a schema and manifest only in this revision.
GitHub remains canonical. Hugging Face is a generated mirror for review.
Current claim status
- No live benchmark score is claimed.
- No leaderboard metric is claimed.
- No benchmark corpus is redistributed here.
- No model-index metrics are published.
Competition-math benchmark scoring remains staged until corpus digest, receipts, reproducible tooling, and judge agreement are present.
Future dataset layout
README.md
MANIFEST.json
benchmark-map.json
schemas/manifest.schema.json
schemas/result-row.schema.json
schemas/receipt-envelope.schema.json
samples/staged/*.jsonl
results/*.jsonl
receipts/*.jsonl
Only schema examples or receipt-backed staged dry-run artifacts may appear before a sealed run exists. Real results require:
- immutable corpus digest;
- raw-score reporting;
- three-judge panel;
- append-only receipt chain;
- unsupported-claim rejection;
- GitHub CI validation.
Validate the current staged manifest with:
npm run hf:test-results:audit
npm run benchmark:audit