The contamination-free part comes from the pretrain, not from no_robots.
no_robots is clean, 9500 train and 500 test, all human-written. But the scarce ingredient is huggyllama/llama-7b underneath it. Llama 1's corpus was scraped before the web filled up with model output. You can always hand-write another 10k demos. You cannot go re-scrape 2022.
Which makes this the right instrument for what we were going back and forth on earlier this month. Helpful-only, no RLHF, means it was never trained to say "as an AI I can't know that." There is no assistant persona to recite. So any self-report it gives has to be a live read of its own state, not a memorized fact about itself.
That is the clean version of the test. Does it carry any calibration signal at all when nothing ever taught it to perform one?
One heads-up on the card: no_robots is cc-by-nc-4.0, and right now the card has no license field. Worth setting before people build on it.
Have you pointed it at its own errors yet, or is 190726 still just the base?