ajay-citadel's picture
Upload README.md with huggingface_hub
faa37b9 verified
|
Raw
History Blame
956 Bytes
metadata
title: CyBench x Kimi K3  Transcript Runs (Inspect View)
emoji: 🕵️
sdk: static

CyBench × Kimi K3 — all runs (Inspect View bundle)

Inspect View bundle of every CyBench run of Moonshot AI's Kimi K3 (2.8T MXFP4 MoE, routed via OpenRouter) collected for the APAC transcript-risk study. One .eval per run batch:

  • pilot_A / pilot_B — routing-config pilot (Moonshot-hosted reasoning-on vs Together reasoning-off)
  • fullA_seedN_unguided / fullA_seedN_subtask — full benchmark passes (40 tasks × 3 seeds × both modes, paper config: 15/5 iterations, 6k/2k tokens)
  • retry_seedN_* — retry passes for tasks whose environments needed the archived-distro repair (see dataset card)

Companion artifacts:

  • Transcripts (native CyBench JSON): ajay-citadel/cybench-kimi-k3-transcripts
  • Harness + patches provenance: see dataset card and PATCHES.md therein

Logs are added as passes complete; refresh to see new batches.