metadata
title: CyBench x Kimi K3 — Transcript Runs (Inspect View)
emoji: 🕵️
sdk: static
CyBench × Kimi K3 — all runs (Inspect View bundle)
Inspect View bundle of every CyBench run of Moonshot AI's Kimi K3 (2.8T MXFP4 MoE,
routed via OpenRouter) collected for the APAC transcript-risk study. One .eval
per run batch:
pilot_A/pilot_B— routing-config pilot (Moonshot-hosted reasoning-on vs Together reasoning-off)fullA_seedN_unguided/fullA_seedN_subtask— full benchmark passes (40 tasks × 3 seeds × both modes, paper config: 15/5 iterations, 6k/2k tokens)retry_seedN_*— retry passes for tasks whose environments needed the archived-distro repair (see dataset card)
Companion artifacts:
- Transcripts (native CyBench JSON):
ajay-citadel/cybench-kimi-k3-transcripts - Harness + patches provenance: see dataset card and PATCHES.md therein
Logs are added as passes complete; refresh to see new batches.