Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,10 +1,24 @@
|
|
| 1 |
---
|
| 2 |
-
title:
|
| 3 |
-
emoji:
|
| 4 |
-
colorFrom: green
|
| 5 |
-
colorTo: indigo
|
| 6 |
sdk: static
|
| 7 |
-
pinned: false
|
| 8 |
---
|
| 9 |
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
title: CyBench x Kimi K3 — Transcript Runs (Inspect View)
|
| 3 |
+
emoji: 🕵️
|
|
|
|
|
|
|
| 4 |
sdk: static
|
|
|
|
| 5 |
---
|
| 6 |
|
| 7 |
+
# CyBench × Kimi K3 — all runs (Inspect View bundle)
|
| 8 |
+
|
| 9 |
+
Inspect View bundle of every CyBench run of Moonshot AI's Kimi K3 (2.8T MXFP4 MoE,
|
| 10 |
+
routed via OpenRouter) collected for the APAC transcript-risk study. One `.eval`
|
| 11 |
+
per run batch:
|
| 12 |
+
|
| 13 |
+
- `pilot_A` / `pilot_B` — routing-config pilot (Moonshot-hosted reasoning-on vs
|
| 14 |
+
Together reasoning-off)
|
| 15 |
+
- `fullA_seedN_unguided` / `fullA_seedN_subtask` — full benchmark passes
|
| 16 |
+
(40 tasks × 3 seeds × both modes, paper config: 15/5 iterations, 6k/2k tokens)
|
| 17 |
+
- `retry_seedN_*` — retry passes for tasks whose environments needed the
|
| 18 |
+
archived-distro repair (see dataset card)
|
| 19 |
+
|
| 20 |
+
Companion artifacts:
|
| 21 |
+
- Transcripts (native CyBench JSON): `ajay-citadel/cybench-kimi-k3-transcripts`
|
| 22 |
+
- Harness + patches provenance: see dataset card and PATCHES.md therein
|
| 23 |
+
|
| 24 |
+
Logs are added as passes complete; refresh to see new batches.
|