Deploy Hallway 8 (multi-arc memory game)
Browse files- README.md +19 -0
- static/audio.js +1 -1
README.md
CHANGED
|
@@ -137,6 +137,24 @@ model included):** [`docs/REPRODUCE.md`](docs/REPRODUCE.md). Audio pipeline:
|
|
| 137 |
[`docs/AUDIO.md`](docs/AUDIO.md). Reviewer panel and its reviews:
|
| 138 |
[`docs/reviews/`](docs/reviews/).
|
| 139 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
---
|
| 141 |
|
| 142 |
## Run it locally
|
|
@@ -303,6 +321,7 @@ templates/ the page
|
|
| 303 |
static/ style (skins), client logic, live audio, the "8" share art
|
| 304 |
scripts/ i18n pipeline, deploy, audio measure/render, screenshot + video
|
| 305 |
tests/ unit/API tests + the headless-browser emulation harness
|
|
|
|
| 306 |
docs/ design, testing, audio, localization, reviews, ownership
|
| 307 |
package.json the Node dev tool-chain (Playwright) for the test/capture tools
|
| 308 |
Dockerfile container image for Hugging Face Spaces and other hosts
|
|
|
|
| 137 |
[`docs/AUDIO.md`](docs/AUDIO.md). Reviewer panel and its reviews:
|
| 138 |
[`docs/reviews/`](docs/reviews/).
|
| 139 |
|
| 140 |
+
## Benchmark an LLM against it
|
| 141 |
+
|
| 142 |
+
Can a language model *remember a place*? [`llm-benchmark/`](llm-benchmark/) points a
|
| 143 |
+
local Hugging Face `transformers` model at the deployed game and measures how many
|
| 144 |
+
rounds it needs to escape each arc, across all **3 arcs × 3 difficulties × 2
|
| 145 |
+
reading levels**. The model plays exactly like a human (the server owns
|
| 146 |
+
correctness; the answer is never in the payload), so it is a genuine memory test.
|
| 147 |
+
|
| 148 |
+
```bash
|
| 149 |
+
cd llm-benchmark && pip install -r requirements.txt
|
| 150 |
+
python benchmark.py --base-url http://127.0.0.1:5000 --agent random # baseline, no model
|
| 151 |
+
python benchmark.py --base-url http://127.0.0.1:5000 \
|
| 152 |
+
--agent llm --model google/gemma-3n-e4b-it # a local model plays
|
| 153 |
+
```
|
| 154 |
+
|
| 155 |
+
See [`llm-benchmark/README.md`](llm-benchmark/README.md) for the agent protocol,
|
| 156 |
+
the sample prompt, and how to plug in your own model.
|
| 157 |
+
|
| 158 |
---
|
| 159 |
|
| 160 |
## Run it locally
|
|
|
|
| 321 |
static/ style (skins), client logic, live audio, the "8" share art
|
| 322 |
scripts/ i18n pipeline, deploy, audio measure/render, screenshot + video
|
| 323 |
tests/ unit/API tests + the headless-browser emulation harness
|
| 324 |
+
llm-benchmark/ drive a local LLM against the deployed game as a memory benchmark
|
| 325 |
docs/ design, testing, audio, localization, reviews, ownership
|
| 326 |
package.json the Node dev tool-chain (Playwright) for the test/capture tools
|
| 327 |
Dockerfile container image for Hugging Face Spaces and other hosts
|
static/audio.js
CHANGED
|
@@ -473,7 +473,7 @@ class Ambience {
|
|
| 473 |
// the whole game sits at one level: the low-frequency coach needs the most
|
| 474 |
// lift (sub-bass reads quiet), the stairwell the least. See
|
| 475 |
// scripts/measure_audio.cjs for how these were derived.
|
| 476 |
-
const BUS = { landing: 1, hallway: 2.2, coach: 1.
|
| 477 |
const busLevel = BUS[skin] || 2.2;
|
| 478 |
if (skin === "coach") this._coach(bus);
|
| 479 |
else if (skin === "stairway") this._stairway(bus);
|
|
|
|
| 473 |
// the whole game sits at one level: the low-frequency coach needs the most
|
| 474 |
// lift (sub-bass reads quiet), the stairwell the least. See
|
| 475 |
// scripts/measure_audio.cjs for how these were derived.
|
| 476 |
+
const BUS = { landing: 1, hallway: 2.2, coach: 1.0, stairway: 1.48 };
|
| 477 |
const busLevel = BUS[skin] || 2.2;
|
| 478 |
if (skin === "coach") this._coach(bus);
|
| 479 |
else if (skin === "stairway") this._stairway(bus);
|