alvations commited on
Commit
f03c908
·
verified ·
1 Parent(s): 6842c6c

Deploy Hallway 8 (multi-arc memory game)

Browse files
Files changed (2) hide show
  1. README.md +19 -0
  2. static/audio.js +1 -1
README.md CHANGED
@@ -137,6 +137,24 @@ model included):** [`docs/REPRODUCE.md`](docs/REPRODUCE.md). Audio pipeline:
137
  [`docs/AUDIO.md`](docs/AUDIO.md). Reviewer panel and its reviews:
138
  [`docs/reviews/`](docs/reviews/).
139
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
140
  ---
141
 
142
  ## Run it locally
@@ -303,6 +321,7 @@ templates/ the page
303
  static/ style (skins), client logic, live audio, the "8" share art
304
  scripts/ i18n pipeline, deploy, audio measure/render, screenshot + video
305
  tests/ unit/API tests + the headless-browser emulation harness
 
306
  docs/ design, testing, audio, localization, reviews, ownership
307
  package.json the Node dev tool-chain (Playwright) for the test/capture tools
308
  Dockerfile container image for Hugging Face Spaces and other hosts
 
137
  [`docs/AUDIO.md`](docs/AUDIO.md). Reviewer panel and its reviews:
138
  [`docs/reviews/`](docs/reviews/).
139
 
140
+ ## Benchmark an LLM against it
141
+
142
+ Can a language model *remember a place*? [`llm-benchmark/`](llm-benchmark/) points a
143
+ local Hugging Face `transformers` model at the deployed game and measures how many
144
+ rounds it needs to escape each arc, across all **3 arcs × 3 difficulties × 2
145
+ reading levels**. The model plays exactly like a human (the server owns
146
+ correctness; the answer is never in the payload), so it is a genuine memory test.
147
+
148
+ ```bash
149
+ cd llm-benchmark && pip install -r requirements.txt
150
+ python benchmark.py --base-url http://127.0.0.1:5000 --agent random # baseline, no model
151
+ python benchmark.py --base-url http://127.0.0.1:5000 \
152
+ --agent llm --model google/gemma-3n-e4b-it # a local model plays
153
+ ```
154
+
155
+ See [`llm-benchmark/README.md`](llm-benchmark/README.md) for the agent protocol,
156
+ the sample prompt, and how to plug in your own model.
157
+
158
  ---
159
 
160
  ## Run it locally
 
321
  static/ style (skins), client logic, live audio, the "8" share art
322
  scripts/ i18n pipeline, deploy, audio measure/render, screenshot + video
323
  tests/ unit/API tests + the headless-browser emulation harness
324
+ llm-benchmark/ drive a local LLM against the deployed game as a memory benchmark
325
  docs/ design, testing, audio, localization, reviews, ownership
326
  package.json the Node dev tool-chain (Playwright) for the test/capture tools
327
  Dockerfile container image for Hugging Face Spaces and other hosts
static/audio.js CHANGED
@@ -473,7 +473,7 @@ class Ambience {
473
  // the whole game sits at one level: the low-frequency coach needs the most
474
  // lift (sub-bass reads quiet), the stairwell the least. See
475
  // scripts/measure_audio.cjs for how these were derived.
476
- const BUS = { landing: 1, hallway: 2.2, coach: 1.2, stairway: 1.48 };
477
  const busLevel = BUS[skin] || 2.2;
478
  if (skin === "coach") this._coach(bus);
479
  else if (skin === "stairway") this._stairway(bus);
 
473
  // the whole game sits at one level: the low-frequency coach needs the most
474
  // lift (sub-bass reads quiet), the stairwell the least. See
475
  // scripts/measure_audio.cjs for how these were derived.
476
+ const BUS = { landing: 1, hallway: 2.2, coach: 1.0, stairway: 1.48 };
477
  const busLevel = BUS[skin] || 2.2;
478
  if (skin === "coach") this._coach(bus);
479
  else if (skin === "stairway") this._stairway(bus);