# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## What This Is A Hugging Face Space (Gradio app) that helps translate the language pack for *Riverstone*, a narrative mystery mobile game. The `app.py` is the deployed entrypoint; the `English/` directory is the source language pack containing all in-game text as `.docx` and `.xlsx` files. ## Running the App ```bash pip install gradio python app.py ``` The app is deployed on Hugging Face Spaces (Gradio SDK 6.16.0, Python 3.13). The `README.md` frontmatter is the HF Space config — don't remove it. `app.py` is the **translation review tool**: it reads/writes the review records in `translations//` only — it makes no LLM calls, so it runs anywhere. ## Translation Pipeline ``` English_JSON ─► pipeline/translate_pack.py ─► translations// (review records) stage 1: TranslateGemma 12B via llama-server :8089 (greedy) stage 2: tone pass via Ollama gemma4 + character_wikis/.md translations// ─► app.py (review) ─► pipeline/export_pack.py ─► German_JSON/ ``` - **`config.py`** — single source of truth: target language (`de`/German by default), paths, server URLs, model names, sampling params. Change language here. - **`pipeline/rules.py`** — decides which strings are translatable and addresses them with key paths; kinds: `dialogue` (gets tone pass), `player`, `other`. Mirrors the "Never translate" rules below. - **`pipeline/dialogue_map.py`** — chat file → character id, from `character_data/` mentions, only for characters with a wiki. Conversations: `message` = character, `options` = player. Filler Chats: type `"1"` = player, `"-1"` = system, anything else = the mapped character/group. - **`pipeline/clients.py`** — model clients. TranslateGemma's jinja template doesn't parse in llama.cpp, so `translate()` posts the raw Gemma turn format to `/completion` with the structured `type:text,source_lang_code:...` prefix and a primed model turn. `restore_placeholders()` un-breaks `Person1`-style tokens the models tend to split. - **`pipeline/translate_pack.py`** — resumable batch driver (`--filter`, `--limit`, `--stage`, `--dry-run`); shared string cache at `translations//.cache.json`. - **`pipeline/export_pack.py`** — builds the target pack; per string: reviewer `final` > `toned` > `mt` > English source. Raises on source drift. - **`pipeline/build_file_context.py`** — parses teammate-maintained `FILE_MAPPING.md` into `file_context.json` (per-file label/role/summary shown in the review UI). Start the translation server with: ```bash llama-server -hf bullerwins/translategemma-12b-it-GGUF:Q4_K_M \ --port 8089 --swa-full --ctx-size 4096 --no-jinja --chat-template gemma ``` Tests: `python -m pytest tests/` (rules + export structure preservation). ## Converting the Language Pack to JSON ```bash python convert.py # English/ -> English_JSON/ python convert.py SRC DST # custom source / destination dirs ``` This uses stdlib only (`zipfile` + `xml.etree`) — no extra dependencies. It skips files whose names contain any of the `skipMarkers` in `converter/utils.py` (e.g. `(Not to be translated)`). ## Converter Architecture The `converter/` package contains the game owner's original processing logic, adapted from cloud (Google Drive/Docs API) to local file reading. **Data flow:** ``` English/*.docx -> converter/docs.py -> readDocElements() -> structured JSON English/*.xlsx -> converter/sheets.py -> structured JSON ``` **`converter/docs.py`** — `.docx` processor. - `_docx_to_body_content(path)` reads a local `.docx` (OOXML zip) and converts it into the Google Docs API dict format (`body.content` with `paragraph`/`table`/`textRun` keys) so the owner's `readDocElements()` can process it unchanged. - `readDocElements()` is the owner's original semantic extractor: converts column headers to `snake_case`, applies `headerReplacements`, detects "array fields" (columns that repeat across continuation rows where the ID cell is empty), and groups rows by their first-column key. - `readLocalDoc(path, file_path)` is the public entry point. **`converter/sheets.py`** — `.xlsx` processor. - Reads raw OOXML cell grid, normalises headers the same way as `readDocElements`, detects array fields by the same empty-first-cell rule, and produces identically structured output. - `readLocalSheet(path, file_path)` is the public entry point. **`converter/utils.py`** — shared helpers: `toSnakeCase`, `toSafeEntityName`, `entityNameReplacements`, `skipMarkers`, `are_keys_sequential`, `getMarkedFields`. **`convert.py`** — walker: recurses `English/`, calls `readLocalDoc` or `readLocalSheet` per file, writes mirrored `.json` files under `English_JSON/`. ## Language Pack Architecture The translation target is the `English/` directory. A translated pack mirrors this entire folder structure, replacing text content while keeping all IDs and structural values intact. **Unit structure:** | Folder | Content | |---|---| | `Initial/` | Episodes 1–2 + all baseline UI | | `U1_MN/` | Unit 1 supplementary (character archive, extra chats) | | `U2_MJ_S4S5/` | Scenes 4–5 of Unit 2 (Sidetrails 3–4) | | `U3_MJ_E3/` | Episode 3 | Each unit has `General/` (UI string tables), `Sequences/` (story objectives), and `Gameplay/` (Adam's Phone, Zoey's Phone, Desk, Other). **File formats:** - `.docx` files: tables where Column 1 is always the `ID` (never translate) and remaining columns are translatable content (`Text`, `Title`, `Caption`, `Subtitle`, `Content`, etc.). - `.xlsx` branching dialogue (Schema A, in `Conversations/`): columns `ID | Message | Next Message ID | Options | Options next ID` — translate `Message` and `Options` only. - `.xlsx` linear chat logs (Schema B, in `Filler Chats/`): columns `ID | type | text | time` — translate `text` only, except values starting with `system::` which are string ID references and must not be changed. **Never translate:** `ID` columns, navigation pointer columns (`Next Message ID`, `Options next ID`), `Audio` columns (mp3 filenames), `system::key` values, `-` (empty/end-of-branch markers), `type`/`time` chat columns, files named `(Not to be translated)`. **Placeholder tokens** like `Person1`, `Person2` in strings are runtime-injected names — keep verbatim, translate only surrounding words. ## JSON Output Format `.docx` tables with 2 columns (`ID | Value`) produce a flat dict: ```json { "messages": "Messages", "calls": "Calls" } ``` `.docx` tables with 3+ columns produce a dict of objects: ```json { "some_id": { "title": "...", "text": "...", "hint": "..." } } ``` `.xlsx` Schema A (branching dialogue) produces a dict keyed by message ID with options as arrays: ```json { "#0": { "message": "Hey Adam", "next": null, "options": ["Hi Tim", "Hey"], "options_next": ["#1a", "#1a"] } } ``` `.xlsx` Schema B (chat log) produces a list (sequential integer IDs collapse to array): ```json [ { "type": "message", "text": "Hello!", "time": "07:15 AM" }, ... ] ``` If a `.docx` has multiple tables, the output is a list of the above dicts (one per table). ## Special Processing Rules Defined in `converter/docs.py`: - **`headerReplacements`**: column name aliases applied after snake_case — e.g. `next_message_id` → `next`, `options_next_id` → `options_next`. - **`avoidTransformMarkers`**: files matching `pixabowl/posts` keep their `id` field un-snake-cased (post IDs like `bGH45f8` must be preserved verbatim). - **`splitMarkers`**: files matching `sidetrails/connections/items` split the `bio` field on newlines into a list. The `file_path` key passed to these markers is the normalised path relative to `English/`, with each component run through `toSafeEntityName` (e.g. `English/U2_MJ_S4S5/Gameplay/Sidetrails/Connections/Items.docx` → `u2mjs4s5/gameplay/sidetrails/connections/items`).