hacktest / CLAUDE.md
bhardwaj08sarthak's picture
Upload 13 files
cbe0298 verified
|
Raw
History Blame Contribute Delete
8 kB

A newer version of the Gradio SDK is available: 6.28.0

Upgrade

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What This Is

A Hugging Face Space (Gradio app) that helps translate the language pack for Riverstone, a narrative mystery mobile game. The app.py is the deployed entrypoint; the English/ directory is the source language pack containing all in-game text as .docx and .xlsx files.

Running the App

pip install gradio
python app.py

The app is deployed on Hugging Face Spaces (Gradio SDK 6.16.0, Python 3.13). The README.md frontmatter is the HF Space config β€” don't remove it.

app.py is the translation review tool: it reads/writes the review records in translations/<lang>/ only β€” it makes no LLM calls, so it runs anywhere.

Translation Pipeline

English_JSON ─► pipeline/translate_pack.py ─► translations/<lang>/ (review records)
                 stage 1: TranslateGemma 12B via llama-server :8089 (greedy)
                 stage 2: tone pass via Ollama gemma4 + character_wikis/<id>.md
translations/<lang>/ ─► app.py (review) ─► pipeline/export_pack.py ─► German_JSON/
  • config.py β€” single source of truth: target language (de/German by default), paths, server URLs, model names, sampling params. Change language here.
  • pipeline/rules.py β€” decides which strings are translatable and addresses them with key paths; kinds: dialogue (gets tone pass), player, other. Mirrors the "Never translate" rules below.
  • pipeline/dialogue_map.py β€” chat file β†’ character id, from character_data/ mentions, only for characters with a wiki. Conversations: message = character, options = player. Filler Chats: type "1" = player, "-1" = system, anything else = the mapped character/group.
  • pipeline/clients.py β€” model clients. TranslateGemma's jinja template doesn't parse in llama.cpp, so translate() posts the raw Gemma turn format to /completion with the structured type:text,source_lang_code:... prefix and a primed model turn. restore_placeholders() un-breaks Person1-style tokens the models tend to split.
  • pipeline/translate_pack.py β€” resumable batch driver (--filter, --limit, --stage, --dry-run); shared string cache at translations/<lang>/.cache.json.
  • pipeline/export_pack.py β€” builds the target pack; per string: reviewer final > toned > mt > English source. Raises on source drift.
  • pipeline/build_file_context.py β€” parses teammate-maintained FILE_MAPPING.md into file_context.json (per-file label/role/summary shown in the review UI).

Start the translation server with:

llama-server -hf bullerwins/translategemma-12b-it-GGUF:Q4_K_M \
    --port 8089 --swa-full --ctx-size 4096 --no-jinja --chat-template gemma

Tests: python -m pytest tests/ (rules + export structure preservation).

Converting the Language Pack to JSON

python convert.py            # English/ -> English_JSON/
python convert.py SRC DST    # custom source / destination dirs

This uses stdlib only (zipfile + xml.etree) β€” no extra dependencies. It skips files whose names contain any of the skipMarkers in converter/utils.py (e.g. (Not to be translated)).

Converter Architecture

The converter/ package contains the game owner's original processing logic, adapted from cloud (Google Drive/Docs API) to local file reading.

Data flow:

English/*.docx  ->  converter/docs.py  ->  readDocElements()  ->  structured JSON
English/*.xlsx  ->  converter/sheets.py                        ->  structured JSON

converter/docs.py β€” .docx processor.

  • _docx_to_body_content(path) reads a local .docx (OOXML zip) and converts it into the Google Docs API dict format (body.content with paragraph/table/textRun keys) so the owner's readDocElements() can process it unchanged.
  • readDocElements() is the owner's original semantic extractor: converts column headers to snake_case, applies headerReplacements, detects "array fields" (columns that repeat across continuation rows where the ID cell is empty), and groups rows by their first-column key.
  • readLocalDoc(path, file_path) is the public entry point.

converter/sheets.py β€” .xlsx processor.

  • Reads raw OOXML cell grid, normalises headers the same way as readDocElements, detects array fields by the same empty-first-cell rule, and produces identically structured output.
  • readLocalSheet(path, file_path) is the public entry point.

converter/utils.py β€” shared helpers: toSnakeCase, toSafeEntityName, entityNameReplacements, skipMarkers, are_keys_sequential, getMarkedFields.

convert.py β€” walker: recurses English/, calls readLocalDoc or readLocalSheet per file, writes mirrored .json files under English_JSON/.

Language Pack Architecture

The translation target is the English/ directory. A translated pack mirrors this entire folder structure, replacing text content while keeping all IDs and structural values intact.

Unit structure:

Folder Content
Initial/ Episodes 1–2 + all baseline UI
U1_MN/ Unit 1 supplementary (character archive, extra chats)
U2_MJ_S4S5/ Scenes 4–5 of Unit 2 (Sidetrails 3–4)
U3_MJ_E3/ Episode 3

Each unit has General/ (UI string tables), Sequences/ (story objectives), and Gameplay/ (Adam's Phone, Zoey's Phone, Desk, Other).

File formats:

  • .docx files: tables where Column 1 is always the ID (never translate) and remaining columns are translatable content (Text, Title, Caption, Subtitle, Content, etc.).
  • .xlsx branching dialogue (Schema A, in Conversations/): columns ID | Message | Next Message ID | Options | Options next ID β€” translate Message and Options only.
  • .xlsx linear chat logs (Schema B, in Filler Chats/): columns ID | type | text | time β€” translate text only, except values starting with system:: which are string ID references and must not be changed.

Never translate: ID columns, navigation pointer columns (Next Message ID, Options next ID), Audio columns (mp3 filenames), system::key values, - (empty/end-of-branch markers), type/time chat columns, files named (Not to be translated).

Placeholder tokens like Person1, Person2 in strings are runtime-injected names β€” keep verbatim, translate only surrounding words.

JSON Output Format

.docx tables with 2 columns (ID | Value) produce a flat dict:

{ "messages": "Messages", "calls": "Calls" }

.docx tables with 3+ columns produce a dict of objects:

{ "some_id": { "title": "...", "text": "...", "hint": "..." } }

.xlsx Schema A (branching dialogue) produces a dict keyed by message ID with options as arrays:

{ "#0": { "message": "Hey Adam", "next": null, "options": ["Hi Tim", "Hey"], "options_next": ["#1a", "#1a"] } }

.xlsx Schema B (chat log) produces a list (sequential integer IDs collapse to array):

[ { "type": "message", "text": "Hello!", "time": "07:15 AM" }, ... ]

If a .docx has multiple tables, the output is a list of the above dicts (one per table).

Special Processing Rules

Defined in converter/docs.py:

  • headerReplacements: column name aliases applied after snake_case β€” e.g. next_message_id β†’ next, options_next_id β†’ options_next.
  • avoidTransformMarkers: files matching pixabowl/posts keep their id field un-snake-cased (post IDs like bGH45f8 must be preserved verbatim).
  • splitMarkers: files matching sidetrails/connections/items split the bio field on newlines into a list.

The file_path key passed to these markers is the normalised path relative to English/, with each component run through toSafeEntityName (e.g. English/U2_MJ_S4S5/Gameplay/Sidetrails/Connections/Items.docx β†’ u2mjs4s5/gameplay/sidetrails/connections/items).