File size: 5,621 Bytes
307bd2c e375637 307bd2c fc616f6 e375637 307bd2c e375637 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 | ---
title: Vernacular
emoji: π»
colorFrom: gray
colorTo: purple
sdk: gradio
sdk_version: 6.16.0
python_version: '3.10'
app_file: app.py
pinned: false
license: mit
short_description: Translate games. Keep every character's voice
---
# Vernacular
Vernacular is a Gradio review tool and character-aware translation pipeline for
*An Elmwood Trail*, a narrative mystery mobile game, developed by our friends at **Techyonic** <https://www.techyonic.co/>. It translates source strings,
uses character voice wikis to rewrite dialogue in the speaker's style, and lets
a reviewer approve, reject, or edit every final line.
The current app workflow also supports direct uploads: users can upload `.docx`,
`.xlsx`, `.json`, `.txt`, or `.csv` files, build/update character wikis, download
or re-upload the wiki JSON bundle, then press **Refresh** to send the uploaded
file through the normal translation and tone workflow.
```
Uploaded doc/xlsx ββΊ converter/docs.py or converter/sheets.py ββΊ English_JSON/Uploaded/
Uploaded wiki JSON ββββββββββββββββββββββββββββββββββββββββββββΊ character_wikis/*.md
Uploaded source ββΊ wiki builder ββΊ character_wikis/*.md ββΊ character_wikis.json
English_JSON/Uploaded/ ββΊ TranslateGemma ββΊ Gemma tone pass + character wiki
β
translations/<lang>/ review records
β
app.py review: approve / reject / edit
```
## Setup
Install the runtime dependencies:
```bash
pip install -r requirements.txt
```
The app uses Hugging Face Transformers in-process, not local llama.cpp or Ollama
servers. The model IDs are configured in `config.py`:
```python
TRANSLATE_MODEL_ID = "google/translategemma-12b-it"
TONE_MODEL_ID = "google/gemma-4-12B-it"
```
For Hugging Face Spaces, add an `HF_TOKEN` secret for reliable downloads and for
access to gated Google Gemma models. The account behind that token must have
accepted the model terms.
Run the app locally:
```bash
python app.py
```
## App Workflow
1. Open the Gradio app.
2. In **Character wiki builder**, upload one or more source files.
3. Optionally upload an existing `character_wikis.json`.
4. Click **Build / update wiki JSON**.
5. Download the generated wiki JSON if you want to keep or reuse it later.
6. Click the existing **Refresh** button.
7. Select the uploaded file from the file dropdown and review its translated
strings as usual.
For `.docx` and `.xlsx` uploads, the app uses the same converter modules as the
original pack conversion:
| Upload type | Conversion path |
|---|---|
| `.docx` | `converter/docs.py` via `readLocalDoc()` |
| `.xlsx` | `converter/sheets.py` via `readLocalSheet()` |
| `.json`, `.txt`, `.csv` | Plain-text fallback wrapped as uploaded chat rows |
If a converted doc/sheet produces no translatable strings, the app falls back to
plain extracted text so the upload still enters the workflow.
## What Uploads Create
Uploaded files are stored in app-managed paths:
| Path | Purpose |
|---|---|
| `uploaded_wiki_sources/` | Plain text JSON used to build/update character wikis |
| `English_JSON/Uploaded/Filler Chats/` | Converted source JSON used by translation/review |
| `character_data/` | Character index entries pointing to uploaded files |
| `character_wikis/*.md` | Per-character markdown wikis used by the tone pass |
| `character_wikis/character_wikis.json` | Downloadable/uploadable wiki bundle |
| `translations/<lang>/Uploaded/Filler Chats/` | Review records produced after Refresh |
## Existing Batch Workflow
The original full-pack workflow is still available:
```bash
python convert.py
python build_character_data.py
python build_character_wikis.py
python -m pipeline.build_file_context
python -m pipeline.translate_pack
python app.py
python -m pipeline.export_pack
```
`pipeline.translate_pack` remains resumable and supports targeted runs:
```bash
python -m pipeline.translate_pack --dry-run
python -m pipeline.translate_pack --filter Initial/
python -m pipeline.translate_pack --stage translate
python -m pipeline.translate_pack --stage tone
```
Note: `build_character_wikis.py` is the older batch wiki builder and still uses
its original local Ollama path. The Gradio upload workflow builds wikis through
the in-process Hugging Face Gemma client instead.
## Changing The Target Language
Edit the target values in `config.py`:
```python
TARGET_LANG_CODE = "de"
TARGET_LANG_NAME = "German"
PACK_NAME = "German_JSON"
```
Then rerun the translation/review/export workflow. TranslateGemma supports many
language pairs, but the selected Hugging Face model and token access must be
available in the runtime.
## Repository Map
| Path | Purpose |
|---|---|
| `app.py` | Gradio app: wiki upload/build, refresh, review UI |
| `requirements.txt` | Space/local Python runtime dependencies |
| `config.py` | Target language, paths, Hugging Face model settings |
| `converter/` | `.docx` / `.xlsx` to structured JSON conversion |
| `English_JSON/` | Source-of-truth JSON files, including uploaded files |
| `character_data/` | Per-character file indexes |
| `character_wikis/` | Markdown wikis and downloadable wiki JSON bundle |
| `pipeline/clients.py` | In-process TranslateGemma and Gemma tone inference |
| `pipeline/translate_pack.py` | Batch translation/tone driver |
| `translations/<lang>/` | Review records used by the app |
| `pipeline/export_pack.py` | Exports approved/edited translations to final pack | |