File size: 5,621 Bytes
307bd2c
 
e375637
 
 
307bd2c
fc616f6
e375637
307bd2c
 
 
 
 
 
e375637
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
---
title: Vernacular
emoji: 😻
colorFrom: gray
colorTo: purple
sdk: gradio
sdk_version: 6.16.0
python_version: '3.10'
app_file: app.py
pinned: false
license: mit
short_description: Translate games. Keep every character's voice
---

# Vernacular

Vernacular is a Gradio review tool and character-aware translation pipeline for
*An Elmwood Trail*, a narrative mystery mobile game, developed by our friends at **Techyonic** <https://www.techyonic.co/>. It translates source strings,
uses character voice wikis to rewrite dialogue in the speaker's style, and lets
a reviewer approve, reject, or edit every final line.

The current app workflow also supports direct uploads: users can upload `.docx`,
`.xlsx`, `.json`, `.txt`, or `.csv` files, build/update character wikis, download
or re-upload the wiki JSON bundle, then press **Refresh** to send the uploaded
file through the normal translation and tone workflow.

```
Uploaded doc/xlsx ─► converter/docs.py or converter/sheets.py ─► English_JSON/Uploaded/
Uploaded wiki JSON ───────────────────────────────────────────► character_wikis/*.md
Uploaded source ─► wiki builder ─► character_wikis/*.md ─► character_wikis.json

English_JSON/Uploaded/ ─► TranslateGemma ─► Gemma tone pass + character wiki
                                               β”‚
                              translations/<lang>/ review records
                                               β”‚
                         app.py review: approve / reject / edit
```

## Setup

Install the runtime dependencies:

```bash
pip install -r requirements.txt
```

The app uses Hugging Face Transformers in-process, not local llama.cpp or Ollama
servers. The model IDs are configured in `config.py`:

```python
TRANSLATE_MODEL_ID = "google/translategemma-12b-it"
TONE_MODEL_ID = "google/gemma-4-12B-it"
```

For Hugging Face Spaces, add an `HF_TOKEN` secret for reliable downloads and for
access to gated Google Gemma models. The account behind that token must have
accepted the model terms.

Run the app locally:

```bash
python app.py
```

## App Workflow

1. Open the Gradio app.
2. In **Character wiki builder**, upload one or more source files.
3. Optionally upload an existing `character_wikis.json`.
4. Click **Build / update wiki JSON**.
5. Download the generated wiki JSON if you want to keep or reuse it later.
6. Click the existing **Refresh** button.
7. Select the uploaded file from the file dropdown and review its translated
   strings as usual.

For `.docx` and `.xlsx` uploads, the app uses the same converter modules as the
original pack conversion:

| Upload type | Conversion path |
|---|---|
| `.docx` | `converter/docs.py` via `readLocalDoc()` |
| `.xlsx` | `converter/sheets.py` via `readLocalSheet()` |
| `.json`, `.txt`, `.csv` | Plain-text fallback wrapped as uploaded chat rows |

If a converted doc/sheet produces no translatable strings, the app falls back to
plain extracted text so the upload still enters the workflow.

## What Uploads Create

Uploaded files are stored in app-managed paths:

| Path | Purpose |
|---|---|
| `uploaded_wiki_sources/` | Plain text JSON used to build/update character wikis |
| `English_JSON/Uploaded/Filler Chats/` | Converted source JSON used by translation/review |
| `character_data/` | Character index entries pointing to uploaded files |
| `character_wikis/*.md` | Per-character markdown wikis used by the tone pass |
| `character_wikis/character_wikis.json` | Downloadable/uploadable wiki bundle |
| `translations/<lang>/Uploaded/Filler Chats/` | Review records produced after Refresh |

## Existing Batch Workflow

The original full-pack workflow is still available:

```bash
python convert.py
python build_character_data.py
python build_character_wikis.py
python -m pipeline.build_file_context
python -m pipeline.translate_pack
python app.py
python -m pipeline.export_pack
```

`pipeline.translate_pack` remains resumable and supports targeted runs:

```bash
python -m pipeline.translate_pack --dry-run
python -m pipeline.translate_pack --filter Initial/
python -m pipeline.translate_pack --stage translate
python -m pipeline.translate_pack --stage tone
```

Note: `build_character_wikis.py` is the older batch wiki builder and still uses
its original local Ollama path. The Gradio upload workflow builds wikis through
the in-process Hugging Face Gemma client instead.

## Changing The Target Language

Edit the target values in `config.py`:

```python
TARGET_LANG_CODE = "de"
TARGET_LANG_NAME = "German"
PACK_NAME = "German_JSON"
```

Then rerun the translation/review/export workflow. TranslateGemma supports many
language pairs, but the selected Hugging Face model and token access must be
available in the runtime.

## Repository Map

| Path | Purpose |
|---|---|
| `app.py` | Gradio app: wiki upload/build, refresh, review UI |
| `requirements.txt` | Space/local Python runtime dependencies |
| `config.py` | Target language, paths, Hugging Face model settings |
| `converter/` | `.docx` / `.xlsx` to structured JSON conversion |
| `English_JSON/` | Source-of-truth JSON files, including uploaded files |
| `character_data/` | Per-character file indexes |
| `character_wikis/` | Markdown wikis and downloadable wiki JSON bundle |
| `pipeline/clients.py` | In-process TranslateGemma and Gemma tone inference |
| `pipeline/translate_pack.py` | Batch translation/tone driver |
| `translations/<lang>/` | Review records used by the app |
| `pipeline/export_pack.py` | Exports approved/edited translations to final pack |