Workspace cleanup — 2026-08-17
This manifest records the conservative post-competition cleanup policy.
Preserved
- All training/evaluation datasets, including
experiment_4,external_eval,rasam-dataset-main,_mh, PAGE-XML/image corpora, andtmp/exp8. - All Git worktrees and their branch history.
- Deployed recognition, segmentation, specialist, and language models.
- The local database and
.envneeded to reproduce the running application; both remain ignored by Git. - Source-retrieval texts, reports, benchmark summaries, and the prepared Hugging Face package.
- The shared backend virtual environment used by the demo launcher.
Removed as regenerable or superseded
- Repeated
_wheelsdownload caches in the four worktrees. - The isolated EasyOCR pilot environment and an unrelated temporary Python environment.
- Playwright, Python bytecode, and pytest caches; browser screenshots/output.
- Untracked competition ZIPs and duplicate unpacked presentation bundles.
- Temporary server logs, OCR book probes, RAC/top-k caches, and abandoned server-patch build directories.
- The pre-competition copy and the pre-segmentation source backup after confirming that the active models and Git history are present.
- Root-level scratch JSON/TXT/log files and stale PID files.
No training dataset, deployed model, language model, source corpus, Git history, or runtime database was intentionally removed.