phoenix-arabic-manuscript-htr / reports /CLEANUP_20260817.md
factlogic's picture
Update publication links for Phoenix model name (part 2)
7211659 verified
|
Raw
History Blame
1.44 kB

Workspace cleanup — 2026-08-17

This manifest records the conservative post-competition cleanup policy.

Preserved

  • All training/evaluation datasets, including experiment_4, external_eval, rasam-dataset-main, _mh, PAGE-XML/image corpora, and tmp/exp8.
  • All Git worktrees and their branch history.
  • Deployed recognition, segmentation, specialist, and language models.
  • The local database and .env needed to reproduce the running application; both remain ignored by Git.
  • Source-retrieval texts, reports, benchmark summaries, and the prepared Hugging Face package.
  • The shared backend virtual environment used by the demo launcher.

Removed as regenerable or superseded

  • Repeated _wheels download caches in the four worktrees.
  • The isolated EasyOCR pilot environment and an unrelated temporary Python environment.
  • Playwright, Python bytecode, and pytest caches; browser screenshots/output.
  • Untracked competition ZIPs and duplicate unpacked presentation bundles.
  • Temporary server logs, OCR book probes, RAC/top-k caches, and abandoned server-patch build directories.
  • The pre-competition copy and the pre-segmentation source backup after confirming that the active models and Git history are present.
  • Root-level scratch JSON/TXT/log files and stale PID files.

No training dataset, deployed model, language model, source corpus, Git history, or runtime database was intentionally removed.