775 MB
2,296 files
Updated 18 days ago
Name
Size
refs_audio
refs_first_frame
refs_id
.gitattributes2.5 kB
xet
README.md4.54 kB
xet
index.json1.77 MB
xet
seq_list.txt2 kB
xet
README.md

UnityShots Benchmark

A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot audio-video generation. Each case is a short cinematic story told across several shots, with a consistent cast whose identity, voice, and world must persist across every cut.

This is the evaluation benchmark released with UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating.

Overview

Sequences 200 stories (story_001–story_200)
Characters 430 named characters with reference identity + voice
Shots / story 3–6 (mean 5.0)
Languages 13 β€” Mandarin, Cantonese, English, German, Spanish, Arabic, Hindi, Bengali, Swahili, Yoruba, Persian, Portuguese, Vietnamese
Cultural regions 6 β€” East Asia, Europe, South/SE Asia, Africa, Latin America, Middle East
Conditioning modes Text-to-Video (T2V), Image-to-Video (I2V), Reference-to-Video (R2V)

Regional & language distribution

  • Regions: East Asia (70) Β· Europe (42) Β· South/SE Asia (26) Β· Africa (23) Β· Latin America (20) Β· Middle East (19)
  • Primary languages: Cantonese (77) Β· Mandarin (73) Β· German (51) Β· English (42) Β· Spanish (40) Β· Swahili (30) Β· Hindi (27) Β· Arabic (27) Β· Bengali (25) Β· Yoruba (19) Β· Persian (14) Β· Portuguese (3) Β· Vietnamese (2)

Dataset structure

UnityShotsBench/
β”œβ”€β”€ index.json                 # all 200 cases: metadata, characters, per-shot scripts
β”œβ”€β”€ seq_list.txt               # ordered list of sequence IDs
β”œβ”€β”€ refs_id/                   # reference identity portraits (for R2V)
β”‚   └── story_XXX/charN.jpg
β”œβ”€β”€ refs_audio/                # reference voice clips, per character (for R2V)
β”‚   └── story_XXX/charN.{wav,mp3}
└── refs_first_frame/          # per-shot first-frame anchors (for I2V)
    └── story_XXX/shot_K.jpg

index.json β€” per-case schema

{
  "sequence_id": "story_001",
  "type": "storytelling",
  "title": "Blueprints and Sweet Potatoes",
  "global_caption": "On a rainy night in Beijing, a veteran taxi driver and a weary architect ...",
  "_region": "east_asia",
  "_category": "general",
  "characters": [
    { "name": "...", "age": 58, "gender": "M", "language": "zh",
      "voice": "...", "appearance": "..." }
  ],
  "shots": [
    { "idx": 0, "duration_sec": 5.0, "transition": "HARD_CUT",
      "shot_caption": "...", "dialogue": "...", "audio": "..." }
  ]
}

Reference files follow the character order in index.json: the i-th character maps to charI.jpg (identity) and charI.wav / charI.mp3 (voice). I2V first frames map shot k to shot_K.jpg.

Conditioning modes

The same story can be evaluated under three input regimes, all covered by this benchmark:

  • T2V β€” per-shot text captions only.
  • I2V β€” the per-shot first-frame anchors in refs_first_frame/.
  • R2V β€” external identity (refs_id/) + reference voice (refs_audio/) per character.

Intended use & evaluation

Designed for fair evaluation of cross-shot identity preservation, voice/timbre consistency, lip-sync, scene continuity, and audio–text alignment in multi-shot generation. Reference identities are public-domain or AI-generated and are provided for academic, non-commercial research only.

Citation

@article{huang2026unityshots,
  title   = {UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating},
  author  = {Huang, Jiehui and Zhang, Yuechen and Xia, Bin and Wang, Jiahao and
             He, Xu and Tang, Zhenchao and Chu, Meng and Tao, Xin and Wan, Pengfei and Jia, Jiaya},
  journal = {arXiv preprint arXiv:2606.21661},
  year    = {2026}
}

License

Released under CC BY-NC 4.0 (non-commercial). Reference identities and voices are for academic research and benchmarking only.

Total size
775 MB
Files
2,296
Last updated
Jul 1
Pre-warmed CDN
US EU US EU

Contributors