Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue Paper • 2609.31948 • Published 8 days ago • 89
view article Article From Coding Agents to Physical Agents: Code as a New Interface Between AI and the Physical World kkakkkka • 15 days ago • 3
view article Article VLANeXt: A Simple and Research-Oriented Codebase for Robotics Research cavanloy • Aug 31 • 23
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published about 1 month ago • 71
view article Article Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States KangLiao • about 1 month ago • 23
Puffin-World Collection A collection covering Puffin-World's models, datasets (Puffin-16M), annotated public datasets, blog, and paper. • 33 items • Updated 22 days ago • 7
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper • 2605.18739 • Published May 18 • 117
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration Paper • 2605.14278 • Published May 14 • 36
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners? Paper • 2603.25823 • Published Mar 26 • 43
MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment Paper • 2512.06628 • Published Dec 7, 2025 • 12
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement Paper • 2511.23475 • Published Nov 28, 2025 • 43
Controllable Layer Decomposition for Reversible Multi-Layer Image Generation Paper • 2511.16249 • Published Nov 20, 2025 • 10
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning Paper • 2510.11026 • Published Oct 13, 2025 • 18
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation Paper • 2509.18824 • Published Sep 23, 2025 • 23
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation Paper • 2507.09862 • Published Jul 14, 2025 • 52