OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 3 days ago • 47
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 3 days ago • 47
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 15 days ago • 57
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 15 days ago • 57
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis Paper • 2609.15309 • Published 18 days ago • 13
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 25 days ago • 77
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Paper • 2608.23383 • Published Aug 24 • 19
view post Post 120 Omni-Rewriter Replay: feed it a local clip, get a formatted H3 prompt, then generate if you want.Gallery: Wayne-King/omni-rewriter-replayCode: https://github.com/WayneJin0918/Omni-Rewriter See translation 👍 1 1 + Reply
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs Paper • 2603.20253 • Published Mar 11 • 2
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Paper • 2604.23781 • Published Apr 26 • 34
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos Paper • 2603.22529 • Published Mar 23 • 7
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting Paper • 2603.14659 • Published Mar 15 • 6