Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 22 days ago • 265
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 22 days ago • 265
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation Paper • 2608.08469 • Published Aug 9 • 2
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published Aug 17 • 9
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published Aug 12 • 43
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Paper • 2607.03451 • Published Jul 3 • 36
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Paper • 2605.05997 • Published May 7 • 18
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Paper • 2605.25979 • Published May 25 • 24