EgoSuite-Open100K Collection The largest fully-annotated open egocentric human dataset. 100,000 hours across 15,000+ tasks and scenes. ⢠3 items ⢠Updated 6 days ago ⢠40
TIPSv2 Collection TIPSv2 foundational vision-language models. Webpage: https://gdm-tipsv2.github.io/ ⢠9 items ⢠Updated Jul 21 ⢠44
VST Collection A comprehensive framework designed to cultivate VLMs with human-like visuospatial abilities. ⢠7 items ⢠Updated Jun 15 ⢠6
Cosmos-Predict2 Collection ā ļø This collection is archived. š https://huggingface.co/collections/nvidia/cosmos-predict25 ⢠10 items ⢠Updated 14 days ago ⢠37
Cosmos World Foundation Model Platform for Physical AI Paper ⢠2501.03575 ⢠Published Jan 7, 2025 ⢠84
Physical AI Collection Collection of open, commercial-grade datasets for physical AI developers ⢠57 items ⢠Updated 14 days ago ⢠179
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper ⢠2501.17161 ⢠Published Jan 28, 2025 ⢠127
PixMo Collection A set of vision-language datasets built by Ai2 and used to train the Molmo family of models. Read more at https://molmo.allenai.org/blog ⢠9 items ⢠Updated Mar 2 ⢠92
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step Paper ⢠2411.10440 ⢠Published Jul 21, 2025 ⢠133
Theia Collection Distilling Diverse Vision Foundation Models for Robot Learning ⢠6 items ⢠Updated Sep 30, 2024 ⢠9
view article Article Metric and Relative Monocular Depth Estimation: An Overview. Fine-Tuning Depth Anything V2 š š Isayoften ⢠Jul 10, 2024 ⢠100
3D-VLA: A 3D Vision-Language-Action Generative World Model Paper ⢠2403.09631 ⢠Published Mar 14, 2024 ⢠12
Minitron Collection A family of compressed models obtained via pruning and knowledge distillation ⢠12 items ⢠Updated 14 days ago ⢠63
OpenResearcher: Unleashing AI for Accelerated Scientific Research Paper ⢠2408.06941 ⢠Published Aug 13, 2024 ⢠32