Deep-VLM

community

AI & ML interests

None defined yet.

Recent Activity

tanhuajie2001 authored a paper 15 days ago

MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles

tanhuajie2001 authored a paper 15 days ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

tanhuajie2001 authored a paper 15 days ago

PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing

View all activity

authored 4 papers 15 days ago

MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles

Paper • 2510.00483 • Published Oct 1, 2025

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing

Paper • 2603.21669 • Published Mar 23 • 1

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 20 days ago • 27

authored 2 papers 16 days ago

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

Paper • 2605.05997 • Published May 7 • 18

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 20 days ago • 27

submitted a paper to Daily Papers 17 days ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 20 days ago • 27

authored a paper 4 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

submitted a paper to Daily Papers 4 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

authored 2 papers 4 months ago

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Paper • 2510.18795 • Published Oct 21, 2025 • 11

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

Paper • 2601.10305 • Published Jan 15 • 37

authored 4 papers 5 months ago

RoboBrain 2.5: Depth in Sight, Time in Mind

Paper • 2601.14352 • Published Jan 20 • 13

Towards Cross-View Point Correspondence in Vision-Language Models

Paper • 2512.04686 • Published Dec 4, 2025

RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration

Paper • 2510.26536 • Published Oct 30, 2025

Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation

Paper • 2512.23703 • Published Dec 29, 2025 • 7

submitted a paper to Daily Papers 6 months ago

Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation

Paper • 2512.23703 • Published Dec 29, 2025 • 7

authored 4 papers 8 months ago

ForCenNet: Foreground-Centric Network for Document Image Rectification

Paper • 2507.19804 • Published Jul 26, 2025 • 12

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Paper • 2509.09118 • Published Sep 11, 2025 • 8

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

Paper • 2510.13515 • Published Oct 15, 2025 • 12

ORID: Organ-Regional Information Driven Framework for Radiology Report Generation

Paper • 2411.13025 • Published Nov 20, 2024 • 2