Models
Datasets
Spaces
Buckets new
Docs
Enterprise
Pricing
- Website
- Community
- Solutions
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2605.12500

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture

sensenova/SenseNova-U1-8B-MoT

Any-to-Any • 18B • Updated 16 days ago • 27.5k • 280
sensenova/SenseNova-U1-8B-MoT-Infographic

Any-to-Any • 18B • Updated 16 days ago • 5.41k • 40
sensenova/SenseNova-U1-8B-MoT-SFT

Any-to-Any • 18B • Updated 17 days ago • 1.62k • 51
sensenova/SenseNova-U1-8B-MoT-LoRAs

Updated 17 days ago • 5

about 3 hours ago

Code as Agent Harness

Paper • 2605.18747 • Published 14 days ago • 210
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191
From Context to Skills: Can Language Models Learn from Context Skillfully?

Paper • 2604.27660 • Published 29 days ago • 166
PhysBrain 1.0 Technical Report

Paper • 2605.15298 • Published 18 days ago • 143

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

Paper • 2603.25746 • Published Mar 26 • 155
TAPS: Task Aware Proposal Distributions for Speculative Sampling

Paper • 2603.27027 • Published Mar 27 • 144
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Paper • 2603.25716 • Published Mar 26 • 156
LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Paper • 2603.27538 • Published Mar 29 • 147

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Paper • 2402.04252 • Published Feb 6, 2024 • 31
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

Paper • 2402.03749 • Published Feb 6, 2024 • 15
ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Paper • 2402.04615 • Published Feb 7, 2024 • 45
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss

Paper • 2402.05008 • Published Feb 7, 2024 • 24

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191
From Context to Skills: Can Language Models Learn from Context Skillfully?

Paper • 2604.27660 • Published 29 days ago • 166
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

Paper • 2605.03849 • Published 27 days ago • 125
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

Paper • 2605.03042 • Published 28 days ago • 124

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture

sensenova/SenseNova-U1-8B-MoT

Any-to-Any • 18B • Updated 16 days ago • 27.5k • 280
sensenova/SenseNova-U1-8B-MoT-Infographic

Any-to-Any • 18B • Updated 16 days ago • 5.41k • 40
sensenova/SenseNova-U1-8B-MoT-SFT

Any-to-Any • 18B • Updated 17 days ago • 1.62k • 51
sensenova/SenseNova-U1-8B-MoT-LoRAs

Updated 17 days ago • 5

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Paper • 2402.04252 • Published Feb 6, 2024 • 31
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

Paper • 2402.03749 • Published Feb 6, 2024 • 15
ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Paper • 2402.04615 • Published Feb 7, 2024 • 45
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss

Paper • 2402.05008 • Published Feb 7, 2024 • 24

about 3 hours ago

Code as Agent Harness

Paper • 2605.18747 • Published 14 days ago • 210
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191
From Context to Skills: Can Language Models Learn from Context Skillfully?

Paper • 2604.27660 • Published 29 days ago • 166
PhysBrain 1.0 Technical Report

Paper • 2605.15298 • Published 18 days ago • 143

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191
From Context to Skills: Can Language Models Learn from Context Skillfully?

Paper • 2604.27660 • Published 29 days ago • 166
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

Paper • 2605.03849 • Published 27 days ago • 125
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

Paper • 2605.03042 • Published 28 days ago • 124

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Paper • 2605.12500 • Published 20 days ago • 191

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

Paper • 2603.25746 • Published Mar 26 • 155
TAPS: Task Aware Proposal Distributions for Speculative Sampling

Paper • 2603.27027 • Published Mar 27 • 144
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Paper • 2603.25716 • Published Mar 26 • 156
LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Paper • 2603.27538 • Published Mar 29 • 147

Company

TOS Privacy About Careers

Website

Models Datasets Spaces Pricing Docs