TokenRouter: Efficient Serving System for Token-Level LLM Routing Paper • 2610.12242 • Published 3 days ago • 123
view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 17 days ago • 68
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 30 days ago • 266
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 19 days ago • 103
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 24 days ago • 57
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 24 days ago • 94
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 24 days ago • 139
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 24 days ago • 78
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 24 days ago • 115
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 24 days ago • 228
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published Sep 9 • 45
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published about 1 month ago • 48
FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation Paper • 2608.24168 • Published Aug 25 • 4
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published Sep 8 • 161