One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation Paper • 2608.25936 • Published 16 days ago • 13
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 8 days ago • 91
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning Paper • 2609.00638 • Published 10 days ago • 72
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall Paper • 2609.01532 • Published 10 days ago • 9
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published 14 days ago • 30
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 11 days ago • 151
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 29 days ago • 15
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 113
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 63
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 106
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published Jul 22 • 36
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness Paper • 2607.19322 • Published Jul 21 • 11
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published Jul 8 • 16
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models Paper • 2607.05803 • Published Jul 7 • 11
LLM-as-a-Verifier: A General-Purpose Verification Framework Paper • 2607.05391 • Published Jul 6 • 19