Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 26 days ago • 81
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Paper • 2607.03451 • Published 26 days ago • 33
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published 22 days ago • 18
Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding Paper • 2607.05722 • Published 22 days ago • 13
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models Paper • 2607.05803 • Published 22 days ago • 10
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages Paper • 2607.05992 • Published 22 days ago • 6
Attending to Multimodal Generation One Token at a Time Paper • 2607.03738 • Published 25 days ago • 10
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published 21 days ago • 88
Automating the Design of Embodied Agent Architectures Paper • 2606.30111 • Published 26 days ago • 13
Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo Paper • 2607.04939 • Published 23 days ago • 5
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published 21 days ago • 26
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Paper • 2607.07386 • Published 21 days ago • 12
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published 25 days ago • 76
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes Paper • 2607.04439 • Published 24 days ago • 63
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published 27 days ago • 238
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 113
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Paper • 2606.31315 • Published 29 days ago • 77
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning Paper • 2607.04425 • Published 24 days ago • 72
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 241
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 177
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 210
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding Paper • 2605.29707 • Published May 28 • 152
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published Jun 27 • 150
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Paper • 2606.13681 • Published Jun 11 • 143
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 130
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 124
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Paper • 2606.16140 • Published Jun 15 • 124
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 123
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 111
OCC-RAG: Optimal Cognitive Core for Faithful Question Answering Paper • 2606.00683 • Published May 30 • 102
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation Paper • 2607.08758 • Published 20 days ago • 41
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition Paper • 2601.16211 • Published 27 days ago • 54
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Paper • 2607.08768 • Published 20 days ago • 34
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published 21 days ago • 24
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published 21 days ago • 16
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Paper • 2607.06987 • Published 21 days ago • 9
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published 24 days ago • 11
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents Paper • 2607.08716 • Published 20 days ago • 15
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs Paper • 2503.16163 • Published Mar 20, 2025
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect Paper • 2403.03853 • Published Mar 6, 2024 • 65
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios Paper • 2603.09983 • Published Feb 12 • 3
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models Paper • 2505.24133 • Published May 30, 2025 • 2
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 20 days ago • 76
Scalable Visual Pretraining for Language Intelligence Paper • 2607.09657 • Published 19 days ago • 57
KronQ: LLM Quantization via Kronecker-Factored Hessian Paper • 2607.07964 • Published 21 days ago • 32
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning Paper • 2607.08393 • Published 20 days ago • 18
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published 21 days ago • 138
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published 14 days ago • 85
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Paper • 2605.30265 • Published May 28 • 23
LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives Paper • 2607.00784 • Published 28 days ago • 2
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Paper • 2606.27313 • Published Jun 25 • 38
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Paper • 2605.22012 • Published May 21 • 46
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Paper • 2606.30168 • Published about 1 month ago
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement Paper • 2606.23539 • Published Jun 22
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Paper • 2605.18160 • Published Jun 2
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Paper • 2607.11849 • Published 16 days ago • 33
Metacognition in LLMs: Foundations, Progress, and Opportunities Paper • 2607.11881 • Published 16 days ago • 29
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 16 days ago • 84
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published 20 days ago • 35
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published 16 days ago • 23
MuScriptor: An Open Model for Multi-Instrument Music Transcription Paper • 2607.08168 • Published 20 days ago • 21
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper • 2607.07769 • Published 21 days ago • 10
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness Paper • 2607.08046 • Published 20 days ago • 14
TernaryLM: Memory-Efficient Language Modeling via Native 1-Bit Quantization with Adaptive Layer-wise Scaling Paper • 2602.07374 • Published Feb 7 • 2
Tequila: Trapping-free Ternary Quantization for Large Language Models Paper • 2509.23809 • Published Sep 28, 2025 • 4
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification Paper • 2601.07892 • Published Jan 12 • 4
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 15 days ago • 227
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Paper • 2607.12395 • Published 15 days ago • 99
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published 14 days ago • 80
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 13 days ago • 71
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 13 days ago • 103
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 13 days ago • 203
Spectral Rewiring for Exploration, Purification, and Model Merging Paper • 2607.03065 • Published 26 days ago • 25
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Paper • 2607.15273 • Published 13 days ago • 17
GRASP: GRanularity-Aware Search Policy for Agentic RAG Paper • 2607.10463 • Published 18 days ago • 9
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Paper • 2607.15277 • Published 13 days ago • 9
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Paper • 2607.11683 • Published 16 days ago • 146
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 13 days ago • 141
From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality Paper • 2607.13196 • Published 15 days ago • 29
Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 12 days ago • 28
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Paper • 2607.14614 • Published 13 days ago • 12
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 12 days ago • 12
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published 9 days ago • 77
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 12 days ago • 43
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints Paper • 2607.18144 • Published 9 days ago • 15
Environment-free Synthetic Data Generation for API-Calling Agents Paper • 2607.16900 • Published 11 days ago • 20
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Paper • 2607.19139 • Published 8 days ago • 73
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 11 days ago • 137
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Paper • 2607.18722 • Published 8 days ago • 35
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models Paper • 2607.01774 • Published 9 days ago • 38
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 7 days ago • 63
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 7 days ago • 31
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models Paper • 2607.19604 • Published 8 days ago • 16
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 6 days ago • 147
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 6 days ago • 59
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 6 days ago • 25
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published 7 days ago • 31
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 7 days ago • 29
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 5 days ago • 35
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 49
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 6 days ago • 21
LAMAR: An Open Language-Aware Multilingual Alignment Reranker Paper • 2607.22042 • Published 5 days ago • 11
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Paper • 2607.22375 • Published 5 days ago • 7
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making Paper • 2607.14277 • Published 14 days ago • 8
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 2 days ago • 67
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published 2 days ago • 63
LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF Image-Text-to-Text • 35B • Updated about 7 hours ago • 99.7k • 196