FinanceHarness: Autonomous Financial Deep Research Framework Paper • 2607.27853 • Published 18 days ago • 11
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models Paper • 2605.06597 • Published May 7 • 16
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Paper • 2509.21882 • Published Sep 26, 2025
Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning Paper • 2509.11420 • Published Sep 14, 2025 • 4
TradingAgents: Multi-Agents LLM Financial Trading Framework Paper • 2412.20138 • Published Dec 28, 2024 • 125
Advancing AI-Scientist Understanding: Making LLM Think Like a Physicist with Interpretable Reasoning Paper • 2504.01911 • Published Apr 2, 2025 • 1
Scito2M: A 2 Million, 30-Year Cross-disciplinary Dataset for Temporal Scientometric Analysis Paper • 2410.09510 • Published Oct 12, 2024 • 1
Memorize and Rank: Elevating Large Language Models for Clinical Diagnosis Prediction Paper • 2501.17326 • Published Jan 28, 2025
AgentReview: Exploring Peer Review Dynamics with LLM Agents Paper • 2406.12708 • Published Jun 18, 2024 • 8
Benchmarking Foundation Models with Language-Model-as-an-Examiner Paper • 2306.04181 • Published Jun 7, 2023
Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web Agents Paper • 2508.01858 • Published Aug 3, 2025 • 20
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts Paper • 2407.04973 • Published Jul 6, 2024
ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding Paper • 2408.11363 • Published Aug 21, 2024
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories Paper • 2502.06111 • Published Feb 10, 2025
Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach Paper • 2505.18882 • Published May 24, 2025 • 15
ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding Paper • 2408.11363 • Published Aug 21, 2024
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts Paper • 2407.04973 • Published Jul 6, 2024