Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models Paper • 2607.22098 • Published 5 days ago • 5
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 5
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 7 days ago • 31
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published 6 days ago • 58
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published 13 days ago • 11
AutoIndex: Learning Representation Programs for Retrieval Paper • 2607.18603 • Published 8 days ago • 10
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Paper • 2607.20379 • Published 7 days ago • 6
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Paper • 2607.10169 • Published 18 days ago • 13
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models Paper • 2607.19604 • Published 8 days ago • 16
SLPO: Scaling Latent Reasoning via a Surrogate Policy Paper • 2607.19691 • Published 7 days ago • 5
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation Paper • 2605.29502 • Published May 28 • 1
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes Paper • 2509.24945 • Published Sep 29, 2025 • 7
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training Paper • 2607.19058 • Published 8 days ago • 6
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Paper • 2607.18110 • Published 9 days ago • 14